File size: 4,089 Bytes
9d398c0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
---
language: en
library_name: setfit
pipeline_tag: text-classification
base_model: sentence-transformers/paraphrase-mpnet-base-v2
datasets:
  - StanfordSCALE/assertions_llm_annotated_talkmoves
metrics:
  - f1
  - precision
  - recall
  - roc_auc
tags:
  - setfit
  - text-classification
  - education
  - classroom-discourse
  - edubehaviors
  - assertion
model-index:
- name: assertion_sentence_has_time_reference
  results:
  - task:
      type: text-classification
    dataset:
      name: assertions_llm_annotated_talkmoves
      type: StanfordSCALE/assertions_llm_annotated_talkmoves
      split: test
    metrics:
    - type: f1
      value: 0.728
      name: F1 (positive class, test)
    - type: precision
      value: 0.7154
      name: Precision (positive class, test)
    - type: recall
      value: 0.741
      name: Recall (positive class, test)
    - type: roc_auc
      value: 0.9241
      name: ROC-AUC (test)

---

# Assertion: sentence has time reference

This classifier was trained for *EduBehaviors: Assertion-based schemas for auditable dialogue coding* and is usable through the Python package `EduBehaviors-kit`. This classifier was trained on an LLM-annotated subset of teacher utterances from the [TalkMoves Dataset](https://github.com/SumnerLab/TalkMoves). See the Datasets section below for more information.

---

## Training Details

### Datasets

| Dataset | Split | Size |
|---------|-------|------|
| [StanfordSCALE/assertions_llm_annotated_talkmoves](https://huggingface.co/datasets/StanfordSCALE/assertions_llm_annotated_talkmoves) | train | 3,432 (53.3%) |
| [StanfordSCALE/assertions_llm_annotated_talkmoves](https://huggingface.co/datasets/StanfordSCALE/assertions_llm_annotated_talkmoves) | dev | 856 (13.3%) |
| [StanfordSCALE/assertions_llm_annotated_talkmoves](https://huggingface.co/datasets/StanfordSCALE/assertions_llm_annotated_talkmoves) | test | 2,146 (33.4%) |

This model's columns are `assertion_sentence_has_time_reference` and `split_sentence_has_time_reference`.

**Base rate** (share of rows labeled as True): 17.3% overall —
17.2% train, 19.2% dev, 16.9% test.

### Labels and annotation

Labels were generated with LLM annotators. Krippendorff's alpha for this assertion is 0.603.


### Hyperparameters

| Parameter | Value |
|-----------|-------|
| Base model (body) | [sentence-transformers/paraphrase-mpnet-base-v2](https://huggingface.co/sentence-transformers/paraphrase-mpnet-base-v2) |
| Head | `LogisticRegression` |
| Body learning rate | 2e-05 |
| Head learning rate | 0.01 |
| Batch size | 16 (contrastive phase) / 32 (head) |
| Epochs | 10 |
| Max steps | 5000 (contrastive phase) |
| Eval max steps | 100 |
| Seed | 20260904 |
| Mixed precision | enabled on GPU |

---

## Evaluation

### Results


| Split | n | Base rate | Precision | Recall | F1 (positive class) | ROC-AUC | Average precision |
|-------|---|-----------|-----------|--------|---------------------|---------|-------------------|
| dev | 856 | 19.2% | 0.687 | 0.695 | 0.691 | 0.888 | 0.704 |
| test | 2,146 | 16.9% | 0.715 | 0.741 | 0.728 | 0.924 | 0.792 |

### Limitations

- **Labels come from LLM annotators, not human coders.** Agreement between annotators with Krippendorff's Alpha is 0.603.
- Trained on teacher utterances only. Behaviour on student speech is untested.


---

## How to Use

### Message Structure

The model was trained on text built as:

```
{utterance}
```

The utterance is passed through as-is.

### Running instructions

```bash
pip install setfit
```

```python
from setfit import SetFitModel

model = SetFitModel.from_pretrained("StanfordSCALE/assertion_sentence_has_time_reference")

text = 'Happy Friday'
model.predict([text])        # -> array([1]) when the assertion holds
model.predict_proba([text])  # -> [[P(no), P(yes)]]
```


## Citation

```bibtex
@misc{assertion_sentence_has_time_reference,
  author = {Stanford SCALE Initiative},
  title  = {Assertion classifier: sentence has time reference},
  year   = {2026},
  url    = {https://huggingface.co/StanfordSCALE/assertion_sentence_has_time_reference}
}
```