Text Classification
setfit
Safetensors
English
mpnet
education
classroom-discourse
edubehaviors
assertion
Eval Results (legacy)
Instructions to use StanfordSCALE/assertion_sentence_has_time_reference with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- setfit
How to use StanfordSCALE/assertion_sentence_has_time_reference with setfit:
from setfit import SetFitModel model = SetFitModel.from_pretrained("StanfordSCALE/assertion_sentence_has_time_reference") preds = model.predict(["i loved the spiderman movie!", "pineapple on pizza is the worst"]) print(preds) - Notebooks
- Google Colab
- Kaggle
File size: 4,089 Bytes
9d398c0 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 | ---
language: en
library_name: setfit
pipeline_tag: text-classification
base_model: sentence-transformers/paraphrase-mpnet-base-v2
datasets:
- StanfordSCALE/assertions_llm_annotated_talkmoves
metrics:
- f1
- precision
- recall
- roc_auc
tags:
- setfit
- text-classification
- education
- classroom-discourse
- edubehaviors
- assertion
model-index:
- name: assertion_sentence_has_time_reference
results:
- task:
type: text-classification
dataset:
name: assertions_llm_annotated_talkmoves
type: StanfordSCALE/assertions_llm_annotated_talkmoves
split: test
metrics:
- type: f1
value: 0.728
name: F1 (positive class, test)
- type: precision
value: 0.7154
name: Precision (positive class, test)
- type: recall
value: 0.741
name: Recall (positive class, test)
- type: roc_auc
value: 0.9241
name: ROC-AUC (test)
---
# Assertion: sentence has time reference
This classifier was trained for *EduBehaviors: Assertion-based schemas for auditable dialogue coding* and is usable through the Python package `EduBehaviors-kit`. This classifier was trained on an LLM-annotated subset of teacher utterances from the [TalkMoves Dataset](https://github.com/SumnerLab/TalkMoves). See the Datasets section below for more information.
---
## Training Details
### Datasets
| Dataset | Split | Size |
|---------|-------|------|
| [StanfordSCALE/assertions_llm_annotated_talkmoves](https://huggingface.co/datasets/StanfordSCALE/assertions_llm_annotated_talkmoves) | train | 3,432 (53.3%) |
| [StanfordSCALE/assertions_llm_annotated_talkmoves](https://huggingface.co/datasets/StanfordSCALE/assertions_llm_annotated_talkmoves) | dev | 856 (13.3%) |
| [StanfordSCALE/assertions_llm_annotated_talkmoves](https://huggingface.co/datasets/StanfordSCALE/assertions_llm_annotated_talkmoves) | test | 2,146 (33.4%) |
This model's columns are `assertion_sentence_has_time_reference` and `split_sentence_has_time_reference`.
**Base rate** (share of rows labeled as True): 17.3% overall —
17.2% train, 19.2% dev, 16.9% test.
### Labels and annotation
Labels were generated with LLM annotators. Krippendorff's alpha for this assertion is 0.603.
### Hyperparameters
| Parameter | Value |
|-----------|-------|
| Base model (body) | [sentence-transformers/paraphrase-mpnet-base-v2](https://huggingface.co/sentence-transformers/paraphrase-mpnet-base-v2) |
| Head | `LogisticRegression` |
| Body learning rate | 2e-05 |
| Head learning rate | 0.01 |
| Batch size | 16 (contrastive phase) / 32 (head) |
| Epochs | 10 |
| Max steps | 5000 (contrastive phase) |
| Eval max steps | 100 |
| Seed | 20260904 |
| Mixed precision | enabled on GPU |
---
## Evaluation
### Results
| Split | n | Base rate | Precision | Recall | F1 (positive class) | ROC-AUC | Average precision |
|-------|---|-----------|-----------|--------|---------------------|---------|-------------------|
| dev | 856 | 19.2% | 0.687 | 0.695 | 0.691 | 0.888 | 0.704 |
| test | 2,146 | 16.9% | 0.715 | 0.741 | 0.728 | 0.924 | 0.792 |
### Limitations
- **Labels come from LLM annotators, not human coders.** Agreement between annotators with Krippendorff's Alpha is 0.603.
- Trained on teacher utterances only. Behaviour on student speech is untested.
---
## How to Use
### Message Structure
The model was trained on text built as:
```
{utterance}
```
The utterance is passed through as-is.
### Running instructions
```bash
pip install setfit
```
```python
from setfit import SetFitModel
model = SetFitModel.from_pretrained("StanfordSCALE/assertion_sentence_has_time_reference")
text = 'Happy Friday'
model.predict([text]) # -> array([1]) when the assertion holds
model.predict_proba([text]) # -> [[P(no), P(yes)]]
```
## Citation
```bibtex
@misc{assertion_sentence_has_time_reference,
author = {Stanford SCALE Initiative},
title = {Assertion classifier: sentence has time reference},
year = {2026},
url = {https://huggingface.co/StanfordSCALE/assertion_sentence_has_time_reference}
}
```
|