Text Classification
setfit
Safetensors
English
mpnet
education
classroom-discourse
edubehaviors
assertion
Eval Results (legacy)
Instructions to use StanfordSCALE/assertion_sentence_has_time_reference with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- setfit
How to use StanfordSCALE/assertion_sentence_has_time_reference with setfit:
from setfit import SetFitModel model = SetFitModel.from_pretrained("StanfordSCALE/assertion_sentence_has_time_reference") preds = model.predict(["i loved the spiderman movie!", "pineapple on pizza is the worst"]) print(preds) - Notebooks
- Google Colab
- Kaggle
|
Download README.md from StanfordSCALE/assertion_sentence_has_time_reference: direct link, hf CLI and curl.
- Browser
- Download file 4.09 kB
-
https://huggingface.co/StanfordSCALE/assertion_sentence_has_time_reference/resolve/main/README.md
- Command line
-
hf download hf://StanfordSCALE/assertion_sentence_has_time_reference/README.md
-
curl -L -o README.md https://huggingface.co/StanfordSCALE/assertion_sentence_has_time_reference/resolve/main/README.md
4.09 kB
| language: en | |
| library_name: setfit | |
| pipeline_tag: text-classification | |
| base_model: sentence-transformers/paraphrase-mpnet-base-v2 | |
| datasets: | |
| - StanfordSCALE/assertions_llm_annotated_talkmoves | |
| metrics: | |
| - f1 | |
| - precision | |
| - recall | |
| - roc_auc | |
| tags: | |
| - setfit | |
| - text-classification | |
| - education | |
| - classroom-discourse | |
| - edubehaviors | |
| - assertion | |
| model-index: | |
| - name: assertion_sentence_has_time_reference | |
| results: | |
| - task: | |
| type: text-classification | |
| dataset: | |
| name: assertions_llm_annotated_talkmoves | |
| type: StanfordSCALE/assertions_llm_annotated_talkmoves | |
| split: test | |
| metrics: | |
| - type: f1 | |
| value: 0.728 | |
| name: F1 (positive class, test) | |
| - type: precision | |
| value: 0.7154 | |
| name: Precision (positive class, test) | |
| - type: recall | |
| value: 0.741 | |
| name: Recall (positive class, test) | |
| - type: roc_auc | |
| value: 0.9241 | |
| name: ROC-AUC (test) | |
| # Assertion: sentence has time reference | |
| This classifier was trained for *EduBehaviors: Assertion-based schemas for auditable dialogue coding* and is usable through the Python package `EduBehaviors-kit`. This classifier was trained on an LLM-annotated subset of teacher utterances from the [TalkMoves Dataset](https://github.com/SumnerLab/TalkMoves). See the Datasets section below for more information. | |
| --- | |
| ## Training Details | |
| ### Datasets | |
| | Dataset | Split | Size | | |
| |---------|-------|------| | |
| | [StanfordSCALE/assertions_llm_annotated_talkmoves](https://huggingface.co/datasets/StanfordSCALE/assertions_llm_annotated_talkmoves) | train | 3,432 (53.3%) | | |
| | [StanfordSCALE/assertions_llm_annotated_talkmoves](https://huggingface.co/datasets/StanfordSCALE/assertions_llm_annotated_talkmoves) | dev | 856 (13.3%) | | |
| | [StanfordSCALE/assertions_llm_annotated_talkmoves](https://huggingface.co/datasets/StanfordSCALE/assertions_llm_annotated_talkmoves) | test | 2,146 (33.4%) | | |
| This model's columns are `assertion_sentence_has_time_reference` and `split_sentence_has_time_reference`. | |
| **Base rate** (share of rows labeled as True): 17.3% overall — | |
| 17.2% train, 19.2% dev, 16.9% test. | |
| ### Labels and annotation | |
| Labels were generated with LLM annotators. Krippendorff's alpha for this assertion is 0.603. | |
| ### Hyperparameters | |
| | Parameter | Value | | |
| |-----------|-------| | |
| | Base model (body) | [sentence-transformers/paraphrase-mpnet-base-v2](https://huggingface.co/sentence-transformers/paraphrase-mpnet-base-v2) | | |
| | Head | `LogisticRegression` | | |
| | Body learning rate | 2e-05 | | |
| | Head learning rate | 0.01 | | |
| | Batch size | 16 (contrastive phase) / 32 (head) | | |
| | Epochs | 10 | | |
| | Max steps | 5000 (contrastive phase) | | |
| | Eval max steps | 100 | | |
| | Seed | 20260904 | | |
| | Mixed precision | enabled on GPU | | |
| --- | |
| ## Evaluation | |
| ### Results | |
| | Split | n | Base rate | Precision | Recall | F1 (positive class) | ROC-AUC | Average precision | | |
| |-------|---|-----------|-----------|--------|---------------------|---------|-------------------| | |
| | dev | 856 | 19.2% | 0.687 | 0.695 | 0.691 | 0.888 | 0.704 | | |
| | test | 2,146 | 16.9% | 0.715 | 0.741 | 0.728 | 0.924 | 0.792 | | |
| ### Limitations | |
| - **Labels come from LLM annotators, not human coders.** Agreement between annotators with Krippendorff's Alpha is 0.603. | |
| - Trained on teacher utterances only. Behaviour on student speech is untested. | |
| --- | |
| ## How to Use | |
| ### Message Structure | |
| The model was trained on text built as: | |
| ``` | |
| {utterance} | |
| ``` | |
| The utterance is passed through as-is. | |
| ### Running instructions | |
| ```bash | |
| pip install setfit | |
| ``` | |
| ```python | |
| from setfit import SetFitModel | |
| model = SetFitModel.from_pretrained("StanfordSCALE/assertion_sentence_has_time_reference") | |
| text = 'Happy Friday' | |
| model.predict([text]) # -> array([1]) when the assertion holds | |
| model.predict_proba([text]) # -> [[P(no), P(yes)]] | |
| ``` | |
| ## Citation | |
| ```bibtex | |
| @misc{assertion_sentence_has_time_reference, | |
| author = {Stanford SCALE Initiative}, | |
| title = {Assertion classifier: sentence has time reference}, | |
| year = {2026}, | |
| url = {https://huggingface.co/StanfordSCALE/assertion_sentence_has_time_reference} | |
| } | |
| ``` | |