Instructions to use ozonetg/bargein-classifier-modernbert-large with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ozonetg/bargein-classifier-modernbert-large with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="ozonetg/bargein-classifier-modernbert-large")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("ozonetg/bargein-classifier-modernbert-large") model = AutoModelForSequenceClassification.from_pretrained("ozonetg/bargein-classifier-modernbert-large", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Barge-in classifier · ModernBERT-large (comparison baseline)
Comparison baseline, not recommended. Use bargein-classifier-ettin-400m instead. It has the same architecture, the same training recipe, the same data and the same speed, and it is better on unseen business types, on shifted calls and on telling a caller who is leaving from one who says goodbye and comes back.
This model answers one question: with everything else fixed, does the pre-trained encoder matter? Ettin uses the ModernBERT architecture with different pre-training data, so ModernBERT-large was fine-tuned with the exact recipe of Ettin-400m; only the base checkpoint changed. It is published so the comparison can be checked and reproduced.
Ettin-400m (recommended) · Live demo · Collection
| ModernBERT-large (this model) | Ettin-400m (recommended) | |
|---|---|---|
| Unseen business types (12,876 exchanges, 2 seeds, paired) | 97.28 (−0.24) | 97.53 |
| Distribution shift, OOD-10k (10,888) | 91.28 (−0.47, p < 0.001) | 91.75 |
| Leaving vs. goodbye-and-back (123 = 41 × 3 seeds) | 87.8% | 97.6% |
| Human-labelled exchanges (158) | 96.73 | 97.47 |
| Held-out business types A (593) | 96.07 | 95.39 |
| Held-out business types B (588) | 94.16 | 93.99 |
| TTS → phone channel → ASR (1,328) | 92.22 | 91.49 |
| Real calls: agreement with the LLM judge (1,425; seed 13) | 79.4 (−0.8, CI −1.9…+0.3) | 80.2 |
| Latency, GPU / CPU p50 (same run) | 8.4 / 112 ms | 8.4 / 112 ms |
The task, inputs, labels, training data and limitations are the same as for Ettin-400m. They are described in full on its model card.
The task in brief
When a caller starts talking while a voice agent is still speaking, the model reads what the agent already said
(agent_said), the rest of its planned line (agent_unsaid) and the caller's ASR words (caller_said). It returns
P(INTERRUPT): INTERRUPT (0.5 or above) means the agent stops and the LLM takes the turn; CONTINUE means it keeps
talking. Inputs must be built with format_input.py, the training-time formatter, and an empty caller
transcript should be treated as CONTINUE without calling the model.
Quickstart
pip install torch "transformers>=4.51" huggingface_hub
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("ozonetg/bargein-classifier-modernbert-large")
sys.path.insert(0, path)
from predict import BargeInClassifier # predict.py ships in this repo
clf = BargeInClassifier(path) # GPU (bf16) if available, else CPU (fp32)
clf.predict(agent_said="Your appointment is on Tuesday at",
agent_unsaid="three pm with Dr. Lee. Does that still work?",
caller_said="wait which doctor")
# {'label': 'INTERRUPT', 'p_interrupt': 0.99..., 'model_called': True}
On CPU fp32, transformers 4.51, 4.57 and 5.17 give identical outputs (0 decision changes on 593 test exchanges).
Head-to-head with Ettin-400m
How the comparison was run
- Same recipe: full fine-tune, lr 2e-5, dropout 0.1 plus R-Drop (α = 1.0), 5 epochs, batch 32, max length 192, mean pooling, the same training file and split. Only the model id differs.
- Same tests and a decision rule fixed in advance:
- unseen business types, 2 seeds, paired against the same-seed Ettin runs;
- all data, 3 seeds (13, 14 and 15, each on 95% of the data), against the Ettin recipe's own 3 seeds;
- one published model per size (seed 13, 99.5% of the data), timed and scored on real calls.
- Checked scoring: the scoring script reproduced all 69 published Ettin reference numbers before scoring anything new.
- p-values: one-sided paired bootstrap, in the direction of the observed difference.
Unseen business types (12,876 exchanges from 33 business types held out of training)
| model | accuracy (mean of 2 seeds) | log-loss |
|---|---|---|
| Ettin-400m | 97.53 | 0.143 |
| ModernBERT-large | 97.28 | 0.153 |
Δ = −0.24 points for ModernBERT-large; the one-sided p-value for a gain is 0.997, so there is no gain. It is also less well calibrated (higher log-loss).
All data, mean of 3 seeds
| test | Ettin-400m | ModernBERT-large | Δ |
|---|---|---|---|
| Human-labelled exchanges (158) | 97.47 | 96.73 | −0.74, p = 0.107 |
| Held-out business types A (593) | 95.39 | 96.07 | +0.67, p = 0.060 |
| Held-out business types B (588) | 93.99 | 94.16 | +0.17, p = 0.318 |
| LLM-written hard cases (150) | 90.00 | 91.33 | +1.33, p = 0.153 |
| Perturbation stress test, accuracy (6,258) | 93.83 | 94.06 | +0.22, p = 0.070 |
| Perturbation stress test, flip rate (lower is better) | 2.10 | 2.07 | −0.03 |
| TTS → phone channel → ASR (1,328) | 91.49 | 92.22 | +0.73, p = 0.012 |
| Distribution shift, OOD-10k (10,888) | 91.75 | 91.28 | −0.47, p < 0.001 |
| Leaving vs. goodbye-and-back (123) | 97.6% | 87.8% | −9.8 |
"LLM-written hard cases" are 150 difficult exchanges written by an LLM; read them with care. Where ModernBERT-large gains, it is on test sets generated the same way as the training data; on shifted calls and on leave vs. goodbye it is behind.
Distribution shift by type (OOD-10k, mean of 3 seeds)
| shift | Ettin-400m | ModernBERT-large | Δ |
|---|---|---|---|
| New business types | 95.56 | 95.38 | −0.18, p = 0.315 |
| Locale | 92.36 | 92.33 | −0.03, p = 0.481 |
| Agent style | 94.71 | 94.22 | −0.49, p = 0.108 |
| Caller population | 90.92 | 90.40 | −0.52, p = 0.133 |
| ASR error profile | 92.20 | 92.01 | −0.18, p = 0.342 |
| Call phase | 95.58 | 95.24 | −0.34, p = 0.164 |
| Timing extremes | 95.77 | 96.04 | +0.27, p = 0.171 |
| Hard meanings | 91.12 | 90.94 | −0.18, p = 0.326 |
| Third party | 80.73 | 79.08 | −1.65, p = 0.007 |
| Mixed | 88.53 | 87.18 | −1.35, p = 0.003 |
Hard sub-types (pooled over 3 seeds)
| sub-type | Ettin-400m | ModernBERT-large | Δ |
|---|---|---|---|
| Interpreter relays the caller's answer (68) | 62.3 | 62.7 | +0.5 |
| Gatekeeper says they will transfer (56) | 75.6 | 63.7 | −11.9 |
| Recogniser clipped the first or last word (143) | 91.6 | 90.0 | −1.6 |
| Recogniser dropped small words (131) | 92.1 | 92.4 | +0.3 |
The published weights (seed 13, one draw each)
| model | human labels (158) | held-out A | held-out B | LLM-written hard | stress acc / flip | TTS → ASR | OOD-10k | leave vs. goodbye (41) |
|---|---|---|---|---|---|---|---|---|
| Ettin-400m (recommended) | 98.10 | 95.11 | 93.71 | 90.67 | 93.59 / 1.98 | 91.11 | 91.84 | 40/41 |
| ModernBERT-large (this model) | 97.47 | 96.29 | 94.05 | 90.00 | 94.23 / 1.82 | 91.87 | 90.94 | 28/41 |
A single model is one draw: the comparison above uses 3-seed means and paired tests, never this table.
Real calls
The same 2,525 barge-in moments from 1,477 real outbound sales calls from a single campaign used on the Ettin-400m card, with the same references (decision at 0.5; 95% bootstrap intervals over calls; Δ is paired on the same resampled calls).
| reference | Ettin-400m | ModernBERT-large | Δ (ModernBERT-large − Ettin-400m) |
|---|---|---|---|
| Rule-unambiguous moments with words (278) | 98.2 (96.0–99.7) | 97.8 (96.0–99.3) | – |
| Agreement with the LLM judge where it is confident, without the opener "hello" (945) | 94.4 (92.7–95.8) | 94.1 (92.4–95.5) | −0.3 (−1.5…+0.8) |
| Agreement with the judge, without the opener "hello" (1,262) | 87.7 (85.8–89.5) | 86.8 (84.8–88.7) | – |
| Agreement with the judge, all moments with words (1,425) | 80.2 (78.2–82.2) | 79.4 (77.4–81.5) | −0.8 (−1.9…+0.3) |
| Stops the agent on an empty transcript (1,100; lower is better) | 10.2 (8.4–12.0) | 11.4 (9.4–13.3) | +1.2 (−0.1…+2.4) |
| Stops the agent on a lone "hello" during the opener (163) | 79.1 (72.7–85.4) | 79.1 (72.7–85.4) | +0.0 (+0.0…+0.0) |
Speed
| model | GPU bf16, p50 (p95) | CPU fp32, 4 threads, p50 (p95) |
|---|---|---|
| Ettin-400m | 8.4 (9.4) ms | 112 (122) ms |
| ModernBERT-large | 8.4 (9.4) ms | 112 (123) ms |
Batch 1, end to end, one run with exclusive use of the GPU (RTX 6000 Ada; Intel Xeon Gold 5412U CPU), all models side by side. Same architecture and size, same speed.
Decision rule
A baseline would have replaced Ettin-400m only if it passed every check:
| check | ModernBERT-large |
|---|---|
| Unseen business types: gain ≥ +0.15 with p < 0.05 | ✗ (−0.24) |
| Human labels not worse (Δ ≥ −0.63, one case) | ✗ (−0.738) |
| Hard-set mix not worse (held-out B and LLM-written cases, Δ ≥ −0.3) | ✓ (+0.75) |
| Perturbation stress not worse | ✓ (+0.22 / flip −0.03) |
| Distribution shift not worse (Δ ≥ −0.2) | ✗ (−0.47) |
| No shift type worse by more than 1.0 | ✗ (worst: third party −1.65) |
| Leave vs. goodbye within 3 points | ✗ (87.8% vs 97.6%) |
| Hard meanings not worse by more than 1.0 | ✓ (−0.18) |
| Latency within 1.2× | ✓ (1.00× GPU, 1.00× CPU) |
Verdict
At identical speed, Ettin transfers better: it wins on unseen business types, on shifted calls (third party and mixed shifts most of all) and on leave vs. goodbye-and-back. ModernBERT-large gains only on test sets generated the same way as the training data, and ties on real calls.
The backbones share an architecture, so the difference is the pre-training data. One reading, not tested here, is that Ettin's pre-training transfers slightly better to calls unlike the training data. Either way, Ettin-400m is the model to use.
Training
| Base model | answerdotai/ModernBERT-large (28 layers, hidden 1024, 395M parameters) |
| Head | sequence classification (mean pooling), 2 labels |
| Fine-tuning | full fine-tune, AdamW (weight decay 0.01), lr 2e-5, linear schedule with 6% warm-up, batch 32, 5 epochs, last checkpoint |
| Regularisation | dropout 0.1 on embeddings, attention, MLP and head, plus R-Drop with α = 1.0 |
| Precision / length | bf16 autocast over fp32 master weights; max length 192 tokens |
| RoPE | ModernBERT's own values are kept: θ = 160,000 for global and 10,000 for local attention (Ettin uses 160,000 for both) |
| Data | the Ettin-400m training set: 64,151 synthetic exchanges, plus 323 held back as a dev monitor |
| Compute | one NVIDIA RTX 6000 Ada, 1.5 h (5,439 s), peak 17.9 GB |
The data is synthetic: LLM-generated multi-business exchanges, labelled by an LLM judge whose rubric was calibrated against about 325 human labels. See the Ettin-400m card for the full pipeline.
Limitations
Everything listed on the Ettin-400m card applies here too, and:
- Leave vs. goodbye-and-back is weaker (87.8% over 3 seeds against 97.6% for Ettin-400m).
- A lone "hello" while the agent's opener plays makes it stop on 79% of real pickups.
- It is a baseline. It was trained and published to measure the effect of the base checkpoint, not for deployment.
Citation
@misc{ozonetg2026bargein,
title = {Barge-in classifier for voice agents},
author = {ozonetg},
year = {2026},
howpublished = {\url{https://huggingface.co/ozonetg/bargein-classifier-ettin-400m}}
}
@misc{modernbert,
title = {Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference},
author = {Benjamin Warner and Antoine Chaffin and Benjamin Clavié and Orion Weller and Oskar Hallström and Said Taghadouini and Alexis Gallagher and Raja Biswas and Faisal Ladhak and Tom Aarsen and Nathan Cooper and Griffin Adams and Jeremy Howard and Iacopo Poli},
year = {2024},
eprint = {2412.13663},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2412.13663}
}
@inproceedings{liang2021rdrop,
title = {R-Drop: Regularized Dropout for Neural Networks},
author = {Xiaobo Liang and Lijun Wu and Juntao Li and Yue Wang and Qi Meng and Tao Qin and Wei Chen and Min Zhang and Tie-Yan Liu},
booktitle = {Advances in Neural Information Processing Systems},
volume = {34},
year = {2021},
url = {https://arxiv.org/abs/2106.14448}
}
Related
- bargein-classifier-ettin-400m: the recommended model of this size.
- bargein-classifier-modernbert-base: the other ModernBERT baseline.
- Live demo: try the Ettin models in your browser.
- Downloads last month
- 6
Model tree for ozonetg/bargein-classifier-modernbert-large
Base model
answerdotai/ModernBERT-largeCollection including ozonetg/bargein-classifier-modernbert-large
Papers for ozonetg/bargein-classifier-modernbert-large
Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference
R-Drop: Regularized Dropout for Neural Networks
Evaluation results
- Accuracy (threshold 0.5) on Held-out business types A (593 exchanges, 8 unseen business types)self-reported96.290
- Accuracy (threshold 0.5) on Held-out business types B (588 exchanges, 20 unseen business types)self-reported94.050
- Accuracy (threshold 0.5) on Human-labelled exchanges (158, not shown to the labelling judge)self-reported97.470
- Accuracy (threshold 0.5) on Perturbation stress test (6,258)self-reported94.230
- Accuracy (threshold 0.5) on TTS -> phone channel -> ASR stress test (1,328)self-reported91.870
- Accuracy (threshold 0.5) on Distribution-shift test OOD-10k (10,888)self-reported90.940
- Accuracy (threshold 0.5) on Real calls, rule-unambiguous moments with words (278)self-reported97.800