Barge-in classifier · ModernBERT-large (comparison baseline)

Comparison baseline, not recommended. Use bargein-classifier-ettin-400m instead. It has the same architecture, the same training recipe, the same data and the same speed, and it is better on unseen business types, on shifted calls and on telling a caller who is leaving from one who says goodbye and comes back.

This model answers one question: with everything else fixed, does the pre-trained encoder matter? Ettin uses the ModernBERT architecture with different pre-training data, so ModernBERT-large was fine-tuned with the exact recipe of Ettin-400m; only the base checkpoint changed. It is published so the comparison can be checked and reproduced.

Ettin-400m (recommended) · Live demo · Collection

ModernBERT-large (this model) Ettin-400m (recommended)
Unseen business types (12,876 exchanges, 2 seeds, paired) 97.28 (−0.24) 97.53
Distribution shift, OOD-10k (10,888) 91.28 (−0.47, p < 0.001) 91.75
Leaving vs. goodbye-and-back (123 = 41 × 3 seeds) 87.8% 97.6%
Human-labelled exchanges (158) 96.73 97.47
Held-out business types A (593) 96.07 95.39
Held-out business types B (588) 94.16 93.99
TTS → phone channel → ASR (1,328) 92.22 91.49
Real calls: agreement with the LLM judge (1,425; seed 13) 79.4 (−0.8, CI −1.9…+0.3) 80.2
Latency, GPU / CPU p50 (same run) 8.4 / 112 ms 8.4 / 112 ms

The task, inputs, labels, training data and limitations are the same as for Ettin-400m. They are described in full on its model card.

The task in brief

When a caller starts talking while a voice agent is still speaking, the model reads what the agent already said (agent_said), the rest of its planned line (agent_unsaid) and the caller's ASR words (caller_said). It returns P(INTERRUPT): INTERRUPT (0.5 or above) means the agent stops and the LLM takes the turn; CONTINUE means it keeps talking. Inputs must be built with format_input.py, the training-time formatter, and an empty caller transcript should be treated as CONTINUE without calling the model.

Quickstart

pip install torch "transformers>=4.51" huggingface_hub
import sys
from huggingface_hub import snapshot_download

path = snapshot_download("ozonetg/bargein-classifier-modernbert-large")
sys.path.insert(0, path)
from predict import BargeInClassifier       # predict.py ships in this repo

clf = BargeInClassifier(path)                # GPU (bf16) if available, else CPU (fp32)
clf.predict(agent_said="Your appointment is on Tuesday at",
            agent_unsaid="three pm with Dr. Lee. Does that still work?",
            caller_said="wait which doctor")
# {'label': 'INTERRUPT', 'p_interrupt': 0.99..., 'model_called': True}

On CPU fp32, transformers 4.51, 4.57 and 5.17 give identical outputs (0 decision changes on 593 test exchanges).

Head-to-head with Ettin-400m

How the comparison was run

  • Same recipe: full fine-tune, lr 2e-5, dropout 0.1 plus R-Drop (α = 1.0), 5 epochs, batch 32, max length 192, mean pooling, the same training file and split. Only the model id differs.
  • Same tests and a decision rule fixed in advance:
    • unseen business types, 2 seeds, paired against the same-seed Ettin runs;
    • all data, 3 seeds (13, 14 and 15, each on 95% of the data), against the Ettin recipe's own 3 seeds;
    • one published model per size (seed 13, 99.5% of the data), timed and scored on real calls.
  • Checked scoring: the scoring script reproduced all 69 published Ettin reference numbers before scoring anything new.
  • p-values: one-sided paired bootstrap, in the direction of the observed difference.

Unseen business types (12,876 exchanges from 33 business types held out of training)

model accuracy (mean of 2 seeds) log-loss
Ettin-400m 97.53 0.143
ModernBERT-large 97.28 0.153

Δ = −0.24 points for ModernBERT-large; the one-sided p-value for a gain is 0.997, so there is no gain. It is also less well calibrated (higher log-loss).

All data, mean of 3 seeds

test Ettin-400m ModernBERT-large Δ
Human-labelled exchanges (158) 97.47 96.73 −0.74, p = 0.107
Held-out business types A (593) 95.39 96.07 +0.67, p = 0.060
Held-out business types B (588) 93.99 94.16 +0.17, p = 0.318
LLM-written hard cases (150) 90.00 91.33 +1.33, p = 0.153
Perturbation stress test, accuracy (6,258) 93.83 94.06 +0.22, p = 0.070
Perturbation stress test, flip rate (lower is better) 2.10 2.07 −0.03
TTS → phone channel → ASR (1,328) 91.49 92.22 +0.73, p = 0.012
Distribution shift, OOD-10k (10,888) 91.75 91.28 −0.47, p < 0.001
Leaving vs. goodbye-and-back (123) 97.6% 87.8% −9.8

"LLM-written hard cases" are 150 difficult exchanges written by an LLM; read them with care. Where ModernBERT-large gains, it is on test sets generated the same way as the training data; on shifted calls and on leave vs. goodbye it is behind.

Distribution shift by type (OOD-10k, mean of 3 seeds)

shift Ettin-400m ModernBERT-large Δ
New business types 95.56 95.38 −0.18, p = 0.315
Locale 92.36 92.33 −0.03, p = 0.481
Agent style 94.71 94.22 −0.49, p = 0.108
Caller population 90.92 90.40 −0.52, p = 0.133
ASR error profile 92.20 92.01 −0.18, p = 0.342
Call phase 95.58 95.24 −0.34, p = 0.164
Timing extremes 95.77 96.04 +0.27, p = 0.171
Hard meanings 91.12 90.94 −0.18, p = 0.326
Third party 80.73 79.08 −1.65, p = 0.007
Mixed 88.53 87.18 −1.35, p = 0.003

Hard sub-types (pooled over 3 seeds)

sub-type Ettin-400m ModernBERT-large Δ
Interpreter relays the caller's answer (68) 62.3 62.7 +0.5
Gatekeeper says they will transfer (56) 75.6 63.7 −11.9
Recogniser clipped the first or last word (143) 91.6 90.0 −1.6
Recogniser dropped small words (131) 92.1 92.4 +0.3

The published weights (seed 13, one draw each)

model human labels (158) held-out A held-out B LLM-written hard stress acc / flip TTS → ASR OOD-10k leave vs. goodbye (41)
Ettin-400m (recommended) 98.10 95.11 93.71 90.67 93.59 / 1.98 91.11 91.84 40/41
ModernBERT-large (this model) 97.47 96.29 94.05 90.00 94.23 / 1.82 91.87 90.94 28/41

A single model is one draw: the comparison above uses 3-seed means and paired tests, never this table.

Real calls

The same 2,525 barge-in moments from 1,477 real outbound sales calls from a single campaign used on the Ettin-400m card, with the same references (decision at 0.5; 95% bootstrap intervals over calls; Δ is paired on the same resampled calls).

reference Ettin-400m ModernBERT-large Δ (ModernBERT-large − Ettin-400m)
Rule-unambiguous moments with words (278) 98.2 (96.0–99.7) 97.8 (96.0–99.3) –
Agreement with the LLM judge where it is confident, without the opener "hello" (945) 94.4 (92.7–95.8) 94.1 (92.4–95.5) −0.3 (−1.5…+0.8)
Agreement with the judge, without the opener "hello" (1,262) 87.7 (85.8–89.5) 86.8 (84.8–88.7) –
Agreement with the judge, all moments with words (1,425) 80.2 (78.2–82.2) 79.4 (77.4–81.5) −0.8 (−1.9…+0.3)
Stops the agent on an empty transcript (1,100; lower is better) 10.2 (8.4–12.0) 11.4 (9.4–13.3) +1.2 (−0.1…+2.4)
Stops the agent on a lone "hello" during the opener (163) 79.1 (72.7–85.4) 79.1 (72.7–85.4) +0.0 (+0.0…+0.0)

Speed

model GPU bf16, p50 (p95) CPU fp32, 4 threads, p50 (p95)
Ettin-400m 8.4 (9.4) ms 112 (122) ms
ModernBERT-large 8.4 (9.4) ms 112 (123) ms

Batch 1, end to end, one run with exclusive use of the GPU (RTX 6000 Ada; Intel Xeon Gold 5412U CPU), all models side by side. Same architecture and size, same speed.

Decision rule

A baseline would have replaced Ettin-400m only if it passed every check:

check ModernBERT-large
Unseen business types: gain ≥ +0.15 with p < 0.05 ✗ (−0.24)
Human labels not worse (Δ ≥ −0.63, one case) ✗ (−0.738)
Hard-set mix not worse (held-out B and LLM-written cases, Δ ≥ −0.3) ✓ (+0.75)
Perturbation stress not worse ✓ (+0.22 / flip −0.03)
Distribution shift not worse (Δ ≥ −0.2) ✗ (−0.47)
No shift type worse by more than 1.0 ✗ (worst: third party −1.65)
Leave vs. goodbye within 3 points ✗ (87.8% vs 97.6%)
Hard meanings not worse by more than 1.0 ✓ (−0.18)
Latency within 1.2× ✓ (1.00× GPU, 1.00× CPU)

Verdict

At identical speed, Ettin transfers better: it wins on unseen business types, on shifted calls (third party and mixed shifts most of all) and on leave vs. goodbye-and-back. ModernBERT-large gains only on test sets generated the same way as the training data, and ties on real calls.

The backbones share an architecture, so the difference is the pre-training data. One reading, not tested here, is that Ettin's pre-training transfers slightly better to calls unlike the training data. Either way, Ettin-400m is the model to use.

Training

Base model answerdotai/ModernBERT-large (28 layers, hidden 1024, 395M parameters)
Head sequence classification (mean pooling), 2 labels
Fine-tuning full fine-tune, AdamW (weight decay 0.01), lr 2e-5, linear schedule with 6% warm-up, batch 32, 5 epochs, last checkpoint
Regularisation dropout 0.1 on embeddings, attention, MLP and head, plus R-Drop with α = 1.0
Precision / length bf16 autocast over fp32 master weights; max length 192 tokens
RoPE ModernBERT's own values are kept: θ = 160,000 for global and 10,000 for local attention (Ettin uses 160,000 for both)
Data the Ettin-400m training set: 64,151 synthetic exchanges, plus 323 held back as a dev monitor
Compute one NVIDIA RTX 6000 Ada, 1.5 h (5,439 s), peak 17.9 GB

The data is synthetic: LLM-generated multi-business exchanges, labelled by an LLM judge whose rubric was calibrated against about 325 human labels. See the Ettin-400m card for the full pipeline.

Limitations

Everything listed on the Ettin-400m card applies here too, and:

  • Leave vs. goodbye-and-back is weaker (87.8% over 3 seeds against 97.6% for Ettin-400m).
  • A lone "hello" while the agent's opener plays makes it stop on 79% of real pickups.
  • It is a baseline. It was trained and published to measure the effect of the base checkpoint, not for deployment.

Citation

@misc{ozonetg2026bargein,
  title        = {Barge-in classifier for voice agents},
  author       = {ozonetg},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/ozonetg/bargein-classifier-ettin-400m}}
}

@misc{modernbert,
  title         = {Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference},
  author        = {Benjamin Warner and Antoine Chaffin and Benjamin Clavié and Orion Weller and Oskar Hallström and Said Taghadouini and Alexis Gallagher and Raja Biswas and Faisal Ladhak and Tom Aarsen and Nathan Cooper and Griffin Adams and Jeremy Howard and Iacopo Poli},
  year          = {2024},
  eprint        = {2412.13663},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2412.13663}
}

@inproceedings{liang2021rdrop,
  title     = {R-Drop: Regularized Dropout for Neural Networks},
  author    = {Xiaobo Liang and Lijun Wu and Juntao Li and Yue Wang and Qi Meng and Tao Qin and Wei Chen and Min Zhang and Tie-Yan Liu},
  booktitle = {Advances in Neural Information Processing Systems},
  volume    = {34},
  year      = {2021},
  url       = {https://arxiv.org/abs/2106.14448}
}

Related

Downloads last month
6
Safetensors
Model size
0.4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ozonetg/bargein-classifier-modernbert-large

Finetuned
(395)
this model

Collection including ozonetg/bargein-classifier-modernbert-large

Papers for ozonetg/bargein-classifier-modernbert-large

Evaluation results

  • Accuracy (threshold 0.5) on Held-out business types A (593 exchanges, 8 unseen business types)
    self-reported
    96.290
  • Accuracy (threshold 0.5) on Held-out business types B (588 exchanges, 20 unseen business types)
    self-reported
    94.050
  • Accuracy (threshold 0.5) on Human-labelled exchanges (158, not shown to the labelling judge)
    self-reported
    97.470
  • Accuracy (threshold 0.5) on Perturbation stress test (6,258)
    self-reported
    94.230
  • Accuracy (threshold 0.5) on TTS -> phone channel -> ASR stress test (1,328)
    self-reported
    91.870
  • Accuracy (threshold 0.5) on Distribution-shift test OOD-10k (10,888)
    self-reported
    90.940
  • Accuracy (threshold 0.5) on Real calls, rule-unambiguous moments with words (278)
    self-reported
    97.800