Barge-in classifier · ModernBERT-base (comparison baseline)

Comparison baseline, not recommended. Use bargein-classifier-ettin-150m instead. It has the same architecture, the same training recipe, the same data and the same speed, and it is better on unseen business types, on shifted calls and on telling a caller who is leaving from one who says goodbye and comes back.

This model answers one question: with everything else fixed, does the pre-trained encoder matter? Ettin uses the ModernBERT architecture with different pre-training data, so ModernBERT-base was fine-tuned with the exact recipe of Ettin-150m; only the base checkpoint changed. It is published so the comparison can be checked and reproduced.

Ettin-150m (recommended) · Live demo · Collection

ModernBERT-base (this model) Ettin-150m (recommended)
Unseen business types (12,876 exchanges, 2 seeds, paired) 96.87 (−0.05) 96.91
Distribution shift, OOD-10k (10,888) 89.18 (−0.30, p = 0.026) 89.48
Leaving vs. goodbye-and-back (123 = 41 × 3 seeds) 78.0% 83.7%
Human-labelled exchanges (158) 95.99 96.62
Held-out business types A (593) 94.66 95.62
Held-out business types B (588) 93.31 92.80
TTS → phone channel → ASR (1,328) 90.84 91.42
Real calls: agreement with the LLM judge (1,425; seed 13) 80.7 (−2.7, CI −4.2…−1.1) 83.4
Latency, GPU / CPU p50 (same run) 6.9 / 37.1 ms 6.9 / 37.5 ms

The task, inputs, labels, training data and limitations are the same as for Ettin-150m. They are described in full on its model card.

The task in brief

When a caller starts talking while a voice agent is still speaking, the model reads what the agent already said (agent_said), the rest of its planned line (agent_unsaid) and the caller's ASR words (caller_said). It returns P(INTERRUPT): INTERRUPT (0.5 or above) means the agent stops and the LLM takes the turn; CONTINUE means it keeps talking. Inputs must be built with format_input.py, the training-time formatter, and an empty caller transcript should be treated as CONTINUE without calling the model.

Quickstart

pip install torch "transformers>=4.51" huggingface_hub
import sys
from huggingface_hub import snapshot_download

path = snapshot_download("ozonetg/bargein-classifier-modernbert-base")
sys.path.insert(0, path)
from predict import BargeInClassifier       # predict.py ships in this repo

clf = BargeInClassifier(path)                # GPU (bf16) if available, else CPU (fp32)
clf.predict(agent_said="Your appointment is on Tuesday at",
            agent_unsaid="three pm with Dr. Lee. Does that still work?",
            caller_said="wait which doctor")
# {'label': 'INTERRUPT', 'p_interrupt': 0.99..., 'model_called': True}

On CPU fp32, transformers 4.51, 4.57 and 5.17 give identical outputs (0 decision changes on 593 test exchanges).

Head-to-head with Ettin-150m

How the comparison was run

  • Same recipe: full fine-tune, lr 5e-5, dropout 0.02 plus R-Drop (α = 1.0), 5 epochs, batch 32, max length 192, mean pooling, the same training file and split. Only the model id differs.
  • Same tests and a decision rule fixed in advance:
    • unseen business types, 2 seeds, paired against the same-seed Ettin runs;
    • all data, 3 seeds (13, 14 and 15, each on 95% of the data), against the Ettin recipe's own 3 seeds;
    • one published model per size (seed 13, 99.5% of the data), timed and scored on real calls.
  • Checked scoring: the scoring script reproduced all 69 published Ettin reference numbers before scoring anything new.
  • p-values: one-sided paired bootstrap, in the direction of the observed difference.

Unseen business types (12,876 exchanges from 33 business types held out of training)

model accuracy (mean of 2 seeds) log-loss
Ettin-150m 96.91 0.235
ModernBERT-base 96.87 0.251

Δ = −0.05 points for ModernBERT-base; the one-sided p-value for a gain is 0.681, so there is no gain. It is also less well calibrated (higher log-loss).

All data, mean of 3 seeds

test Ettin-150m ModernBERT-base Δ
Human-labelled exchanges (158) 96.62 95.99 −0.63, p = 0.129
Held-out business types A (593) 95.62 94.66 −0.96, p = 0.015
Held-out business types B (588) 92.80 93.31 +0.51, p = 0.161
LLM-written hard cases (150) 88.22 87.33 −0.89, p = 0.208
Perturbation stress test, accuracy (6,258) 92.52 92.36 −0.16, p = 0.171
Perturbation stress test, flip rate (lower is better) 2.56 2.50 −0.06
TTS → phone channel → ASR (1,328) 91.42 90.84 −0.58, p = 0.066
Distribution shift, OOD-10k (10,888) 89.48 89.18 −0.30, p = 0.026
Leaving vs. goodbye-and-back (123) 83.7% 78.0% −5.7

"LLM-written hard cases" are 150 difficult exchanges written by an LLM; read them with care. Where ModernBERT-base gains, it is on test sets generated the same way as the training data; on shifted calls and on leave vs. goodbye it is behind.

Distribution shift by type (OOD-10k, mean of 3 seeds)

shift Ettin-150m ModernBERT-base Δ
New business types 93.88 93.63 −0.24, p = 0.278
Locale 90.15 89.29 −0.86, p = 0.029
Agent style 93.80 93.07 −0.73, p = 0.027
Caller population 88.77 87.72 −1.05, p = 0.020
ASR error profile 89.83 89.68 −0.15, p = 0.390
Call phase 93.86 93.98 +0.12, p = 0.382
Timing extremes 95.01 95.04 +0.03, p = 0.472
Hard meanings 88.71 88.00 −0.70, p = 0.073
Third party 75.72 76.49 +0.76, p = 0.132
Mixed 85.07 84.88 −0.18, p = 0.374

Hard sub-types (pooled over 3 seeds)

sub-type Ettin-150m ModernBERT-base Δ
Interpreter relays the caller's answer (68) 53.9 58.8 +4.9
Gatekeeper says they will transfer (56) 66.1 72.6 +6.5
Recogniser clipped the first or last word (143) 84.8 86.5 +1.6
Recogniser dropped small words (131) 90.1 90.3 +0.3

The published weights (seed 13, one draw each)

model human labels (158) held-out A held-out B LLM-written hard stress acc / flip TTS → ASR OOD-10k leave vs. goodbye (41)
Ettin-150m (recommended) 98.10 95.28 93.37 90.67 92.94 / 2.70 91.57 89.76 36/41
ModernBERT-base (this model) 96.20 95.11 92.18 87.33 92.01 / 2.81 90.66 88.82 34/41

A single model is one draw: the comparison above uses 3-seed means and paired tests, never this table.

Real calls

The same 2,525 barge-in moments from 1,477 real outbound sales calls from a single campaign used on the Ettin-150m card, with the same references (decision at 0.5; 95% bootstrap intervals over calls; Δ is paired on the same resampled calls).

reference Ettin-150m ModernBERT-base Δ (ModernBERT-base − Ettin-150m)
Rule-unambiguous moments with words (278) 97.5 (95.2–99.3) 97.5 (95.2–99.3) –
Agreement with the LLM judge where it is confident, without the opener "hello" (945) 94.5 (93.0–95.9) 94.7 (93.2–96.0) +0.2 (−1.1…+1.5)
Agreement with the judge, without the opener "hello" (1,262) 87.8 (85.8–89.6) 88.3 (86.4–90.1) –
Agreement with the judge, all moments with words (1,425) 83.4 (81.5–85.4) 80.7 (78.7–82.8) −2.7 (−4.2…−1.1) (significant)
Stops the agent on an empty transcript (1,100; lower is better) 13.1 (11.1–15.2) 13.0 (11.0–15.0) −0.1 (−1.6…+1.4)
Stops the agent on a lone "hello" during the opener (163) 52.1 (44.0–59.8) 79.1 (72.7–85.4) +27.0 (+19.9…+34.6) (significant)

Speed

model GPU bf16, p50 (p95) CPU fp32, 4 threads, p50 (p95)
Ettin-150m 6.9 (7.8) ms 37.5 (43.8) ms
ModernBERT-base 6.9 (7.8) ms 37.1 (44.3) ms

Batch 1, end to end, one run with exclusive use of the GPU (RTX 6000 Ada; Intel Xeon Gold 5412U CPU), all models side by side. Same architecture and size, same speed.

Decision rule

A baseline would have replaced Ettin-150m only if it passed every check:

check ModernBERT-base
Unseen business types: gain ≥ +0.15 with p < 0.05 ✗ (−0.05)
Human labels not worse (Δ ≥ −0.63, one case) ✗ (−0.633)
Hard-set mix not worse (held-out B and LLM-written cases, Δ ≥ −0.3) ✓ (−0.19)
Perturbation stress not worse ✓ (−0.16 / flip −0.06)
Distribution shift not worse (Δ ≥ −0.2) ✗ (−0.30)
No shift type worse by more than 1.0 ✗ (worst: caller population −1.05)
Leave vs. goodbye within 3 points ✗ (78.0% vs 83.7%)
Leave vs. goodbye ≥ 85% (CPU-model bar) ✗ (78.0%)
Hard meanings not worse by more than 1.0 ✓ (−0.70)
Latency within 1.2× ✓ (1.00× GPU, 0.99× CPU)

Verdict

At identical speed, Ettin transfers better: ModernBERT-base ties on unseen business types and is behind on shifted calls, leave vs. goodbye-and-back, human labels and real calls. On real calls it stops on 79% of pickup "hello"s, against 52% for Ettin-150m.

The backbones share an architecture, so the difference is the pre-training data. One reading, not tested here, is that Ettin's pre-training transfers slightly better to calls unlike the training data. Either way, Ettin-150m is the model to use.

Training

Base model answerdotai/ModernBERT-base (22 layers, hidden 768, 150M parameters)
Head sequence classification (mean pooling), 2 labels
Fine-tuning full fine-tune, AdamW (weight decay 0.01), lr 5e-5, linear schedule with 6% warm-up, batch 32, 5 epochs, last checkpoint
Regularisation dropout 0.02 on embeddings, attention, MLP and head, plus R-Drop with α = 1.0
Precision / length bf16 autocast over fp32 master weights; max length 192 tokens
RoPE ModernBERT's own values are kept: θ = 160,000 for global and 10,000 for local attention (Ettin uses 160,000 for both)
Data the Ettin-150m training set: 64,151 synthetic exchanges, plus 323 held back as a dev monitor
Compute one NVIDIA RTX 6000 Ada, 51 min (3,075 s), peak 8.0 GB

The data is synthetic: LLM-generated multi-business exchanges, labelled by an LLM judge whose rubric was calibrated against about 325 human labels. See the Ettin-150m card for the full pipeline.

Limitations

Everything listed on the Ettin-150m card applies here too, and:

  • Leave vs. goodbye-and-back is weaker (78.0% over 3 seeds against 83.7% for Ettin-150m).
  • A lone "hello" while the agent's opener plays makes it stop on 79% of real pickups.
  • It is a baseline. It was trained and published to measure the effect of the base checkpoint, not for deployment.

Citation

@misc{ozonetg2026bargein,
  title        = {Barge-in classifier for voice agents},
  author       = {ozonetg},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/ozonetg/bargein-classifier-ettin-150m}}
}

@misc{modernbert,
  title         = {Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference},
  author        = {Benjamin Warner and Antoine Chaffin and Benjamin Clavié and Orion Weller and Oskar Hallström and Said Taghadouini and Alexis Gallagher and Raja Biswas and Faisal Ladhak and Tom Aarsen and Nathan Cooper and Griffin Adams and Jeremy Howard and Iacopo Poli},
  year          = {2024},
  eprint        = {2412.13663},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2412.13663}
}

@inproceedings{liang2021rdrop,
  title     = {R-Drop: Regularized Dropout for Neural Networks},
  author    = {Xiaobo Liang and Lijun Wu and Juntao Li and Yue Wang and Qi Meng and Tao Qin and Wei Chen and Min Zhang and Tie-Yan Liu},
  booktitle = {Advances in Neural Information Processing Systems},
  volume    = {34},
  year      = {2021},
  url       = {https://arxiv.org/abs/2106.14448}
}

Related

Downloads last month
6
Safetensors
Model size
0.1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ozonetg/bargein-classifier-modernbert-base

Finetuned
(1537)
this model

Collection including ozonetg/bargein-classifier-modernbert-base

Papers for ozonetg/bargein-classifier-modernbert-base

Evaluation results

  • Accuracy (threshold 0.5) on Held-out business types A (593 exchanges, 8 unseen business types)
    self-reported
    95.110
  • Accuracy (threshold 0.5) on Held-out business types B (588 exchanges, 20 unseen business types)
    self-reported
    92.180
  • Accuracy (threshold 0.5) on Human-labelled exchanges (158, not shown to the labelling judge)
    self-reported
    96.200
  • Accuracy (threshold 0.5) on Perturbation stress test (6,258)
    self-reported
    92.010
  • Accuracy (threshold 0.5) on TTS -> phone channel -> ASR stress test (1,328)
    self-reported
    90.660
  • Accuracy (threshold 0.5) on Distribution-shift test OOD-10k (10,888)
    self-reported
    88.820
  • Accuracy (threshold 0.5) on Real calls, rule-unambiguous moments with words (278)
    self-reported
    97.500