intenter / README.md
Jinou6502's picture
card: correct stale benchmark info to v668 measurements (KLUE dev 85.7/84.6, KLUE train 12,500, internal gates, regressions, latency, production status)
387aac6 verified
|
Raw History Blame Contribute Delete
18 kB
---
license: other
license_name: elda-community-license-1.0
license_link: LICENSE
language:
- ko
- en
- ja
library_name: transformers
pipeline_tag: token-classification
datasets:
- klue/klue
- AmazonScience/massive
tags:
- korean
- japanese
- multilingual
- intent-detection
- speech-act
- named-entity-recognition
- span-extraction
- conversational-ai
- real-time
- encoder
- modernbert
- byte-level-bpe
- long-context
extra_gated_heading: Access to Intenter
extra_gated_description: >-
Intenter is released for research, evaluation, internal validation and
education. Tell us who you are and we will grant access.
extra_gated_prompt: >-
By requesting access you agree to the Elda Community License 1.0:
no commercial use, no redistribution of the weights or derivatives,
and attribution as "Built with Elda". Commercial licensing is available
on request.
extra_gated_fields:
Name: text
Email: text
Affiliation: text
Intended use: text
I agree to the Elda Community License: checkbox
---
# Elda-Intenter β€” real-time speech-act + entity extraction in a single pass
**Korean-first, multilingual (ko Β· ja Β· en).** 307M parameters Β· **8,192-token window** Β· single forward pass.
> ⚠ Latency is **not stated for this build.** The previous release measured `p50 27.8 ms` on a
> 12-layer / 512-token backbone; this one is 22 layers with a 8,192 window, so that number does not
> carry over and we have not re-measured it.
One forward pass returns four aligned channels: what the speaker is *doing* (speech act), what they
are *talking about* (entities), and the attribute/predicate structure binding them. Not a
general-purpose NER model β€” the perception layer of a live conversational system.
```
"κ°„νŒμ˜ ν’ˆκ²©μ΄λΌλŠ” μŒμ‹μ μ΄ μžˆλ‹€λ©°, κ·Έ μŒμ‹μ μ€ 어디에 μžˆλƒκ³ "
intent NARRATE Β· DESCRIBE Β· QUESTION (three spans, one pass)
entity κ°„νŒμ˜ ν’ˆκ²© β†’ ORGANIZATION
μŒμ‹μ  β†’ ORGANIZATION
signals is_pure_negation=0 Β· is_claim=0 Β· polarity=0 Β· depends_on_prev=1
```
---
## What is different here β€” **byte-level spans, so an unseen name is still a span**
This build runs on a **byte-level BPE** tokenizer (Gemma-2 vocabulary, 256k). That is a decoder-style
property in an encoder, and it is the point of the model, not a detail.
```
28 shop names that are not in any index, tokenized:
previous build (SentencePiece unigram) 20 / 28 become [UNK]
this build (byte-level BPE) 0 / 28 become [UNK]
```
A name that becomes `[UNK]` cannot be a span, and a missing span is worse than a wrong type: the
layer downstream has **no string to look up**. With byte fallback every surface is representable β€”
rare syllables, emoji, `(@_!)--/` punctuation runs β€” so the span survives even when the name is new.
```
"🍜 λ„ˆμšΈμ‚¬μΈ κ°„νŒ 재질 μ•Œλ €μ€˜" β†’ λ„ˆμšΈμ‚¬μΈ / ORGANIZATION Β· REQUEST
"ㅸㅿ뷁곡방 μ–΄λ””μ•Ό" β†’ ㅸㅿ뷁곡방 / ORGANIZATION Β· QUESTION
```
*Both are shop names that do not exist; the second is deliberately unpronounceable. Neither
produces an `[UNK]` token on this build.*
β˜… **Representable is not the same as emitted.** On a held-out probe of synthetic shop names built
from syllables the previous tokenizer could not encode, anchor reach went **0% β†’ 95%** in the
question frame β€” but the same names inside an emoji-heavy frame still abstained 70–75% of the time.
Tokenization removes the floor; the span still has to be taught. That gap is this model's open work.
---
## Results
All sets below are excluded from training by construction.
### MASSIVE (dev, re-annotated) β€” conversational speech, same five peers
Voice-assistant utterances, our actual traffic shape. Same axes, same particle rule; peers emit
KLUE's six types, so only the shared axes are comparable and `type` is omitted.
| | abstention ↓ | span recall | boundary | precision | anchor reach |
|---|---|---|---|---|---|
| **ko** ours *(469 sent.)* | **6.9** | **93.1** | **95.2** | 96.5 | **84.3** |
| ko β€” 5 peers | 33.7 – 67.9 | 32.1 – 66.3 | 69.9 – 86.4 | 89.4 – 94.7 | 6.6 – 72.7 |
| **ja** ours *(749 sent.)* | **6.9** | **93.1** | **97.3** | 95.5 | **83.4** |
| ja β€” 5 peers | 30.1 – 76.2 | 23.8 – 69.9 | 50.8 – 84.0 | 72.3 – 93.9 | 18.8 – 48.0 |
| **en** ours *(796 sent.)* | **8.5** | **91.5** | **96.6** | 84.5 | **89.9** |
| en β€” 5 peers | 38.6 – 69.4 | 30.6 – 61.4 | 12.5 – 63.3 | 90.7 – 94.8 | 28.5 – 64.0 |
We lead every axis on ko and ja, and all but `precision` on en. Entity F1 under our own scorer:
**ja 79.7 Β· ko 85.5 Β· en 74.5** (one trailing function word allowed).
*Not comparable to published MASSIVE scores: that benchmark is intent classification plus slot
filling with 55 slot types. We map 21 onto 11 of our entity types; peers are KLUE-NER models run
outside their training domain. Zero training overlap on dev is verified by the evaluator.*
### KLUE-NER (dev Β· 600 sentences / 1,687 gold), entity-level F1
| | entity F1 | boundary F1 |
|---|---|---|
| exact match | 84.6 | 86.9 |
| allowing one trailing Korean particle | **85.7** | **88.0** |
*Korean particles attach to the noun. The second row accepts one trailing particle when the start
offset matches. ⚠ On this build the two rows are only 1.1 pp apart, because the model now emits
bare nouns β€” see* ***Korean particles*** *under Limitations.*
Measured with our own scorer on the same inputs, a KLUE-NER specialist fine-tuned on the full KLUE
train set scores **83.4** (`soddokayo/klue-roberta-large-klue-ner`). This model sees **12,500 KLUE-NER
training sentences** (of 21,008) and is otherwise a general conversational model.
| build | KLUE F1 (one particle) | note |
|---|---|---|
| previous release (12-layer, 512) | 63.9 | SentencePiece unigram |
| this release (22-layer, 8,192) | **85.7** | byte-level BPE + KLUE-NER training sentences |
*The jump is two changes at once β€” backbone and training data β€” so it is not a clean ablation of
either.*
### KLUE-NER β€” the same five peers, on their own training domain
Every peer is a dedicated KLUE-NER model fine-tuned on all 21,008 KLUE-NER training sentences,
emitting exactly KLUE's six types. We see **12,500** of them, emit 17 types projected onto those six,
and run four channels plus six auxiliary heads in the same pass.
> Re-measured on this build (`v668`) with the same five peers and the same scorer.
| axis | **ours** | KF-DeBERTa | KLUE-BERT | RoBERTa-large | KoELECTRA | RoBERTa-base |
|---|---|---|---|---|---|---|
| type accuracy | **96.7** | **98.2** | 97.7 | 98.1 | 95.1 | 92.2 |
| abstention *(lower better)* | **5.8** | **2.0** | 2.4 | 13.6 | 18.3 | 41.1 |
| span recall | **94.2** | **98.0** | 97.6 | 86.4 | 81.7 | 58.9 |
| boundary | **96.3** | **97.3** | 96.2 | 96.2 | 91.9 | 62.9 |
| precision | **96.4** | 97.2 | 97.4 | **98.0** | 96.0 | 96.7 |
| anchor reach | **87.9** | **94.8** | 94.6 | 69.8 | 54.8 | 13.0 |
**We still lose every axis to the best peer**, but the margins are now small: 2nd of six on
`boundary`, 3rd on `abstention`, `span recall` and `anchor reach`, 4th on `type`, 5th on `precision`.
Note the peer spread β€” anchor reach runs 13.0 to 94.8, so a single-peer table could have been picked
to favour us.
*Axes.* **abstention** β€” nothing emitted at a gold position. **span recall** β€” something emitted
there. **boundary** β€” returned span contains the gold string. **precision** β€” emitted spans landing
on gold, over spans carrying a KLUE-mappable type (1,597/1,657); over *all* our spans it is 95.6,
since some carry types KLUE does not annotate (WORK, ARTIFACTS, TERM, EVENT, DURATION).
**anchor reach** β€” span returned under a type that opens a downstream lookup (`LOCATION` /
`ORGANIZATION`; 447 of 1,687 gold here).
**KLUE-NER news text is 19.9 % of our training mix** (12,500 of 62,971 sentences); the rest is wiki,
conversational Korean and synthetic templates. Whether these peers hold up on conversational Korean is untested
here.
*Reported by others, on their own runs β€” not our ruler, not verified by us:* KLUE-RoBERTa-large
90.8 Β· XLM-R-large 85.9 Β· KR-BERT-base 77.2 Β· mBERT-base 73.2 ([KLUE
paper](https://arxiv.org/abs/2105.09680)). Our re-scoring of a KLUE-NER RoBERTa-large on this split
gave 83.4, not 90.8.
### Frozen internal gates
Measured on this build (`v668`), **per surface, not summed**. Rows marked β€” were not re-run on this
build; we leave them blank rather than carry an older build's number forward.
| | |
|---|---|
| Role nouns β€” detection / typing (405 slots) | **404 / 397** |
| Referent expressions, held out (23 slots) | **22 / 23** Β· typed 21 |
| Safety floor β€” base model / serving form | **12 / 12 Β· 14 / 14** |
| Conversational-speech detection | β€” |
| Possessive structure β€” owner kept (40 surfaces) | **37 / 40** |
| Discourse signals β€” false positives on short utterances | β€” |
| Discourse signals β€” polarity grid | β€” |
| Formal-register interrogatives | **12 / 14** |
| Proper-name span probe, third-party frames (80 slots) | **75 / 80** |
| Name-span probe, third-party (698 slots) | β€” |
| Span-convention compliance (126 slots) | **126 / 126** |
*The proper-name probe uses frames supplied by a separate team and was measured here by us; the
698-slot name-span probe and span-convention compliance are run by that team on their own
infrastructure (the former was not re-run on this build).*
---
## Why an encoder
| | Elda-Intenter | typical small-LLM extraction |
|---|---|---|
| Parameters | **307M** | 500M – 4B |
| Latency (single, incl. heads) | *not measured on this build* | hundreds of ms |
| Latency (batch of 16) | **118.5 ms** | seconds |
| Serving precision | fp32 (2 workers, 6.1 GB GPU) | varies |
| Output | fixed contract, char offsets | free text to be parsed |
Measured on the production serving path (RTX 3060, fp32, heads attached). Span extraction is
classification over token pairs, and a bidirectional encoder sees the whole utterance at once.
---
## Output contract
Four **Global Pointer** channels over one shared **mmBERT-base** backbone (22 layers, 8,192-token
window, byte-level BPE), plus binary signal
heads and one span-attribute head. All channels are character spans over the original string.
| channel | what it carries |
|---|---|
| `intents` | speech act over the utterance or clause β€” 18 types incl. AGREE / DISAGREE / QUESTION / COMMAND / NARRATE / DESCRIBE / REQUEST |
| `entities` | 17 entity types Γ— subtypes, plus `about_speaker` per span |
| `attributes` | modifiers bound to an entity |
| `predicates` | what is asserted of it |
| `signals` | four binary discourse signals, plus two three-valued judgements, each with probabilities |
`about_speaker` is `true` / `false` / `null` (`null` = abstains). `signals` also carries `place_q`
("did the speaker point at one specific place or business?") and `needs_map` ("must something
outside the model be consulted to answer?"), each `"yes"` / `"no"` / `null` with a probability.
**No threshold is applied to any of them** β€” label and probability are both emitted and the cut is
the consumer's. See `ABOUT_SPEAKER_OUTPUT_CONTRACT.md` and `SIGNALS_JUDGE_OUTPUT_CONTRACT.md`.
Spans may nest (γ€Œμ œ μΉœκ΅¬γ€ and γ€ŒμΉœκ΅¬γ€). Key by span offsets, not by surface.
## Usage
No `from_pretrained` one-liner β€” custom four-channel span architecture.
```python
import json, sys
from pathlib import Path
d = Path("path/to/this/repo")
sys.path.insert(0, str(d))
import modeling_btrack_4ch_w4 as M
m = M.load(str(d), "cuda") # base + 5 gap-fill heads + about_speaker + W4 + 2 judge heads
print(json.dumps(M.infer(m, "판ꡐ 카카였 본사 μ–΄λ””μ•Ό"), ensure_ascii=False, indent=1))
# model.safetensors carries the same weights in the same dtype, if you prefer it:
# from safetensors.torch import load_file
# m.load_state_dict(load_file(d / "model.safetensors"))
```
`backbone_config.json` is included β€” the repository is self-contained. `md5sum -c MD5SUMS`
verifies a download in full; `model.pt` is bit-identical to the file answering live traffic.
Serve in **fp32** (weights stored bf16). **INT8 collapses this model** β€” measured.
---
## Limitations
* **Korean first.** ja and en are supported and measured; Korean gets the curated material.
* **Not a general NER model.** Type inventory and span scope come from a conversational product.
* **Korean particles β€” this build trims them.** Earlier releases kept the particle inside the span
(`μ–΄λΆ€κ°€`); this one emits the bare noun (`μ–΄λΆ€`). We measured the shift rather than intending it:
the gap between exact match and *one-particle-allowed* scoring collapsed from **8.6 pp** to
**1.1 pp** on this build, and on a held-out probe 79.5 % of the spans that changed were particle-bearing spans
replaced by their bare form (97 matched pairs, **zero** cases of a noun being cut short). The
training mix now includes KLUE-NER, whose gold excludes particles, and the model followed that
convention. Whether to pull it back is an open decision for us β€” reverting costs KLUE score.
* **News text is no longer the weak spot** β€” 12,500 KLUE-NER news sentences are in the mix, and KLUE F1
went 63.9 β†’ 85.7. That gain came with the losses listed under *Known regressions* below.
* **`is_pure_negation`** has false positives no threshold recovers (8 fp / 8 fn on 3,193 rows at
Ο„=0.7). A soft signal, not a filter.
* **Boundaries, not senses.** Entity linking is a separate layer, not in this repository.
### Known regressions (vs. the previous production model)
This build changed the backbone, so some axes moved. We publish the numbers rather than the summary.
| gate | previous production | previous release | **this build** |
|---|---|---|---|
| risk-utterance floor, serving form *(14 lines)* | **14 / 14** | 14 / 14 | **14 / 14** |
| risk-utterance floor, base model *(12 lines)* | **12 / 12** | 12 / 12 | **12 / 12** |
| referential expressions, held out *(23 slots)* | **23 / 23** | 23 / 23 | 22 / 23 |
| connective-clause abstention *(50 cells, lower is better)* | **0 / 50** | 0 / 50 | 2 / 50 |
| role nouns, detected *(405 slots)* | 401 | **404** | **404** |
| role noun γ€Œκ²½λΉ„λ³‘γ€ typed as person *(46 slots)* | **46** | 45 | 39 |
| map-intent judge heads, consumer sheet *(place_q Β· needs_map, of 115)* | **114 Β· 114** | 112 Β· 113 | 110 Β· 112 |
| entity-set drift probe, net *(boundary moves excluded)* | baseline | βˆ’128 | **βˆ’65** |
The safety floor holds at production parity on both forms. What we gave up: one held-out referential
slot, two connective-clause abstentions, and the role noun γ€Œκ²½λΉ„λ³‘γ€, which loses person typing in
several of our probes at once β€” that is our open work. The drift probe roughly halved its loss
against the previous release.
**Latency** β€” measured by the serving team on their GPU host: single request p50 **28.7 β†’ 30.5 ms**,
16-request batch **126.5 β†’ 114.3 ms** against the previous production model. Not re-measured on
other hardware.
---
## This repository
Build `45b7cc4c`. Refreshed roughly monthly.
This build was promoted to production on 2026-10-06. Production has since moved to the **same
weights** with two updated auxiliary heads (re-issued map-intent judge heads and a type-cover head,
2026-10-07); those heads are not in this repository. Two output contracts ship alongside the weights
(`ABOUT_SPEAKER_OUTPUT_CONTRACT.md`, `SIGNALS_JUDGE_OUTPUT_CONTRACT.md`).
## Attribution β€” third-party training data
Parts of the training corpus are adapted from publicly licensed datasets. Their licenses apply to
those parts and are reproduced here.
**KLUE** β€” *KLUE: Korean Language Understanding Evaluation*, Park et al., 2021.
Licensed under [CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/).
Source: <https://github.com/KLUE-benchmark/KLUE> Β· paper: [arXiv:2105.09680](https://arxiv.org/abs/2105.09680).
> **Changes we made (this is an adaptation, not a copy).** 12,500 sentences from the KLUE-NER
> *training* split were re-annotated under our own span convention and type inventory: KLUE's six
> entity types were mapped onto our seventeen, sentences overlapping our held-out probes were
> removed, and in 192 sentences common-noun person words that KLUE leaves unannotated (e.g. μ‚¬λžŒ) were
> labeled as persons under our convention. Intent, attribute and predicate channels carry no KLUE labels and are masked out during
> training. The **dev** split is used for evaluation only and appears nowhere in training β€” the
> evaluator verifies this and refuses to run otherwise.
**MASSIVE** β€” *MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset*,
FitzGerald et al., Amazon Science. Licensed under
[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/).
Source: <https://github.com/alexa/massive>.
> **Changes we made.** Korean and Japanese rows were re-annotated under our span convention; 21 of
> MASSIVE's 55 slot types were mapped onto 11 of our entity types. MASSIVE's intent labels are
> **not used** β€” that channel is masked out per source. Dev rows used for evaluation are excluded
> from training by construction.
Everything else in the corpus is our own work. Training data and evaluation sets are not
distributed with this repository.
## License & access
Released under the **Elda Community License 1.0** (see `LICENSE`).
* βœ… Research, evaluation, internal validation, education β€” free of charge
* ✳ Attribution: **"Built with Elda"**
* β›” Commercial use and redistribution require a separate agreement
Access is gated: tell us who you are and access is granted automatically. Training data and
evaluation sets are not distributed.
## Citation
```bibtex
@software{intenter2026,
title = {Elda-Intenter: real-time multilingual speech-act and entity extraction},
author = {Elda AI},
year = {2026},
url = {https://huggingface.co/Elda-AI/intenter}
}
```