Instructions to use Elda-AI/intenter with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Elda-AI/intenter with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="Elda-AI/intenter")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Elda-AI/intenter", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download README.md from Elda-AI/intenter: direct link, hf CLI and curl.
- Browser
- Download file 18 kB
-
https://huggingface.co/Elda-AI/intenter/resolve/main/README.md
- Command line
-
hf download hf://Elda-AI/intenter/README.md
-
curl -L -H "Authorization: Bearer $HF_TOKEN" -o README.md https://huggingface.co/Elda-AI/intenter/resolve/main/README.md
license: other
license_name: elda-community-license-1.0
license_link: LICENSE
language:
- ko
- en
- ja
library_name: transformers
pipeline_tag: token-classification
datasets:
- klue/klue
- AmazonScience/massive
tags:
- korean
- japanese
- multilingual
- intent-detection
- speech-act
- named-entity-recognition
- span-extraction
- conversational-ai
- real-time
- encoder
- modernbert
- byte-level-bpe
- long-context
extra_gated_heading: Access to Intenter
extra_gated_description: >-
Intenter is released for research, evaluation, internal validation and
education. Tell us who you are and we will grant access.
extra_gated_prompt: >-
By requesting access you agree to the Elda Community License 1.0: no
commercial use, no redistribution of the weights or derivatives, and
attribution as "Built with Elda". Commercial licensing is available on
request.
extra_gated_fields:
Name: text
Email: text
Affiliation: text
Intended use: text
I agree to the Elda Community License: checkbox
Elda-Intenter β real-time speech-act + entity extraction in a single pass
Korean-first, multilingual (ko Β· ja Β· en). 307M parameters Β· 8,192-token window Β· single forward pass.
β Latency is not stated for this build. The previous release measured
p50 27.8 mson a 12-layer / 512-token backbone; this one is 22 layers with a 8,192 window, so that number does not carry over and we have not re-measured it.
One forward pass returns four aligned channels: what the speaker is doing (speech act), what they are talking about (entities), and the attribute/predicate structure binding them. Not a general-purpose NER model β the perception layer of a live conversational system.
"κ°νμ ν격μ΄λΌλ μμμ μ΄ μλ€λ©°, κ·Έ μμμ μ μ΄λμ μλκ³ "
intent NARRATE Β· DESCRIBE Β· QUESTION (three spans, one pass)
entity κ°νμ ν격 β ORGANIZATION
μμμ β ORGANIZATION
signals is_pure_negation=0 Β· is_claim=0 Β· polarity=0 Β· depends_on_prev=1
What is different here β byte-level spans, so an unseen name is still a span
This build runs on a byte-level BPE tokenizer (Gemma-2 vocabulary, 256k). That is a decoder-style property in an encoder, and it is the point of the model, not a detail.
28 shop names that are not in any index, tokenized:
previous build (SentencePiece unigram) 20 / 28 become [UNK]
this build (byte-level BPE) 0 / 28 become [UNK]
A name that becomes [UNK] cannot be a span, and a missing span is worse than a wrong type: the
layer downstream has no string to look up. With byte fallback every surface is representable β
rare syllables, emoji, (@_!)--/ punctuation runs β so the span survives even when the name is new.
"π λμΈμ¬μΈ κ°ν μ¬μ§ μλ €μ€" β λμΈμ¬μΈ / ORGANIZATION Β· REQUEST
"γ
Έγ
Ώλ·κ³΅λ°© μ΄λμΌ" β γ
Έγ
Ώλ·κ³΅λ°© / ORGANIZATION Β· QUESTION
Both are shop names that do not exist; the second is deliberately unpronounceable. Neither
produces an [UNK] token on this build.
β Representable is not the same as emitted. On a held-out probe of synthetic shop names built from syllables the previous tokenizer could not encode, anchor reach went 0% β 95% in the question frame β but the same names inside an emoji-heavy frame still abstained 70β75% of the time. Tokenization removes the floor; the span still has to be taught. That gap is this model's open work.
Results
All sets below are excluded from training by construction.
MASSIVE (dev, re-annotated) β conversational speech, same five peers
Voice-assistant utterances, our actual traffic shape. Same axes, same particle rule; peers emit
KLUE's six types, so only the shared axes are comparable and type is omitted.
| abstention β | span recall | boundary | precision | anchor reach | |
|---|---|---|---|---|---|
| ko ours (469 sent.) | 6.9 | 93.1 | 95.2 | 96.5 | 84.3 |
| ko β 5 peers | 33.7 β 67.9 | 32.1 β 66.3 | 69.9 β 86.4 | 89.4 β 94.7 | 6.6 β 72.7 |
| ja ours (749 sent.) | 6.9 | 93.1 | 97.3 | 95.5 | 83.4 |
| ja β 5 peers | 30.1 β 76.2 | 23.8 β 69.9 | 50.8 β 84.0 | 72.3 β 93.9 | 18.8 β 48.0 |
| en ours (796 sent.) | 8.5 | 91.5 | 96.6 | 84.5 | 89.9 |
| en β 5 peers | 38.6 β 69.4 | 30.6 β 61.4 | 12.5 β 63.3 | 90.7 β 94.8 | 28.5 β 64.0 |
We lead every axis on ko and ja, and all but precision on en. Entity F1 under our own scorer:
ja 79.7 Β· ko 85.5 Β· en 74.5 (one trailing function word allowed).
Not comparable to published MASSIVE scores: that benchmark is intent classification plus slot filling with 55 slot types. We map 21 onto 11 of our entity types; peers are KLUE-NER models run outside their training domain. Zero training overlap on dev is verified by the evaluator.
KLUE-NER (dev Β· 600 sentences / 1,687 gold), entity-level F1
| entity F1 | boundary F1 | |
|---|---|---|
| exact match | 84.6 | 86.9 |
| allowing one trailing Korean particle | 85.7 | 88.0 |
Korean particles attach to the noun. The second row accepts one trailing particle when the start offset matches. β On this build the two rows are only 1.1 pp apart, because the model now emits bare nouns β see Korean particles under Limitations.
Measured with our own scorer on the same inputs, a KLUE-NER specialist fine-tuned on the full KLUE
train set scores 83.4 (soddokayo/klue-roberta-large-klue-ner). This model sees 12,500 KLUE-NER
training sentences (of 21,008) and is otherwise a general conversational model.
| build | KLUE F1 (one particle) | note |
|---|---|---|
| previous release (12-layer, 512) | 63.9 | SentencePiece unigram |
| this release (22-layer, 8,192) | 85.7 | byte-level BPE + KLUE-NER training sentences |
The jump is two changes at once β backbone and training data β so it is not a clean ablation of either.
KLUE-NER β the same five peers, on their own training domain
Every peer is a dedicated KLUE-NER model fine-tuned on all 21,008 KLUE-NER training sentences, emitting exactly KLUE's six types. We see 12,500 of them, emit 17 types projected onto those six, and run four channels plus six auxiliary heads in the same pass.
Re-measured on this build (
v668) with the same five peers and the same scorer.
| axis | ours | KF-DeBERTa | KLUE-BERT | RoBERTa-large | KoELECTRA | RoBERTa-base |
|---|---|---|---|---|---|---|
| type accuracy | 96.7 | 98.2 | 97.7 | 98.1 | 95.1 | 92.2 |
| abstention (lower better) | 5.8 | 2.0 | 2.4 | 13.6 | 18.3 | 41.1 |
| span recall | 94.2 | 98.0 | 97.6 | 86.4 | 81.7 | 58.9 |
| boundary | 96.3 | 97.3 | 96.2 | 96.2 | 91.9 | 62.9 |
| precision | 96.4 | 97.2 | 97.4 | 98.0 | 96.0 | 96.7 |
| anchor reach | 87.9 | 94.8 | 94.6 | 69.8 | 54.8 | 13.0 |
We still lose every axis to the best peer, but the margins are now small: 2nd of six on
boundary, 3rd on abstention, span recall and anchor reach, 4th on type, 5th on precision.
Note the peer spread β anchor reach runs 13.0 to 94.8, so a single-peer table could have been picked
to favour us.
Axes. abstention β nothing emitted at a gold position. span recall β something emitted
there. boundary β returned span contains the gold string. precision β emitted spans landing
on gold, over spans carrying a KLUE-mappable type (1,597/1,657); over all our spans it is 95.6,
since some carry types KLUE does not annotate (WORK, ARTIFACTS, TERM, EVENT, DURATION).
anchor reach β span returned under a type that opens a downstream lookup (LOCATION /
ORGANIZATION; 447 of 1,687 gold here).
KLUE-NER news text is 19.9 % of our training mix (12,500 of 62,971 sentences); the rest is wiki, conversational Korean and synthetic templates. Whether these peers hold up on conversational Korean is untested here.
Reported by others, on their own runs β not our ruler, not verified by us: KLUE-RoBERTa-large 90.8 Β· XLM-R-large 85.9 Β· KR-BERT-base 77.2 Β· mBERT-base 73.2 (KLUE paper). Our re-scoring of a KLUE-NER RoBERTa-large on this split gave 83.4, not 90.8.
Frozen internal gates
Measured on this build (v668), per surface, not summed. Rows marked β were not re-run on this
build; we leave them blank rather than carry an older build's number forward.
| Role nouns β detection / typing (405 slots) | 404 / 397 |
| Referent expressions, held out (23 slots) | 22 / 23 Β· typed 21 |
| Safety floor β base model / serving form | 12 / 12 Β· 14 / 14 |
| Conversational-speech detection | β |
| Possessive structure β owner kept (40 surfaces) | 37 / 40 |
| Discourse signals β false positives on short utterances | β |
| Discourse signals β polarity grid | β |
| Formal-register interrogatives | 12 / 14 |
| Proper-name span probe, third-party frames (80 slots) | 75 / 80 |
| Name-span probe, third-party (698 slots) | β |
| Span-convention compliance (126 slots) | 126 / 126 |
The proper-name probe uses frames supplied by a separate team and was measured here by us; the 698-slot name-span probe and span-convention compliance are run by that team on their own infrastructure (the former was not re-run on this build).
Why an encoder
| Elda-Intenter | typical small-LLM extraction | |
|---|---|---|
| Parameters | 307M | 500M β 4B |
| Latency (single, incl. heads) | not measured on this build | hundreds of ms |
| Latency (batch of 16) | 118.5 ms | seconds |
| Serving precision | fp32 (2 workers, 6.1 GB GPU) | varies |
| Output | fixed contract, char offsets | free text to be parsed |
Measured on the production serving path (RTX 3060, fp32, heads attached). Span extraction is classification over token pairs, and a bidirectional encoder sees the whole utterance at once.
Output contract
Four Global Pointer channels over one shared mmBERT-base backbone (22 layers, 8,192-token window, byte-level BPE), plus binary signal heads and one span-attribute head. All channels are character spans over the original string.
| channel | what it carries |
|---|---|
intents |
speech act over the utterance or clause β 18 types incl. AGREE / DISAGREE / QUESTION / COMMAND / NARRATE / DESCRIBE / REQUEST |
entities |
17 entity types Γ subtypes, plus about_speaker per span |
attributes |
modifiers bound to an entity |
predicates |
what is asserted of it |
signals |
four binary discourse signals, plus two three-valued judgements, each with probabilities |
about_speaker is true / false / null (null = abstains). signals also carries place_q
("did the speaker point at one specific place or business?") and needs_map ("must something
outside the model be consulted to answer?"), each "yes" / "no" / null with a probability.
No threshold is applied to any of them β label and probability are both emitted and the cut is
the consumer's. See ABOUT_SPEAKER_OUTPUT_CONTRACT.md and SIGNALS_JUDGE_OUTPUT_CONTRACT.md.
Spans may nest (γμ μΉκ΅¬γ and γμΉκ΅¬γ). Key by span offsets, not by surface.
Usage
No from_pretrained one-liner β custom four-channel span architecture.
import json, sys
from pathlib import Path
d = Path("path/to/this/repo")
sys.path.insert(0, str(d))
import modeling_btrack_4ch_w4 as M
m = M.load(str(d), "cuda") # base + 5 gap-fill heads + about_speaker + W4 + 2 judge heads
print(json.dumps(M.infer(m, "νκ΅ μΉ΄μΉ΄μ€ λ³Έμ¬ μ΄λμΌ"), ensure_ascii=False, indent=1))
# model.safetensors carries the same weights in the same dtype, if you prefer it:
# from safetensors.torch import load_file
# m.load_state_dict(load_file(d / "model.safetensors"))
backbone_config.json is included β the repository is self-contained. md5sum -c MD5SUMS
verifies a download in full; model.pt is bit-identical to the file answering live traffic.
Serve in fp32 (weights stored bf16). INT8 collapses this model β measured.
Limitations
- Korean first. ja and en are supported and measured; Korean gets the curated material.
- Not a general NER model. Type inventory and span scope come from a conversational product.
- Korean particles β this build trims them. Earlier releases kept the particle inside the span
(
μ΄λΆκ°); this one emits the bare noun (μ΄λΆ). We measured the shift rather than intending it: the gap between exact match and one-particle-allowed scoring collapsed from 8.6 pp to 1.1 pp on this build, and on a held-out probe 79.5 % of the spans that changed were particle-bearing spans replaced by their bare form (97 matched pairs, zero cases of a noun being cut short). The training mix now includes KLUE-NER, whose gold excludes particles, and the model followed that convention. Whether to pull it back is an open decision for us β reverting costs KLUE score. - News text is no longer the weak spot β 12,500 KLUE-NER news sentences are in the mix, and KLUE F1 went 63.9 β 85.7. That gain came with the losses listed under Known regressions below.
is_pure_negationhas false positives no threshold recovers (8 fp / 8 fn on 3,193 rows at Ο=0.7). A soft signal, not a filter.- Boundaries, not senses. Entity linking is a separate layer, not in this repository.
Known regressions (vs. the previous production model)
This build changed the backbone, so some axes moved. We publish the numbers rather than the summary.
| gate | previous production | previous release | this build |
|---|---|---|---|
| risk-utterance floor, serving form (14 lines) | 14 / 14 | 14 / 14 | 14 / 14 |
| risk-utterance floor, base model (12 lines) | 12 / 12 | 12 / 12 | 12 / 12 |
| referential expressions, held out (23 slots) | 23 / 23 | 23 / 23 | 22 / 23 |
| connective-clause abstention (50 cells, lower is better) | 0 / 50 | 0 / 50 | 2 / 50 |
| role nouns, detected (405 slots) | 401 | 404 | 404 |
| role noun γκ²½λΉλ³γ typed as person (46 slots) | 46 | 45 | 39 |
| map-intent judge heads, consumer sheet (place_q Β· needs_map, of 115) | 114 Β· 114 | 112 Β· 113 | 110 Β· 112 |
| entity-set drift probe, net (boundary moves excluded) | baseline | β128 | β65 |
The safety floor holds at production parity on both forms. What we gave up: one held-out referential slot, two connective-clause abstentions, and the role noun γκ²½λΉλ³γ, which loses person typing in several of our probes at once β that is our open work. The drift probe roughly halved its loss against the previous release.
Latency β measured by the serving team on their GPU host: single request p50 28.7 β 30.5 ms, 16-request batch 126.5 β 114.3 ms against the previous production model. Not re-measured on other hardware.
This repository
Build 45b7cc4c. Refreshed roughly monthly.
This build was promoted to production on 2026-10-06. Production has since moved to the same
weights with two updated auxiliary heads (re-issued map-intent judge heads and a type-cover head,
2026-10-07); those heads are not in this repository. Two output contracts ship alongside the weights
(ABOUT_SPEAKER_OUTPUT_CONTRACT.md, SIGNALS_JUDGE_OUTPUT_CONTRACT.md).
Attribution β third-party training data
Parts of the training corpus are adapted from publicly licensed datasets. Their licenses apply to those parts and are reproduced here.
KLUE β KLUE: Korean Language Understanding Evaluation, Park et al., 2021. Licensed under CC BY-SA 4.0. Source: https://github.com/KLUE-benchmark/KLUE Β· paper: arXiv:2105.09680.
Changes we made (this is an adaptation, not a copy). 12,500 sentences from the KLUE-NER training split were re-annotated under our own span convention and type inventory: KLUE's six entity types were mapped onto our seventeen, sentences overlapping our held-out probes were removed, and in 192 sentences common-noun person words that KLUE leaves unannotated (e.g. μ¬λ) were labeled as persons under our convention. Intent, attribute and predicate channels carry no KLUE labels and are masked out during training. The dev split is used for evaluation only and appears nowhere in training β the evaluator verifies this and refuses to run otherwise.
MASSIVE β MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset, FitzGerald et al., Amazon Science. Licensed under CC BY 4.0. Source: https://github.com/alexa/massive.
Changes we made. Korean and Japanese rows were re-annotated under our span convention; 21 of MASSIVE's 55 slot types were mapped onto 11 of our entity types. MASSIVE's intent labels are not used β that channel is masked out per source. Dev rows used for evaluation are excluded from training by construction.
Everything else in the corpus is our own work. Training data and evaluation sets are not distributed with this repository.
License & access
Released under the Elda Community License 1.0 (see LICENSE).
- β Research, evaluation, internal validation, education β free of charge
- β³ Attribution: "Built with Elda"
- β Commercial use and redistribution require a separate agreement
Access is gated: tell us who you are and access is granted automatically. Training data and evaluation sets are not distributed.
Citation
@software{intenter2026,
title = {Elda-Intenter: real-time multilingual speech-act and entity extraction},
author = {Elda AI},
year = {2026},
url = {https://huggingface.co/Elda-AI/intenter}
}