topic-scope-v3
The model behind the topic_scope detector in flowx-border. Given a message and a taxonomy a policy supplies, it answers which node the message is about, or "none of these", with a calibrated probability.
The same design as flowxai/topic-scope-v2, on XLM-RoBERTa large instead of base: same corpus, same head, same calibration, and a training config that differs only in the learning rate and batch layout the larger encoder needs. Both are compared below on the same rows, with the bi-encoder both replace, flowxai/topic-scope.
How it runs
XLM-RoBERTa large, fine-tuned, plus a two-layer decision head trained from scratch. The late encoding: the message, the question and each node are encoded as separate short sequences, and meet only in the head. Node states are therefore computed once per taxonomy and cached; a scan encodes the message and runs the head.
| file | what | bytes | sha256 |
|---|---|---|---|
onnx/model.int8.onnx |
encoder, embedding Gather in int8, outputs last_hidden_state | 1,466,403,115 | 483552daf200d566... |
onnx/head.onnx |
decision head, fp32, takes hidden states and returns one logit per option | 105,030,983 | 69a14a213934d71e... |
decision_config.json |
question text, the none option, token budgets, none offset, temperatures | 1,020 | beaeab650cc02998... |
tokenizer.json |
XLM-R tokenizer | 17,098,086 | acbd420e2269cdc1... |
config.json |
encoder config | 743 | d44ba2614aaeeb01... |
Token budgets: message 128, question 48, each node 64. A node is rendered as {key}: {text}, and the none option is always offered as none: none of these topics; the message is about something else. Its logit gets -0.5 before calibration. Validation fitted -0.25; -0.5 is shipped (changed 2026-10-01): topic-scope-v2 ships -0.5, and on 400 hand-written probes this model scores 0.752 at -0.5 against 0.728 at -0.25, while held-out accuracy moves by under 0.3 points (training repository, reports/topic_scope_v3_none_offset.json). Temperatures are as fitted. Probabilities are a softmax at a temperature fitted per option-count bucket on validation rows, listed in decision_config.json.
Options are packed without padding for the head: it has no position embedding, so this is the function it was trained as, over fewer positions.
Evaluation
Synthetic, and held out by deployment. The corpus is generated: 299 taxonomies and 135,747 messages in the 26 languages, written by claude-sonnet-5 (most messages) and gpt-5.6-luna, and every validation and test row checked by a judge model, a different model from the writer for 97% of messages. Splits are by taxonomy. unseen_type taxonomies come from deployment types no training taxonomy came from (airline, legal services, telecom); unseen_schema taxonomies are new taxonomies of trained types. Neither is a measurement on real traffic, and the numbers below should be read as how the model transfers to taxonomies it has not seen, not as a production error rate.
Accuracy is top-1 over the offered nodes plus none, after the none offset. The comparison is the model this one replaces, on the same rows, with its none bar (0.875) chosen on validation.
| test | this model | topic-scope-v2 | flowxai/topic-scope | rows | ECE |
|---|---|---|---|---|---|
| unseen_type | 0.859 | 0.804 | 0.479 | 11,497 | 0.0118 |
| unseen_schema | 0.887 | 0.808 | 0.471 | 11,071 | 0.0141 |
| v1 topic_scope test, recast | 0.973 | 442 | 0.0145 |
Two training seeds. The published weights are seed 42. A second run on the same config and data, seed 1337, was evaluated on the same rows at the same none offset and not exported. The spread between them is the honest error bar on the table above.
| test | published seed | seed 1337 |
|---|---|---|
| topic_scope_v2 unseen_type | 0.859 | 0.869 |
| topic_scope_v2 unseen_schema | 0.887 | 0.887 |
| policy_questions unseen_type | 0.807 | 0.823 |
| policy_questions unseen_schema | 0.817 | 0.802 |
| policy minimal pairs, unseen_type | 0.632 | 0.668 |
Hand-written probes
Every number above comes from synthetic rows written by the same models that wrote the training set. As a check from outside that style, 400 probes were written by hand: short, plain questions in 10 languages (en, ro, de, fr, es, it, pl, nl, pt, tr) against small taxonomies of 3 allowed and 3 disallowed nodes (banking, twice with two description styles, insurance and telecom), run through the library at threshold 0.5. in passes when nothing fires and the answer is not none; dis when the expected disallowed node fires; out, about nothing in the taxonomy, when nothing fires. The probes ship with the library as a test fixture.
| model | in | dis | out | all |
|---|---|---|---|---|
| topic-scope-v2 | 127/200 | 76/120 | 80/80 | 0.708 |
| topic-scope-v3 at -0.25 | 124/200 | 87/120 | 80/80 | 0.728 |
| topic-scope-v3 at -0.5 | 130/200 | 91/120 | 80/80 | 0.752 |
Paired exact McNemar against topic-scope-v2: p=0.0474 over all 400 (46 probes only this model gets right, 28 only v2), p=0.0107 on the disallowed ones. The shipped offset was chosen after these probes were run, so they are not an independent test of that choice; the held-out synthetic rows, which moved by under 0.3 points, are.
Per language
Every language is reported, the weak ones included. Maltese is the weakest, as it is for every model on this base: Maltese is not in XLM-RoBERTa's pretraining set, and fine-tuning data cannot give the encoder a representation it never had. Irish is second.
| language | unseen_type | topic-scope-v2 | rows | unseen_schema | topic-scope-v2 | rows |
|---|---|---|---|---|---|---|
| az | 0.816 | 0.753 | 413 | 0.902 | 0.816 | 705 |
| bg | 0.825 | 0.804 | 445 | 0.899 | 0.812 | 485 |
| cs | 0.862 | 0.801 | 412 | 0.856 | 0.808 | 500 |
| da | 0.844 | 0.833 | 551 | 0.880 | 0.794 | 209 |
| de | 0.824 | 0.771 | 546 | 0.889 | 0.839 | 217 |
| el | 0.857 | 0.814 | 462 | 0.911 | 0.789 | 123 |
| en | 0.828 | 0.792 | 471 | 0.842 | 0.742 | 120 |
| es | 0.899 | 0.851 | 496 | 0.846 | 0.802 | 343 |
| et | 0.851 | 0.814 | 489 | 0.857 | 0.796 | 343 |
| fi | 0.885 | 0.851 | 495 | 0.885 | 0.798 | 573 |
| fr | 0.888 | 0.854 | 493 | 0.906 | 0.825 | 576 |
| ga | 0.808 | 0.710 | 427 | 0.848 | 0.676 | 732 |
| hr | 0.893 | 0.842 | 430 | 0.897 | 0.834 | 735 |
| hu | 0.859 | 0.830 | 383 | 0.895 | 0.843 | 477 |
| it | 0.895 | 0.847 | 372 | 0.905 | 0.839 | 465 |
| lt | 0.861 | 0.816 | 403 | 0.899 | 0.822 | 365 |
| lv | 0.895 | 0.805 | 389 | 0.887 | 0.797 | 364 |
| mt | 0.750 | 0.624 | 412 | 0.769 | 0.556 | 117 |
| nl | 0.862 | 0.818 | 406 | 0.839 | 0.750 | 112 |
| pl | 0.876 | 0.790 | 434 | 0.887 | 0.793 | 300 |
| pt | 0.895 | 0.851 | 449 | 0.874 | 0.824 | 301 |
| ro | 0.861 | 0.797 | 454 | 0.903 | 0.838 | 401 |
| sk | 0.896 | 0.793 | 425 | 0.929 | 0.837 | 406 |
| sl | 0.882 | 0.814 | 414 | 0.904 | 0.848 | 709 |
| sv | 0.888 | 0.823 | 419 | 0.896 | 0.840 | 686 |
| tr | 0.840 | 0.784 | 407 | 0.891 | 0.826 | 707 |
Per register
What kind of message it was. _hole rows had their correct node withheld from the options, so the right answer is none. negated_near_miss is a message that matches a node's words but falls under what the node excludes, and it is the weakest register on unseen types.
| register | unseen_type | topic-scope-v2 | rows | unseen_schema | topic-scope-v2 | rows |
|---|---|---|---|---|---|---|
| in_domain_uncovered | 0.883 | 0.865 | 814 | 0.912 | 0.872 | 693 |
| in_node | 0.869 | 0.806 | 4661 | 0.905 | 0.819 | 4485 |
| in_node_hole | 0.835 | 0.801 | 1521 | 0.848 | 0.766 | 1531 |
| negated_near_miss | 0.728 | 0.667 | 672 | 0.844 | 0.771 | 741 |
| negated_near_miss_hole | 0.759 | 0.782 | 87 | 0.905 | 0.778 | 63 |
| out_of_taxonomy | 0.980 | 0.979 | 817 | 0.997 | 0.995 | 645 |
| sibling_near_miss | 0.864 | 0.778 | 2212 | 0.886 | 0.795 | 2132 |
| sibling_near_miss_hole | 0.801 | 0.738 | 713 | 0.789 | 0.694 | 781 |
Export verification
300 test rows, the torch fp32 model on its training layout against the ONNX pipeline (onnxruntime CPU, one thread, int8 encoder, packed head): 0 changed their answer, and the largest logit difference was 0.13505. The head took 30.84 ms at p50 and 69.85 ms at p95 on those rows, one thread, on an Apple M-series laptop; the library's budget test measures the whole detector.
Latency
The whole topic_scope detector through the library, Apple M5 Max, one thread, onnxruntime CPU provider, measured 2026-09-30: tests/test_budgets.py p95 over 20 runs at REFERENCE_INPUT (396 characters), after one warm run; nodes are real descriptions from tests/fixtures/topic_scope/typed_26_languages.json, one disallowed and the rest allowed; node states cached per taxonomy as in a deployment. The shipped base model was measured beside it in the same session. Budget 300 ms.
| nodes | this model, p95 | topic-scope-v2, p95 |
|---|---|---|
| 20 | 113.7 ms | 57.4 ms |
| 30 | 145.7 ms | 82.7 ms |
| 40 | 180.8 ms | 102.3 ms |
Base at 40 nodes reads 102.3 here against 104.1 recorded in the library's MEASURED_MS, so the machine was quiet.
Limits
- Trained on taxonomies of 4 to 37 nodes. Head cost grows with the square of the total node text, so the library shortlists larger taxonomies to the nearest nodes by the model's own similarity term first, and records that it did.
- A node description past 64 tokens is cut. Write the part that decides first.
- The head was trained on four question families (topic_scope v1 and v2, injection and policy questions). The library asks it the topic question only. Its policy-question accuracy is not good enough to ship: 0.63 of minimal pairs on unseen types get both halves right, against a shipping bar of 0.80.
- Short plain in-scope questions get "none of these" too often. On the hand-written probes, about 35% of questions about an allowed node do: 73 of 200 for topic-scope-v2, 70 of 200 for topic-scope-v3 at -0.5. None of them fires a disallowed node, so a policy that logs none, the library default, refuses nothing; under
options.on_none: blockeach is a refusal. Neither the none offset nor a longer node description fixes it. Measure the rate on your own traffic before blocking on none. - Deterministic: no sampling. The same message, taxonomy and weights give the same answer.
Training
Source run: configs/typed_decisions_a3_large.yaml, base FacebookAI/xlm-roberta-large, 2 epochs, seed 42. Code and the plan it follows: docs/typed-decisions-plan.md in the training repository.
- Downloads last month
- -
Model tree for flowxai/topic-scope-v3
Base model
FacebookAI/xlm-roberta-large