topic-scope-v2
The model behind the topic_scope detector in flowx-border. Given a message and a taxonomy a policy supplies, it answers which node the message is about, or "none of these", with a calibrated probability.
It replaces flowxai/topic-scope, a bi-encoder that scored cosine similarity to each node. That model compares meanings and cannot read a node described by exclusion, and it has no way to say that a message is about none of the nodes. This one reads the message, the question and every node together in a small decision head.
How it runs
XLM-RoBERTa base, fine-tuned, plus a two-layer decision head trained from scratch. The late encoding: the message, the question and each node are encoded as separate short sequences, and meet only in the head. Node states are therefore computed once per taxonomy and cached; a scan encodes the message and runs the head.
| file | what | bytes | sha256 |
|---|---|---|---|
onnx/model.int8.onnx |
encoder, embedding Gather in int8, outputs last_hidden_state | 532,929,271 | fb08eeecd7967d8d... |
onnx/head.onnx |
decision head, fp32, takes hidden states and returns one logit per option | 59,118,915 | 3b750a812d9fb28f... |
decision_config.json |
question text, the none option, token budgets, none offset, temperatures | 1,019 | b3b24d99c9a42012... |
tokenizer.json |
XLM-R tokenizer | 17,098,086 | acbd420e2269cdc1... |
config.json |
encoder config | 742 | 0af0984d803b4e13... |
Token budgets: message 128, question 48, each node 64. A node is rendered as {key}: {text}, and the none option is always offered as none: none of these topics; the message is about something else. Its logit gets -0.5 before calibration, the offset fitted on validation. Probabilities are a softmax at a temperature fitted per option-count bucket on validation rows, listed in decision_config.json.
Options are packed without padding for the head: it has no position embedding, so this is the function it was trained as, over fewer positions.
Evaluation
Synthetic, and held out by deployment. The corpus is generated: 299 taxonomies and 135,747 messages in the 26 languages, written by claude-sonnet-5 (most messages) and gpt-5.6-luna, and every validation and test row checked by a judge model, a different model from the writer for 97% of messages. Splits are by taxonomy. unseen_type taxonomies come from deployment types no training taxonomy came from (airline, legal services, telecom); unseen_schema taxonomies are new taxonomies of trained types. Neither is a measurement on real traffic, and the numbers below should be read as how the model transfers to taxonomies it has not seen, not as a production error rate.
Accuracy is top-1 over the offered nodes plus none, after the none offset. The comparison is the model this one replaces, on the same rows, with its none bar (0.875) chosen on validation.
| test | this model | flowxai/topic-scope | rows | ECE |
|---|---|---|---|---|
| unseen_type | 0.804 | 0.479 | 11,497 | 0.0084 |
| unseen_schema | 0.808 | 0.471 | 11,071 | 0.0187 |
| v1 topic_scope test, recast | 0.962 | 442 | 0.0238 |
Per language
Every language is reported, the weak ones included. Maltese is the weakest, as it is for every model on this base: Maltese is not in XLM-RoBERTa's pretraining set, and fine-tuning data cannot give the encoder a representation it never had. Irish is second.
| language | unseen_type | flowxai/topic-scope | rows | unseen_schema | flowxai/topic-scope | rows |
|---|---|---|---|---|---|---|
| az | 0.753 | 0.448 | 413 | 0.816 | 0.458 | 705 |
| bg | 0.804 | 0.449 | 445 | 0.812 | 0.491 | 485 |
| cs | 0.801 | 0.468 | 412 | 0.808 | 0.458 | 500 |
| da | 0.833 | 0.463 | 551 | 0.794 | 0.478 | 209 |
| de | 0.771 | 0.447 | 546 | 0.839 | 0.530 | 217 |
| el | 0.814 | 0.450 | 462 | 0.789 | 0.528 | 123 |
| en | 0.792 | 0.473 | 471 | 0.742 | 0.408 | 120 |
| es | 0.851 | 0.502 | 496 | 0.802 | 0.501 | 343 |
| et | 0.814 | 0.425 | 489 | 0.796 | 0.481 | 343 |
| fi | 0.851 | 0.451 | 495 | 0.798 | 0.499 | 573 |
| fr | 0.854 | 0.483 | 493 | 0.825 | 0.481 | 576 |
| ga | 0.710 | 0.384 | 427 | 0.676 | 0.380 | 732 |
| hr | 0.842 | 0.502 | 430 | 0.834 | 0.456 | 735 |
| hu | 0.830 | 0.517 | 383 | 0.843 | 0.468 | 477 |
| it | 0.847 | 0.559 | 372 | 0.839 | 0.460 | 465 |
| lt | 0.816 | 0.501 | 403 | 0.822 | 0.499 | 365 |
| lv | 0.805 | 0.501 | 389 | 0.797 | 0.459 | 364 |
| mt | 0.624 | 0.405 | 412 | 0.556 | 0.410 | 117 |
| nl | 0.818 | 0.542 | 406 | 0.750 | 0.473 | 112 |
| pl | 0.790 | 0.514 | 434 | 0.793 | 0.510 | 300 |
| pt | 0.851 | 0.521 | 449 | 0.824 | 0.452 | 301 |
| ro | 0.797 | 0.502 | 454 | 0.838 | 0.486 | 401 |
| sk | 0.793 | 0.529 | 425 | 0.837 | 0.497 | 406 |
| sl | 0.814 | 0.473 | 414 | 0.848 | 0.478 | 709 |
| sv | 0.823 | 0.527 | 419 | 0.840 | 0.490 | 686 |
| tr | 0.784 | 0.445 | 407 | 0.826 | 0.467 | 707 |
Per register
What kind of message it was. _hole rows had their correct node withheld from the options, so the right answer is none. negated_near_miss is a message that matches a node's words but falls under what the node excludes, and it is the weakest register on unseen types.
| register | unseen_type | rows | unseen_schema | rows |
|---|---|---|---|---|
| in_domain_uncovered | 0.865 | 814 | 0.872 | 693 |
| in_node | 0.806 | 4661 | 0.819 | 4485 |
| in_node_hole | 0.801 | 1521 | 0.766 | 1531 |
| negated_near_miss | 0.667 | 672 | 0.771 | 741 |
| negated_near_miss_hole | 0.782 | 87 | 0.778 | 63 |
| out_of_taxonomy | 0.979 | 817 | 0.995 | 645 |
| sibling_near_miss | 0.778 | 2212 | 0.795 | 2132 |
| sibling_near_miss_hole | 0.738 | 713 | 0.694 | 781 |
Export verification
300 test rows, the torch fp32 model on its training layout against the ONNX pipeline (onnxruntime CPU, one thread, int8 encoder, packed head): 0 changed their answer, and the largest logit difference was 0.16385. The head took 20.43 ms at p50 and 47.24 ms at p95 on those rows, one thread, on an Apple M-series laptop; the library's budget test measures the whole detector.
Limits
- Trained on taxonomies of 4 to 37 nodes. Head cost grows with the square of the total node text, so the library shortlists larger taxonomies to the nearest nodes by the model's own similarity term first, and records that it did.
- A node description past 64 tokens is cut. Write the part that decides first.
- The head was trained on four question families (topic_scope v1 and v2, injection and policy questions). The library asks it the topic question only. Its policy-question accuracy is not good enough to ship: about half of minimal pairs get both halves right.
- Deterministic: no sampling. The same message, taxonomy and weights give the same answer.
Training
Source run: configs/typed_decisions_a3.yaml, base FacebookAI/xlm-roberta-base, 2 epochs, seed 42. Code and the plan it follows: docs/typed-decisions-plan.md in the training repository.
- Downloads last month
- -
Model tree for flowxai/topic-scope-v2
Base model
FacebookAI/xlm-roberta-base