|
Download README.md from SamCodeManMk2/tron-1b: direct link, hf CLI and curl.
- Browser
- Download file 6.29 kB
-
https://huggingface.co/SamCodeManMk2/tron-1b/resolve/main/README.md
- Command line
-
hf download hf://SamCodeManMk2/tron-1b/README.md
-
curl -L -o README.md https://huggingface.co/SamCodeManMk2/tron-1b/resolve/main/README.md
6.29 kB
| license: cc-by-nc-4.0 | |
| language: | |
| - en | |
| library_name: troncore | |
| base_model: jhu-clsp/ettin-encoder-1b | |
| pipeline_tag: zero-shot-classification | |
| tags: | |
| - decision-model | |
| - zero-shot-classification | |
| - text-classification | |
| - intent-classification | |
| - prompt-injection | |
| - calibration | |
| - modernbert | |
| # Tron-1B | |
| **Tron-1B answers typed questions about text or JSON in a single pass: choose one option, rate on a scale, or answer yes/no, with calibrated probabilities.** It is built for the fast decisions around an AI application: routing, triage, safety screening, intent detection, and workflow automation. | |
| - **Accurate:** beats Jev 1.13.0 on all four of its published benchmarks (below). | |
| - **Fast:** 16.7 ms per decision (p50) on one GPU. | |
| - **Calibrated:** its confidence scores track how often it is right, so you can gate actions on them. | |
| - **Any question, at request time:** you write the question and the options; no retraining for new label sets. | |
|  | |
| ## Try it | |
| [](https://colab.research.google.com/#fileId=https://huggingface.co/SamCodeManMk2/tron-1b/blob/main/tron-1b-quickstart.ipynb) | |
| The [quickstart notebook](https://huggingface.co/SamCodeManMk2/tron-1b/blob/main/tron-1b-quickstart.ipynb) runs on Colab's free T4 GPU: routing, safety screening, all 77 Banking77 intents, | |
| a speed test, and a form for your own questions. | |
| ## Quickstart | |
| ```bash | |
| pip install https://huggingface.co/SamCodeManMk2/tron-1b/resolve/main/troncore-1.0.0-py3-none-any.whl | |
| ``` | |
| ```python | |
| from troncore import Engine | |
| eng = Engine("SamCodeManMk2/tron-1b") # downloads once, then runs locally (GPU if available) | |
| answers = eng.decide( | |
| {"subject": "Duplicate charge on invoice #4411", | |
| "body": "We were billed twice for March. Refund the duplicate today or we cancel our plan."}, | |
| { | |
| "department": {"type": "choice", "instructions": "Which team should handle this?", | |
| "criteria": {"billing": "invoices, payments, refunds", "technical": "bugs, outages", | |
| "sales": "pricing, new contracts", "other": "anything else"}}, | |
| "urgency": {"type": "score", "instructions": "How urgent is this?", | |
| "criteria": ["not urgent", "this week", "today", "critical"]}, | |
| "churn_risk": {"type": "yesno", "instructions": "Does the customer threaten to cancel?"}, | |
| }, | |
| ) | |
| answers["department"]["choice"] # 'billing' | |
| answers["churn_risk"]["probability"] # P(yes) | |
| answers["department"]["confidence"] # calibrated confidence of the top option | |
| ``` | |
| Every answer includes `probabilities`, `confidence`, `margin` (top minus second) and `entropy`. Pass | |
| `min_confidence=0.8` to get `abstain: true` on uncertain answers, so they can be routed to a person or a larger model. | |
| Questions with more than 128 options are handled automatically by scoring them in rounds. Long inputs can be read | |
| in sliding windows with `windows="all"`. | |
| ### Question types | |
| | type | options | answer | | |
| |---|---|---| | |
| | `choice` | `criteria`: list of labels, or dict label → description | `choice` + a probability per label | | |
| | `score` | `criteria`: ordered list of levels | `score` (expected level), `level`, a probability per level | | |
| | `yesno` (alias `noul`) | fixed no / yes | `answer` (bool) + `probability` of yes | | |
| ## Results | |
|  | |
| | Benchmark | Tron-1B | Jev 1.13.0 | | |
| |---|---|---| | |
| | Banking77 (intent, 77 labels)¹ | **94.0** | 87.0 | | |
| | AG News (topic) | **93.9** | 91.0 | | |
| | typed-decisions (2,000 business decisions) | **79.6** | 72.7 | | |
| | DAIR emotion | **92.9** | 48.0 | | |
| | Latency, p50 per decision² | **16.7 ms** | 236–276 ms | | |
| **Tasks Tron-1B never trained on** (whole task families held out of training): | |
| | Task | Accuracy | | |
| |---|---| | |
| | IMDB sentiment | 94.7 | | |
| | Prompt-injection detection (deepset) | 90.5 | | |
| | MASSIVE intent (59 labels) | 88.4 | | |
| | XNLI entailment (English) | 86.7 | | |
| | BoolQ reading comprehension | 79.2 | | |
| | CLINC intent (151 labels) | 62.8 | | |
| | SST-5 graded sentiment | 54.8 | | |
| ¹ Jev's published Banking77 result used 72 labels; Tron's uses all 77. Jev figures are its published numbers, not re-run by us. | |
| ² Tron: one question on one NVIDIA GB10 GPU. Jev: independently measured p50 through its hosted API, which includes network time. | |
| Tron-1B was trained on the training splits of Banking77, AG News and DAIR emotion and evaluated on their official | |
| test splits. All results measured 2026-09-28 with troncore 1.0. | |
| ## How it works | |
| Tron-1B is a 1.1B-parameter bidirectional encoder (initialised from | |
| [ettin-encoder-1b](https://huggingface.co/jhu-clsp/ettin-encoder-1b)) with a decision head. For each question, the | |
| question, the input and every option are encoded together. Each option is pooled into a vector, and a small | |
| attention layer compares the options with each other and with the question before they are scored. This is what lets | |
| it separate close labels such as "card not arrived" and "card delivery estimate". Probabilities are calibrated per | |
| question type and option count. | |
| It was trained on about 1.2 million typed decisions built from public classification, entailment, safety, routing, | |
| reasoning, preference and business-workflow datasets. | |
| ## Limitations | |
| - **It decides; it does not write.** Answers are always one of the options you give it. | |
| - **Very large unseen label sets are its weakest area** (62.8% on CLINC's 151 intents zero-shot). For big label | |
| sets, short descriptive label names help, and so does a few hundred examples of fine-tuning. | |
| - **Long inputs:** accuracy is best under about 2,000 tokens of input. Send the relevant section rather than a whole | |
| document, or use `windows="all"`. | |
| - **English first.** It handles other languages, but was mostly trained and evaluated in English. | |
| - **Not a safety guarantee.** Use its safety and injection judgements as one layer of defence, with confidence | |
| gating, not as the only one. | |
| ## License | |
| The model weights are released under [CC BY-NC 4.0](https://creativecommons.org/licenses/by-nc/4.0/): free to use, | |
| share and adapt for non-commercial purposes, with attribution. The `troncore` runtime is Apache-2.0. The base | |
| encoder is MIT-licensed; see `NOTICE`. | |