Instructions to use Falconsai/LightDec_V2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Falconsai/LightDec_V2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("zero-shot-classification", model="Falconsai/LightDec_V2")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Falconsai/LightDec_V2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
LightDec_V2 (Long)
LightDec_V2 is a small, fast, calibrated decision model from Falcons.ai. Given a piece of state (an email, a ticket, an agent trace, a contract, a log, a JSON record) and one or more typed questions, it picks an answer from a closed set of options in a single forward pass and returns calibrated probabilities, so your application knows when to trust the answer and when to defer to a human or a larger model.
It is the long-context successor to Falconsai/LightDec. Every call now uses a 2,048-token window (LightDec v1 used 512 tokens by default and only switched to 2,048 for questions with more than 24 options), so long email threads, contracts and service logs fit without truncation.
| Architecture | FalconDec (encoder + permutation-equivariant option head) |
| Backbone | jhu-clsp/ettin-encoder-150m |
| Parameters | ~0.16B (fp16 weights ≈ 319 MB; int8 export available) |
| Context | 2,048 tokens per call |
| Question types | choice, noul (yes/no), score (ordinal levels) |
| Options per question | 2 to 96 in one pass; more via an automatic tournament |
| Test accuracy | 78.4% micro / 78.9% macro over 31,990 test items, 73 tasks |
| Calibration | ECE 0.029 (per-type, per-option-count temperature scaling) |
| Latency | ~10 ms per call on GPU (fp16), ~48 ms on CPU (int8) |
| Version | 1.0.0 (trained from the pretrained backbone, not fine-tuned from LightDec v1) |
What it does
LightDec_V2 answers the kind of small, structured questions that sit inside real systems: Which queue does this ticket go to? Does this email ask for a refund? How angry is the customer, 0–2? Is this agent step wrong? Is this prompt a jailbreak? Does this invoice violate the policy? It is not a generative model and never produces free text; it only chooses among the options you give it, which makes its output easy to validate, log and act on.
Three question types are supported, and several questions about the same state can be asked in one call:
choice: pick one of N labelled options (options can carry descriptions, e.g.{"billing": "Charges, invoices, refunds"}).noul: a yes/no judgment about a statement (returnsp_true).score: an ordinal scale such as["Calm", "Annoyed", "Furious"](also returns anexpected_level).
Every answer includes a confidence and a defer flag. With the default defer_threshold of 0.7, the model answered 68.0% of test questions and was 90.6% accurate on those; the remainder are flagged for review.
How it works
The question, the options and the state are packed into one sequence:
[CLS] question [SEP] [MASK] option_1 [MASK] option_2 ... [MASK] option_k [SEP] state [SEP]
The encoder reads the whole sequence once. The hidden vector at each [MASK] marker represents one option; it is combined with the [CLS] context and a question-type embedding, then passed through a 2-layer, 8-head set transformer in which options attend to each other without positional encoding, so the result does not depend on option order. An MLP produces one logit per option and a softmax turns them into a distribution. Finally a temperature, learned per question type and per option-count bucket (2, 3–5, 6–12, 13+ options) and stored in falcondec_config.json, calibrates the probabilities.
Questions with more than 96 options are resolved with a tournament: options are scored in chunks, the best of each chunk advance, and a final pass ranks the survivors.
Usage
The model ships with its own loader, falcondec_modeling.py, which requires torch, transformers, safetensors and huggingface_hub.
import importlib.util
from huggingface_hub import hf_hub_download
path = hf_hub_download("Falconsai/LightDec_V2", "falcondec_modeling.py")
spec = importlib.util.spec_from_file_location("falcondec_modeling", path)
fd = importlib.util.module_from_spec(spec)
spec.loader.exec_module(fd)
model, tok = fd.load_falcondec("Falconsai/LightDec_V2") # GPU if available, else CPU
state = {
"email": {
"subject": "Charged twice this month",
"body": "I was billed twice for my Pro plan in March. Please refund the duplicate charge. This is the second time!",
},
"customer": {"plan": "Pro", "customer_since": "2021"},
}
questions = {
"topic": {
"type": "choice",
"instructions": "What is this support email primarily about?",
"criteria": {
"billing": "Charges, invoices, payments, refunds or subscription costs",
"technical": "Something in the product is broken, failing or erroring",
"feature_request": "Asking for a new feature or an improvement",
"account": "Login, password, profile, seats or account settings",
"none": "None of these",
},
},
"refund": {"type": "noul", "instructions": "The email asks for money to be returned or credited"},
"anger": {"type": "score", "instructions": "How angry does the customer sound?",
"criteria": ["Calm", "Annoyed", "Furious"]},
}
out = fd.decide(model, tok, state, questions)
for key, r in out["answers"].items():
print(key, r["choice"], round(r["confidence"], 3), "DEFER" if r["defer"] else "")
Each result contains choice, choice_text, confidence, the full probs distribution and defer; noul answers add p_true and score answers add expected_level. Pass defer_threshold= to decide to trade coverage for accuracy. For lower-level control, fd.score_items(model, tok, items) scores raw {"state", "question", "options", "type"} items in batches.
CPU / small footprint. If the repository includes the int8 export (as LightDec v1 does in compact-int8/), load that folder the same way; the loader dequantizes the per-channel int8 weights automatically.
Evaluation
All numbers come from falcondec_report.json. The test split has 31,990 items across 73 tasks in 11 domains. Tasks marked held-out are flagged as such in the training report.
Overall
| Metric | Value |
|---|---|
| Micro accuracy | 78.4% |
| Macro accuracy (mean over tasks) | 78.9% |
| Negative log-likelihood | 0.563 |
| Brier score | 0.295 |
| Expected calibration error | 0.029 |
| AURC (area under risk–coverage) | 0.068 |
Mean absolute error on score questions (levels) |
0.51 |
| Coverage at defer threshold 0.7 | 68.0% |
| Accuracy on covered questions | 90.6% |
By domain
| Domain | Tasks | Test items | Accuracy |
|---|---|---|---|
| long_context | 3 | 450 | 100.0% |
| support | 3 | 1,505 | 98.2% |
| code | 9 | 3,854 | 91.1% |
| intents | 4 | 1,800 | 85.0% |
| agentic | 6 | 1,563 | 82.1% |
| policy | 11 | 5,500 | 80.6% |
| guardrails | 4 | 1,294 | 79.8% |
| workflows | 4 | 2,000 | 77.3% |
| reasoning | 18 | 9,224 | 71.8% |
| tev1_benchmark | 7 | 2,800 | 65.4% |
| classification | 4 | 2,000 | 61.2% |
Head-to-head against the proof_v2 baseline
On 72 shared tasks (up to 120 items each), LightDec_V2 averaged 79.6% against 46.7% for the proof_v2 baseline, winning on 70 tasks, tying on 1 and losing on 1 (policy/invoice_total_transfer).
Per-task head-to-head
| Task | n | LightDec_V2 | proof_v2 | Δ |
|---|---|---|---|---|
long/email_thread |
120 | 100.0% | 5.0% | +95.0 |
agenttrek/next_action_type |
120 | 87.5% | 7.5% | +80.0 |
policy/table_extreme_transfer |
120 | 92.5% | 14.2% | +78.3 |
long/contract_clause |
120 | 100.0% | 23.3% | +76.7 |
policy/access_control_transfer |
120 | 100.0% | 23.3% | +76.7 |
long/service_log |
120 | 100.0% | 24.2% | +75.8 |
policy/count_threshold_transfer |
120 | 86.7% | 10.8% | +75.8 |
bigclonebench/clone |
120 | 96.7% | 24.2% | +72.5 |
policy/return_window_transfer |
120 | 100.0% | 34.2% | +65.8 |
snli/contradicts |
120 | 100.0% | 35.8% | +64.2 |
policy/refund_approval_transfer |
120 | 95.8% | 32.5% | +63.3 |
gsm8k/math |
120 | 73.3% | 12.5% | +60.8 |
hotpotqa/retrieve |
120 | 87.5% | 29.2% | +58.3 |
civil_comments/toxic |
120 | 93.3% | 35.8% | +57.5 |
snli/nli |
120 | 90.8% | 34.2% | +56.7 |
triage/support_email |
120 | 95.8% | 39.2% | +56.7 |
mnli/claim |
120 | 87.5% | 31.7% | +55.8 |
scitail/support |
120 | 95.8% | 42.5% | +53.3 |
agenttrek/finish_now |
120 | 78.3% | 25.8% | +52.5 |
tev1_test/ag_news |
120 | 92.5% | 42.5% | +50.0 |
typed_decisions/agent_trace_observability |
120 | 80.8% | 33.3% | +47.5 |
hotpotqa/comparison_yes_no |
26 | 92.3% | 46.2% | +46.2 |
typed_decisions/customer_service |
120 | 74.2% | 28.3% | +45.8 |
policy/table_compare_transfer |
120 | 91.7% | 49.2% | +42.5 |
jailbreak/detect |
120 | 98.3% | 56.7% | +41.7 |
ag_news/topic |
120 | 85.0% | 45.0% | +40.0 |
openbookqa/mcq |
120 | 65.0% | 28.3% | +36.7 |
tev1_test/mnli |
120 | 75.0% | 39.2% | +35.8 |
counsel/critique_quality |
120 | 60.8% | 25.8% | +35.0 |
typed_decisions/invoice_processing |
120 | 81.7% | 47.5% | +34.2 |
humaneval/completion |
119 | 84.9% | 51.3% | +33.6 |
clinc150/intent |
120 | 97.5% | 64.2% | +33.3 |
hellaswag/continuation |
120 | 60.8% | 29.2% | +31.7 |
yelp/score |
120 | 62.5% | 32.5% | +30.0 |
mbpp/bugspot |
120 | 85.8% | 57.5% | +28.3 |
tev1_test/boolq |
120 | 84.2% | 55.8% | +28.3 |
commonsense_qa/mcq |
120 | 68.3% | 40.8% | +27.5 |
anli/nli |
120 | 59.2% | 32.5% | +26.7 |
typed_decisions/security_incidents |
120 | 75.8% | 50.0% | +25.8 |
arc_challenge/mcq |
120 | 53.3% | 30.8% | +22.5 |
policy/free_shipping_transfer |
120 | 92.5% | 70.8% | +21.7 |
agentharm/refuse |
120 | 66.7% | 45.0% | +21.7 |
massive_en/intent |
120 | 95.0% | 74.2% | +20.8 |
arc_easy/mcq |
120 | 65.0% | 45.8% | +19.2 |
tev1_test/banking77 |
120 | 71.7% | 55.0% | +16.7 |
boolq/yes_no |
120 | 82.5% | 66.7% | +15.8 |
policy/table_count_transfer |
120 | 31.7% | 15.8% | +15.8 |
mbpp/solution |
120 | 100.0% | 85.0% | +15.0 |
policy/invoice_overdue_transfer |
120 | 80.8% | 65.8% | +15.0 |
tev1_test/routing |
120 | 44.2% | 29.2% | +15.0 |
bitext/category |
120 | 100.0% | 85.8% | +14.2 |
qasc/mcq |
120 | 99.2% | 85.8% | +13.3 |
winogrande/blank |
120 | 69.2% | 55.8% | +13.3 |
emotion/6way |
120 | 49.2% | 36.7% | +12.5 |
tev1_test/sst5 |
120 | 39.2% | 27.5% | +11.7 |
aqua_rat/math |
120 | 35.0% | 24.2% | +10.8 |
sst5/score |
120 | 38.3% | 28.3% | +10.0 |
devign/vulnerability |
120 | 66.7% | 56.7% | +10.0 |
mmlu/mcq |
120 | 40.8% | 30.8% | +10.0 |
sciq/mcq |
120 | 96.7% | 86.7% | +10.0 |
tev1_test/policy |
120 | 50.0% | 40.8% | +9.2 |
codexglue/func_name |
120 | 99.2% | 91.7% | +7.5 |
policy/sla_urgency_transfer |
120 | 51.7% | 44.2% | +7.5 |
bitext/route |
120 | 100.0% | 92.5% | +7.5 |
counsel/step_has_error |
120 | 82.5% | 75.0% | +7.5 |
codexglue/code_to_doc |
120 | 100.0% | 93.3% | +6.7 |
snli/must_be_true |
120 | 99.2% | 92.5% | +6.7 |
banking77/intent |
120 | 92.5% | 86.7% | +5.8 |
prompt_injections/detect |
116 | 59.5% | 54.3% | +5.2 |
codexglue/doc_to_code |
120 | 98.3% | 97.5% | +0.8 |
codexglue/lang_id |
120 | 100.0% | 100.0% | +0.0 |
policy/invoice_total_transfer |
120 | 45.0% | 52.5% | -7.5 |
Support-email triage (bundled benchmark)
benchmarks/triage_support_email_test.jsonl contains 101 support emails, each with five typed questions and reference answers. It doubles as a worked example of the input format.
| Question | Type | Accuracy |
|---|---|---|
topic |
choice (5-way) | 99.0% |
refund |
noul (yes/no) | 100.0% |
breakage |
score (4 levels) | 99.0% |
anger |
score (3 levels) | 92.1% |
judgment |
choice (3-way) | 82.2% |
All five answers together route an email to the correct handling pile 94.1% of the time.
Long context
Three long-document tasks exercise the 2,048-token window (150 items each): long/contract_clause, long/email_thread and long/service_log. LightDec_V2 scored 100% on all three. These are synthetic, in-distribution tasks, so treat them as a check that long inputs are read end to end rather than as a measure of general long-document reasoning.
TEV1 transfer benchmark
TEV1 is a separate decision benchmark. Some of its tasks draw on the same public sources as LightDec's training data; the table marks which.
| Task | n | Accuracy | Overlaps LightDec training source |
|---|---|---|---|
tev1_test/ag_news |
150 | 92.0% | yes |
tev1_test/banking77 |
200 | 72.0% | no |
tev1_test/boolq |
200 | 83.5% | yes |
tev1_test/mnli |
300 | 72.3% | yes |
tev1_test/policy |
1200 | 52.9% | no |
tev1_test/routing |
600 | 45.8% | no |
tev1_test/sst5 |
150 | 39.3% | no |
Overall TEV1 accuracy is 58.4% (2,800 items), and 51.8% on the tasks with no source overlap. For reference, the report lists published results for the 4B-parameter TEV1 model of 88% on its main decisions set and 100% on its policy transfer set; those are different splits and a model roughly 25× larger, so the figures are context rather than a like-for-like comparison.
Full per-task test results (73 tasks)
| Domain | Task | n | Accuracy | Chance | ECE | Held-out |
|---|---|---|---|---|---|---|
| agentic | agenttrek/finish_now |
151 | 78.8% | 50.0% | 0.075 | |
| agentic | agenttrek/next_action_type |
487 | 86.4% | 19.1% | 0.092 | |
| agentic | counsel/critique_quality |
201 | 62.7% | 33.3% | 0.271 | |
| agentic | counsel/step_has_error |
201 | 83.1% | 50.0% | 0.148 | |
| agentic | hotpotqa/comparison_yes_no |
26 | 92.3% | 50.0% | 0.082 | |
| agentic | hotpotqa/retrieve |
497 | 89.3% | 16.7% | 0.043 | |
| classification | ag_news/topic |
500 | 90.0% | 25.0% | 0.052 | |
| classification | emotion/6way |
500 | 47.2% | 16.7% | 0.211 | ✓ |
| classification | sst5/score |
500 | 40.6% | 20.0% | 0.069 | ✓ |
| classification | yelp/score |
500 | 66.8% | 20.0% | 0.091 | |
| code | bigclonebench/clone |
500 | 96.2% | 50.0% | 0.031 | |
| code | codexglue/code_to_doc |
504 | 99.0% | 26.2% | 0.010 | |
| code | codexglue/doc_to_code |
504 | 97.6% | 27.9% | 0.013 | |
| code | codexglue/func_name |
467 | 95.5% | 25.5% | 0.039 | |
| code | codexglue/lang_id |
504 | 100.0% | 23.5% | 0.001 | |
| code | devign/vulnerability |
500 | 63.2% | 50.0% | 0.070 | |
| code | humaneval/completion |
119 | 84.9% | 39.4% | 0.140 | ✓ |
| code | mbpp/bugspot |
256 | 85.5% | 40.6% | 0.037 | |
| code | mbpp/solution |
500 | 97.8% | 25.0% | 0.020 | |
| guardrails | agentharm/refuse |
416 | 70.0% | 50.0% | 0.089 | ✓ |
| guardrails | civil_comments/toxic |
500 | 92.6% | 50.0% | 0.051 | |
| guardrails | jailbreak/detect |
262 | 97.3% | 50.0% | 0.031 | |
| guardrails | prompt_injections/detect |
116 | 59.5% | 50.0% | 0.333 | ✓ |
| intents | banking77/intent |
500 | 92.2% | 32.1% | 0.037 | ✓ |
| intents | banking77/intent_77 |
300 | 58.0% | 1.3% | 0.241 | ✓ |
| intents | clinc150/intent |
500 | 97.0% | 12.4% | 0.015 | |
| intents | massive_en/intent |
500 | 93.0% | 13.0% | 0.028 | |
| long_context | long/contract_clause |
150 | 100.0% | 20.0% | 0.000 | |
| long_context | long/email_thread |
150 | 100.0% | 25.0% | 0.000 | |
| long_context | long/service_log |
150 | 100.0% | 20.0% | 0.001 | |
| policy | policy/access_control_transfer |
500 | 100.0% | 33.3% | 0.002 | |
| policy | policy/count_threshold_transfer |
500 | 88.0% | 10.5% | 0.048 | |
| policy | policy/free_shipping_transfer |
500 | 93.8% | 50.0% | 0.041 | |
| policy | policy/invoice_overdue_transfer |
500 | 82.4% | 50.0% | 0.020 | |
| policy | policy/invoice_total_transfer |
500 | 47.4% | 50.0% | 0.082 | |
| policy | policy/refund_approval_transfer |
500 | 96.2% | 33.3% | 0.019 | |
| policy | policy/return_window_transfer |
500 | 100.0% | 33.3% | 0.026 | |
| policy | policy/sla_urgency_transfer |
500 | 60.6% | 25.0% | 0.131 | |
| policy | policy/table_compare_transfer |
500 | 92.8% | 50.0% | 0.049 | |
| policy | policy/table_count_transfer |
500 | 32.4% | 12.1% | 0.541 | |
| policy | policy/table_extreme_transfer |
500 | 93.2% | 12.2% | 0.053 | |
| reasoning | anli/nli |
498 | 48.6% | 33.3% | 0.146 | |
| reasoning | aqua_rat/math |
247 | 34.0% | 20.0% | 0.080 | |
| reasoning | arc_challenge/mcq |
500 | 49.0% | 25.0% | 0.168 | ✓ |
| reasoning | arc_easy/mcq |
500 | 61.2% | 25.0% | 0.122 | ✓ |
| reasoning | boolq/yes_no |
500 | 82.2% | 50.0% | 0.080 | |
| reasoning | commonsense_qa/mcq |
493 | 64.3% | 20.0% | 0.143 | |
| reasoning | gsm8k/math |
500 | 70.0% | 25.0% | 0.085 | |
| reasoning | hellaswag/continuation |
500 | 57.6% | 25.0% | 0.075 | |
| reasoning | mmlu/mcq |
500 | 39.0% | 25.0% | 0.155 | ✓ |
| reasoning | mnli/claim |
500 | 86.0% | 33.3% | 0.081 | |
| reasoning | openbookqa/mcq |
500 | 57.2% | 25.0% | 0.218 | |
| reasoning | qasc/mcq |
500 | 98.6% | 12.5% | 0.008 | |
| reasoning | sciq/mcq |
498 | 95.4% | 25.0% | 0.021 | |
| reasoning | scitail/support |
500 | 96.2% | 50.0% | 0.030 | |
| reasoning | snli/contradicts |
500 | 99.0% | 33.3% | 0.019 | |
| reasoning | snli/must_be_true |
500 | 98.6% | 33.3% | 0.043 | |
| reasoning | snli/nli |
988 | 89.5% | 33.3% | 0.108 | |
| reasoning | winogrande/blank |
500 | 66.2% | 50.0% | 0.116 | |
| support | bitext/category |
500 | 100.0% | 16.1% | 0.001 | |
| support | bitext/route |
500 | 100.0% | 19.0% | 0.000 | |
| support | triage/support_email |
505 | 94.5% | 32.3% | 0.046 | |
| tev1_benchmark | tev1_test/ag_news |
150 | 92.0% | 25.0% | 0.098 | ✓ |
| tev1_benchmark | tev1_test/banking77 |
200 | 72.0% | 17.8% | 0.137 | ✓ |
| tev1_benchmark | tev1_test/boolq |
200 | 83.5% | 50.0% | 0.050 | ✓ |
| tev1_benchmark | tev1_test/mnli |
300 | 72.3% | 33.3% | 0.075 | ✓ |
| tev1_benchmark | tev1_test/policy |
1200 | 52.9% | 33.3% | 0.059 | ✓ |
| tev1_benchmark | tev1_test/routing |
600 | 45.8% | 20.0% | 0.038 | ✓ |
| tev1_benchmark | tev1_test/sst5 |
150 | 39.3% | 20.0% | 0.108 | ✓ |
| workflows | typed_decisions/agent_trace_observability |
500 | 73.4% | 30.0% | 0.222 | |
| workflows | typed_decisions/customer_service |
500 | 76.4% | 28.0% | 0.210 | |
| workflows | typed_decisions/invoice_processing |
500 | 82.6% | 35.0% | 0.195 | |
| workflows | typed_decisions/security_incidents |
500 | 76.8% | 34.0% | 0.235 |
Training
| Initialization | Pretrained jhu-clsp/ettin-encoder-150m (trained from scratch on top of the backbone; no LightDec v1 weights) |
| Data | 1,023,814 training / 24,200 validation / 31,990 test examples |
| Sources | Public classification, NLI, QA, code, intent, guardrail and agent-trace datasets converted into typed decisions, plus synthetic policy-transfer, workflow, support-triage and long-context tasks and TEV1 builders. mind2web was excluded; no custom data was added. |
| Preset | long (2,048-token sequences, 1,500 long-context training tasks) |
| Objective | Cross-entropy with spherical-score (0.5) and ranked-probability-score (1.0) terms for ordinal questions; 8% "none of the above" augmentation; task sampling α = 0.5 |
| Optimizer | Learning rate 8e-5 (encoder) / 6e-4 (head), layer-wise decay 0.9, weight decay 0.01, 6% warmup, gradient clip 1.0, EMA 0.999 |
| Batching | Token-budget batches of 65,536 tokens (batch size 128) |
| Schedule | 10 epochs; checkpoint from epoch 4 selected on validation macro accuracy |
| Calibration | Temperature per (question type × option-count bucket) fit on validation |
| Compute | 292 minutes on one NVIDIA RTX PRO 6000 Blackwell Server Edition |
| Software | PyTorch 2.9.0 (CUDA 13.0), Transformers 5.17.0, Python 3.12 |
| Epoch | Train loss | Train acc | Val macro acc | Val NLL |
|---|---|---|---|---|
| 1 | 0.706 | 73.4% | 81.9% | 0.408 |
| 2 | 0.431 | 84.7% | 84.5% | 0.362 |
| 3 | 0.343 | 87.9% | 85.0% | 0.392 |
| 4 (selected) | 0.277 | 90.4% | 85.2% | 0.453 |
| 5 | 0.220 | 92.6% | 84.9% | 0.537 |
| 6 | 0.170 | 94.5% | 84.7% | 0.671 |
| 7 | 0.127 | 96.0% | 84.5% | 0.843 |
| 8 | 0.095 | 97.1% | 84.5% | 1.034 |
| 9 | 0.071 | 97.9% | 84.4% | 1.211 |
| 10 | 0.057 | 98.4% | 84.4% | 1.419 |
Validation accuracy peaked at epoch 4 while validation NLL kept rising after epoch 2, so the selected checkpoint is somewhat overconfident before calibration. The fitted temperatures (roughly 1.6 to 2.5) correct for this; they are part of the checkpoint and applied automatically by decide and score_items.
Speed
Measured with decide on a single state; latency grows only slightly as more questions are asked about the same state.
| Device | Weights | Questions per call | p50 | p95 |
|---|---|---|---|---|
| cuda | fp16 | 1 | 10.05 ms | 10.43 ms |
| cuda | fp16 | 5 | 11.32 ms | 11.41 ms |
| cuda | fp16 | 10 | 12.57 ms | 12.69 ms |
| cpu fp32 | int8 file | 1 | 48.4 ms | 48.8 ms |
GPU throughput: about 2,589 decisions per second with batching.
Intended use
LightDec_V2 is built for high-volume, low-latency decisions inside software: ticket and email routing, intent detection, triage scoring, content and prompt-safety screening, policy and rule checks over structured records, agent step verification and "should the agent stop now" checks, and pre-filtering before a more expensive LLM call. The calibrated confidence and defer flag are designed for human-in-the-loop and cascade setups in which uncertain cases escalate.
Limitations
- Closed-set only. The model can only choose among the options you supply. If the right answer is not listed, it will still pick something; include a
noneoption when that can happen. - Weak at multi-step reasoning and arithmetic. Knowledge- and math-heavy tasks are well below the rest: MMLU 39.0%, AQuA-RAT 34.0%, ANLI 48.6%, ARC-Challenge 49.0%. Some policy tasks that require computing over a table are also weak:
policy/table_count_transferscores 32.4% with poor calibration (ECE 0.54), andpolicy/invoice_total_transfer(47.4%) is below chance. Use a larger model for anything that needs counting or totalling. - Fine-grained and subjective labels. Accuracy drops on 77-way banking intents (58.0%), six-way emotion (47.2%) and five-level sentiment (SST-5 40.6%, Yelp 66.8%).
- Transfer to unfamiliar rule sets is limited. On TEV1 tasks without source overlap, accuracy is 51.8%, and routing and policy transfer (45.8% and 52.9%) are the weakest; validate on your own policies before relying on it.
- Guardrail coverage is partial. Jailbreak detection is strong (97.3%), but prompt-injection detection (59.5%) and harmful-request refusal (70.0%) are not reliable enough to be a sole safety layer.
- Calibration is aggregate. Overall ECE is low, but a few tasks (for example
counsel/critique_quality,prompt_injections/detect, the table-count policy task) are noticeably miscalibrated; check calibration on your own distribution before tuningdefer_threshold. - English only, and questions are truncated to 96 tokens and each option to 24 tokens by default (
option_tokenscan raise the per-option budget). - Long-context scores are synthetic. The 100% long-context results are in-distribution checks, not evidence of general long-document reasoning.
Files
| File | Description |
|---|---|
model.safetensors |
fp16 weights |
falcondec_config.json |
Head configuration, special-token ids, defer threshold and calibration temperatures |
falcondec_modeling.py |
Model definition, loader (load_falcondec) and inference API (decide, score_items) |
falcondec_report.json |
Full training and evaluation report |
encoder/, tokenizer/ |
Ettin encoder config and tokenizer |
benchmarks/triage_support_email_test.jsonl |
101-email support-triage benchmark with reference answers |
Citation
@misc{falconsai_lightdec_v2_2026,
title = {LightDec_V2: A Long-Context, Calibrated Decision Model},
author = {Falcons.ai},
year = {2026},
howpublished = {\url{https://huggingface.co/Falconsai/LightDec_V2}}
}
Built on the Ettin encoder from JHU CLSP (jhu-clsp/ettin-encoder-150m).
This card is generated from the surgical record itself; the package's
lineage.intoto.jsonl is the signed source of truth (verify it free at
the Surgeon's public verifier or with the bundled verify_attestation.py).
Architecture
- Identification: NLP · Small Language Model (SLM) (98% confidence)
- Source format:
safetensors· Intended task: not declared config.json: synthesized from the anatomy (no source config.json)- Source license: apache-2.0
- Lineage chain: 1 surgery (no prior attestation reachable) · Falconsai/LightDec_V2
- Post-surgery totals: 159,654,157 parameters · 168 tensors
- Compute estimate: 15.02439 GFLOPs (comparison metric, not a measurement)
Provenance & operations
- Parents: Falconsai/LightDec_V2/model.safetensors
- Operations performed: load×1
- Weight merges recorded: 0
- Quantized tensors (F32→F16): 0
Surgery Log (ordered)
- load — hub:Falconsai/LightDec_V2/model.safetensors (319.3 MB, safetensors)
Validation
- Tissue imaging: not run
- Structural integrity is testable offline via the packaged
load_and_test.py.
Compliance note
The signed attestation + this card together document model composition, modification history, and validation evidence — the record structure technical-documentation obligations (e.g. EU AI Act Annex IV) ask for. This is evidence, not legal advice.
Operated with Model Surgeon — verify this package at https://surgeon.falcons.ai/verify © 2026 FALCONS.AI — Model Surgeon record format. The model weights remain their owner's.
- Downloads last month
- -