LightDec_V2 (Long)

View in Model Surgeon

LightDec_V2 is a small, fast, calibrated decision model from Falcons.ai. Given a piece of state (an email, a ticket, an agent trace, a contract, a log, a JSON record) and one or more typed questions, it picks an answer from a closed set of options in a single forward pass and returns calibrated probabilities, so your application knows when to trust the answer and when to defer to a human or a larger model.

It is the long-context successor to Falconsai/LightDec. Every call now uses a 2,048-token window (LightDec v1 used 512 tokens by default and only switched to 2,048 for questions with more than 24 options), so long email threads, contracts and service logs fit without truncation.

Architecture FalconDec (encoder + permutation-equivariant option head)
Backbone jhu-clsp/ettin-encoder-150m
Parameters ~0.16B (fp16 weights ≈ 319 MB; int8 export available)
Context 2,048 tokens per call
Question types choice, noul (yes/no), score (ordinal levels)
Options per question 2 to 96 in one pass; more via an automatic tournament
Test accuracy 78.4% micro / 78.9% macro over 31,990 test items, 73 tasks
Calibration ECE 0.029 (per-type, per-option-count temperature scaling)
Latency ~10 ms per call on GPU (fp16), ~48 ms on CPU (int8)
Version 1.0.0 (trained from the pretrained backbone, not fine-tuned from LightDec v1)

What it does

LightDec_V2 answers the kind of small, structured questions that sit inside real systems: Which queue does this ticket go to? Does this email ask for a refund? How angry is the customer, 0–2? Is this agent step wrong? Is this prompt a jailbreak? Does this invoice violate the policy? It is not a generative model and never produces free text; it only chooses among the options you give it, which makes its output easy to validate, log and act on.

Three question types are supported, and several questions about the same state can be asked in one call:

  • choice: pick one of N labelled options (options can carry descriptions, e.g. {"billing": "Charges, invoices, refunds"}).
  • noul: a yes/no judgment about a statement (returns p_true).
  • score: an ordinal scale such as ["Calm", "Annoyed", "Furious"] (also returns an expected_level).

Every answer includes a confidence and a defer flag. With the default defer_threshold of 0.7, the model answered 68.0% of test questions and was 90.6% accurate on those; the remainder are flagged for review.

How it works

The question, the options and the state are packed into one sequence:

[CLS] question [SEP] [MASK] option_1 [MASK] option_2 ... [MASK] option_k [SEP] state [SEP]

The encoder reads the whole sequence once. The hidden vector at each [MASK] marker represents one option; it is combined with the [CLS] context and a question-type embedding, then passed through a 2-layer, 8-head set transformer in which options attend to each other without positional encoding, so the result does not depend on option order. An MLP produces one logit per option and a softmax turns them into a distribution. Finally a temperature, learned per question type and per option-count bucket (2, 3–5, 6–12, 13+ options) and stored in falcondec_config.json, calibrates the probabilities.

Questions with more than 96 options are resolved with a tournament: options are scored in chunks, the best of each chunk advance, and a final pass ranks the survivors.

Usage

The model ships with its own loader, falcondec_modeling.py, which requires torch, transformers, safetensors and huggingface_hub.

import importlib.util
from huggingface_hub import hf_hub_download

path = hf_hub_download("Falconsai/LightDec_V2", "falcondec_modeling.py")
spec = importlib.util.spec_from_file_location("falcondec_modeling", path)
fd = importlib.util.module_from_spec(spec)
spec.loader.exec_module(fd)

model, tok = fd.load_falcondec("Falconsai/LightDec_V2")   # GPU if available, else CPU

state = {
    "email": {
        "subject": "Charged twice this month",
        "body": "I was billed twice for my Pro plan in March. Please refund the duplicate charge. This is the second time!",
    },
    "customer": {"plan": "Pro", "customer_since": "2021"},
}

questions = {
    "topic": {
        "type": "choice",
        "instructions": "What is this support email primarily about?",
        "criteria": {
            "billing": "Charges, invoices, payments, refunds or subscription costs",
            "technical": "Something in the product is broken, failing or erroring",
            "feature_request": "Asking for a new feature or an improvement",
            "account": "Login, password, profile, seats or account settings",
            "none": "None of these",
        },
    },
    "refund": {"type": "noul", "instructions": "The email asks for money to be returned or credited"},
    "anger": {"type": "score", "instructions": "How angry does the customer sound?",
              "criteria": ["Calm", "Annoyed", "Furious"]},
}

out = fd.decide(model, tok, state, questions)
for key, r in out["answers"].items():
    print(key, r["choice"], round(r["confidence"], 3), "DEFER" if r["defer"] else "")

Each result contains choice, choice_text, confidence, the full probs distribution and defer; noul answers add p_true and score answers add expected_level. Pass defer_threshold= to decide to trade coverage for accuracy. For lower-level control, fd.score_items(model, tok, items) scores raw {"state", "question", "options", "type"} items in batches.

CPU / small footprint. If the repository includes the int8 export (as LightDec v1 does in compact-int8/), load that folder the same way; the loader dequantizes the per-channel int8 weights automatically.

Evaluation

All numbers come from falcondec_report.json. The test split has 31,990 items across 73 tasks in 11 domains. Tasks marked held-out are flagged as such in the training report.

Overall

Metric Value
Micro accuracy 78.4%
Macro accuracy (mean over tasks) 78.9%
Negative log-likelihood 0.563
Brier score 0.295
Expected calibration error 0.029
AURC (area under risk–coverage) 0.068
Mean absolute error on score questions (levels) 0.51
Coverage at defer threshold 0.7 68.0%
Accuracy on covered questions 90.6%

By domain

Domain Tasks Test items Accuracy
long_context 3 450 100.0%
support 3 1,505 98.2%
code 9 3,854 91.1%
intents 4 1,800 85.0%
agentic 6 1,563 82.1%
policy 11 5,500 80.6%
guardrails 4 1,294 79.8%
workflows 4 2,000 77.3%
reasoning 18 9,224 71.8%
tev1_benchmark 7 2,800 65.4%
classification 4 2,000 61.2%

Head-to-head against the proof_v2 baseline

On 72 shared tasks (up to 120 items each), LightDec_V2 averaged 79.6% against 46.7% for the proof_v2 baseline, winning on 70 tasks, tying on 1 and losing on 1 (policy/invoice_total_transfer).

Per-task head-to-head
Task n LightDec_V2 proof_v2 Δ
long/email_thread 120 100.0% 5.0% +95.0
agenttrek/next_action_type 120 87.5% 7.5% +80.0
policy/table_extreme_transfer 120 92.5% 14.2% +78.3
long/contract_clause 120 100.0% 23.3% +76.7
policy/access_control_transfer 120 100.0% 23.3% +76.7
long/service_log 120 100.0% 24.2% +75.8
policy/count_threshold_transfer 120 86.7% 10.8% +75.8
bigclonebench/clone 120 96.7% 24.2% +72.5
policy/return_window_transfer 120 100.0% 34.2% +65.8
snli/contradicts 120 100.0% 35.8% +64.2
policy/refund_approval_transfer 120 95.8% 32.5% +63.3
gsm8k/math 120 73.3% 12.5% +60.8
hotpotqa/retrieve 120 87.5% 29.2% +58.3
civil_comments/toxic 120 93.3% 35.8% +57.5
snli/nli 120 90.8% 34.2% +56.7
triage/support_email 120 95.8% 39.2% +56.7
mnli/claim 120 87.5% 31.7% +55.8
scitail/support 120 95.8% 42.5% +53.3
agenttrek/finish_now 120 78.3% 25.8% +52.5
tev1_test/ag_news 120 92.5% 42.5% +50.0
typed_decisions/agent_trace_observability 120 80.8% 33.3% +47.5
hotpotqa/comparison_yes_no 26 92.3% 46.2% +46.2
typed_decisions/customer_service 120 74.2% 28.3% +45.8
policy/table_compare_transfer 120 91.7% 49.2% +42.5
jailbreak/detect 120 98.3% 56.7% +41.7
ag_news/topic 120 85.0% 45.0% +40.0
openbookqa/mcq 120 65.0% 28.3% +36.7
tev1_test/mnli 120 75.0% 39.2% +35.8
counsel/critique_quality 120 60.8% 25.8% +35.0
typed_decisions/invoice_processing 120 81.7% 47.5% +34.2
humaneval/completion 119 84.9% 51.3% +33.6
clinc150/intent 120 97.5% 64.2% +33.3
hellaswag/continuation 120 60.8% 29.2% +31.7
yelp/score 120 62.5% 32.5% +30.0
mbpp/bugspot 120 85.8% 57.5% +28.3
tev1_test/boolq 120 84.2% 55.8% +28.3
commonsense_qa/mcq 120 68.3% 40.8% +27.5
anli/nli 120 59.2% 32.5% +26.7
typed_decisions/security_incidents 120 75.8% 50.0% +25.8
arc_challenge/mcq 120 53.3% 30.8% +22.5
policy/free_shipping_transfer 120 92.5% 70.8% +21.7
agentharm/refuse 120 66.7% 45.0% +21.7
massive_en/intent 120 95.0% 74.2% +20.8
arc_easy/mcq 120 65.0% 45.8% +19.2
tev1_test/banking77 120 71.7% 55.0% +16.7
boolq/yes_no 120 82.5% 66.7% +15.8
policy/table_count_transfer 120 31.7% 15.8% +15.8
mbpp/solution 120 100.0% 85.0% +15.0
policy/invoice_overdue_transfer 120 80.8% 65.8% +15.0
tev1_test/routing 120 44.2% 29.2% +15.0
bitext/category 120 100.0% 85.8% +14.2
qasc/mcq 120 99.2% 85.8% +13.3
winogrande/blank 120 69.2% 55.8% +13.3
emotion/6way 120 49.2% 36.7% +12.5
tev1_test/sst5 120 39.2% 27.5% +11.7
aqua_rat/math 120 35.0% 24.2% +10.8
sst5/score 120 38.3% 28.3% +10.0
devign/vulnerability 120 66.7% 56.7% +10.0
mmlu/mcq 120 40.8% 30.8% +10.0
sciq/mcq 120 96.7% 86.7% +10.0
tev1_test/policy 120 50.0% 40.8% +9.2
codexglue/func_name 120 99.2% 91.7% +7.5
policy/sla_urgency_transfer 120 51.7% 44.2% +7.5
bitext/route 120 100.0% 92.5% +7.5
counsel/step_has_error 120 82.5% 75.0% +7.5
codexglue/code_to_doc 120 100.0% 93.3% +6.7
snli/must_be_true 120 99.2% 92.5% +6.7
banking77/intent 120 92.5% 86.7% +5.8
prompt_injections/detect 116 59.5% 54.3% +5.2
codexglue/doc_to_code 120 98.3% 97.5% +0.8
codexglue/lang_id 120 100.0% 100.0% +0.0
policy/invoice_total_transfer 120 45.0% 52.5% -7.5

Support-email triage (bundled benchmark)

benchmarks/triage_support_email_test.jsonl contains 101 support emails, each with five typed questions and reference answers. It doubles as a worked example of the input format.

Question Type Accuracy
topic choice (5-way) 99.0%
refund noul (yes/no) 100.0%
breakage score (4 levels) 99.0%
anger score (3 levels) 92.1%
judgment choice (3-way) 82.2%

All five answers together route an email to the correct handling pile 94.1% of the time.

Long context

Three long-document tasks exercise the 2,048-token window (150 items each): long/contract_clause, long/email_thread and long/service_log. LightDec_V2 scored 100% on all three. These are synthetic, in-distribution tasks, so treat them as a check that long inputs are read end to end rather than as a measure of general long-document reasoning.

TEV1 transfer benchmark

TEV1 is a separate decision benchmark. Some of its tasks draw on the same public sources as LightDec's training data; the table marks which.

Task n Accuracy Overlaps LightDec training source
tev1_test/ag_news 150 92.0% yes
tev1_test/banking77 200 72.0% no
tev1_test/boolq 200 83.5% yes
tev1_test/mnli 300 72.3% yes
tev1_test/policy 1200 52.9% no
tev1_test/routing 600 45.8% no
tev1_test/sst5 150 39.3% no

Overall TEV1 accuracy is 58.4% (2,800 items), and 51.8% on the tasks with no source overlap. For reference, the report lists published results for the 4B-parameter TEV1 model of 88% on its main decisions set and 100% on its policy transfer set; those are different splits and a model roughly 25× larger, so the figures are context rather than a like-for-like comparison.

Full per-task test results (73 tasks)
Domain Task n Accuracy Chance ECE Held-out
agentic agenttrek/finish_now 151 78.8% 50.0% 0.075
agentic agenttrek/next_action_type 487 86.4% 19.1% 0.092
agentic counsel/critique_quality 201 62.7% 33.3% 0.271
agentic counsel/step_has_error 201 83.1% 50.0% 0.148
agentic hotpotqa/comparison_yes_no 26 92.3% 50.0% 0.082
agentic hotpotqa/retrieve 497 89.3% 16.7% 0.043
classification ag_news/topic 500 90.0% 25.0% 0.052
classification emotion/6way 500 47.2% 16.7% 0.211 ✓
classification sst5/score 500 40.6% 20.0% 0.069 ✓
classification yelp/score 500 66.8% 20.0% 0.091
code bigclonebench/clone 500 96.2% 50.0% 0.031
code codexglue/code_to_doc 504 99.0% 26.2% 0.010
code codexglue/doc_to_code 504 97.6% 27.9% 0.013
code codexglue/func_name 467 95.5% 25.5% 0.039
code codexglue/lang_id 504 100.0% 23.5% 0.001
code devign/vulnerability 500 63.2% 50.0% 0.070
code humaneval/completion 119 84.9% 39.4% 0.140 ✓
code mbpp/bugspot 256 85.5% 40.6% 0.037
code mbpp/solution 500 97.8% 25.0% 0.020
guardrails agentharm/refuse 416 70.0% 50.0% 0.089 ✓
guardrails civil_comments/toxic 500 92.6% 50.0% 0.051
guardrails jailbreak/detect 262 97.3% 50.0% 0.031
guardrails prompt_injections/detect 116 59.5% 50.0% 0.333 ✓
intents banking77/intent 500 92.2% 32.1% 0.037 ✓
intents banking77/intent_77 300 58.0% 1.3% 0.241 ✓
intents clinc150/intent 500 97.0% 12.4% 0.015
intents massive_en/intent 500 93.0% 13.0% 0.028
long_context long/contract_clause 150 100.0% 20.0% 0.000
long_context long/email_thread 150 100.0% 25.0% 0.000
long_context long/service_log 150 100.0% 20.0% 0.001
policy policy/access_control_transfer 500 100.0% 33.3% 0.002
policy policy/count_threshold_transfer 500 88.0% 10.5% 0.048
policy policy/free_shipping_transfer 500 93.8% 50.0% 0.041
policy policy/invoice_overdue_transfer 500 82.4% 50.0% 0.020
policy policy/invoice_total_transfer 500 47.4% 50.0% 0.082
policy policy/refund_approval_transfer 500 96.2% 33.3% 0.019
policy policy/return_window_transfer 500 100.0% 33.3% 0.026
policy policy/sla_urgency_transfer 500 60.6% 25.0% 0.131
policy policy/table_compare_transfer 500 92.8% 50.0% 0.049
policy policy/table_count_transfer 500 32.4% 12.1% 0.541
policy policy/table_extreme_transfer 500 93.2% 12.2% 0.053
reasoning anli/nli 498 48.6% 33.3% 0.146
reasoning aqua_rat/math 247 34.0% 20.0% 0.080
reasoning arc_challenge/mcq 500 49.0% 25.0% 0.168 ✓
reasoning arc_easy/mcq 500 61.2% 25.0% 0.122 ✓
reasoning boolq/yes_no 500 82.2% 50.0% 0.080
reasoning commonsense_qa/mcq 493 64.3% 20.0% 0.143
reasoning gsm8k/math 500 70.0% 25.0% 0.085
reasoning hellaswag/continuation 500 57.6% 25.0% 0.075
reasoning mmlu/mcq 500 39.0% 25.0% 0.155 ✓
reasoning mnli/claim 500 86.0% 33.3% 0.081
reasoning openbookqa/mcq 500 57.2% 25.0% 0.218
reasoning qasc/mcq 500 98.6% 12.5% 0.008
reasoning sciq/mcq 498 95.4% 25.0% 0.021
reasoning scitail/support 500 96.2% 50.0% 0.030
reasoning snli/contradicts 500 99.0% 33.3% 0.019
reasoning snli/must_be_true 500 98.6% 33.3% 0.043
reasoning snli/nli 988 89.5% 33.3% 0.108
reasoning winogrande/blank 500 66.2% 50.0% 0.116
support bitext/category 500 100.0% 16.1% 0.001
support bitext/route 500 100.0% 19.0% 0.000
support triage/support_email 505 94.5% 32.3% 0.046
tev1_benchmark tev1_test/ag_news 150 92.0% 25.0% 0.098 ✓
tev1_benchmark tev1_test/banking77 200 72.0% 17.8% 0.137 ✓
tev1_benchmark tev1_test/boolq 200 83.5% 50.0% 0.050 ✓
tev1_benchmark tev1_test/mnli 300 72.3% 33.3% 0.075 ✓
tev1_benchmark tev1_test/policy 1200 52.9% 33.3% 0.059 ✓
tev1_benchmark tev1_test/routing 600 45.8% 20.0% 0.038 ✓
tev1_benchmark tev1_test/sst5 150 39.3% 20.0% 0.108 ✓
workflows typed_decisions/agent_trace_observability 500 73.4% 30.0% 0.222
workflows typed_decisions/customer_service 500 76.4% 28.0% 0.210
workflows typed_decisions/invoice_processing 500 82.6% 35.0% 0.195
workflows typed_decisions/security_incidents 500 76.8% 34.0% 0.235

Training

Initialization Pretrained jhu-clsp/ettin-encoder-150m (trained from scratch on top of the backbone; no LightDec v1 weights)
Data 1,023,814 training / 24,200 validation / 31,990 test examples
Sources Public classification, NLI, QA, code, intent, guardrail and agent-trace datasets converted into typed decisions, plus synthetic policy-transfer, workflow, support-triage and long-context tasks and TEV1 builders. mind2web was excluded; no custom data was added.
Preset long (2,048-token sequences, 1,500 long-context training tasks)
Objective Cross-entropy with spherical-score (0.5) and ranked-probability-score (1.0) terms for ordinal questions; 8% "none of the above" augmentation; task sampling α = 0.5
Optimizer Learning rate 8e-5 (encoder) / 6e-4 (head), layer-wise decay 0.9, weight decay 0.01, 6% warmup, gradient clip 1.0, EMA 0.999
Batching Token-budget batches of 65,536 tokens (batch size 128)
Schedule 10 epochs; checkpoint from epoch 4 selected on validation macro accuracy
Calibration Temperature per (question type × option-count bucket) fit on validation
Compute 292 minutes on one NVIDIA RTX PRO 6000 Blackwell Server Edition
Software PyTorch 2.9.0 (CUDA 13.0), Transformers 5.17.0, Python 3.12
Epoch Train loss Train acc Val macro acc Val NLL
1 0.706 73.4% 81.9% 0.408
2 0.431 84.7% 84.5% 0.362
3 0.343 87.9% 85.0% 0.392
4 (selected) 0.277 90.4% 85.2% 0.453
5 0.220 92.6% 84.9% 0.537
6 0.170 94.5% 84.7% 0.671
7 0.127 96.0% 84.5% 0.843
8 0.095 97.1% 84.5% 1.034
9 0.071 97.9% 84.4% 1.211
10 0.057 98.4% 84.4% 1.419

Validation accuracy peaked at epoch 4 while validation NLL kept rising after epoch 2, so the selected checkpoint is somewhat overconfident before calibration. The fitted temperatures (roughly 1.6 to 2.5) correct for this; they are part of the checkpoint and applied automatically by decide and score_items.

Speed

Measured with decide on a single state; latency grows only slightly as more questions are asked about the same state.

Device Weights Questions per call p50 p95
cuda fp16 1 10.05 ms 10.43 ms
cuda fp16 5 11.32 ms 11.41 ms
cuda fp16 10 12.57 ms 12.69 ms
cpu fp32 int8 file 1 48.4 ms 48.8 ms

GPU throughput: about 2,589 decisions per second with batching.

Intended use

LightDec_V2 is built for high-volume, low-latency decisions inside software: ticket and email routing, intent detection, triage scoring, content and prompt-safety screening, policy and rule checks over structured records, agent step verification and "should the agent stop now" checks, and pre-filtering before a more expensive LLM call. The calibrated confidence and defer flag are designed for human-in-the-loop and cascade setups in which uncertain cases escalate.

Limitations

  • Closed-set only. The model can only choose among the options you supply. If the right answer is not listed, it will still pick something; include a none option when that can happen.
  • Weak at multi-step reasoning and arithmetic. Knowledge- and math-heavy tasks are well below the rest: MMLU 39.0%, AQuA-RAT 34.0%, ANLI 48.6%, ARC-Challenge 49.0%. Some policy tasks that require computing over a table are also weak: policy/table_count_transfer scores 32.4% with poor calibration (ECE 0.54), and policy/invoice_total_transfer (47.4%) is below chance. Use a larger model for anything that needs counting or totalling.
  • Fine-grained and subjective labels. Accuracy drops on 77-way banking intents (58.0%), six-way emotion (47.2%) and five-level sentiment (SST-5 40.6%, Yelp 66.8%).
  • Transfer to unfamiliar rule sets is limited. On TEV1 tasks without source overlap, accuracy is 51.8%, and routing and policy transfer (45.8% and 52.9%) are the weakest; validate on your own policies before relying on it.
  • Guardrail coverage is partial. Jailbreak detection is strong (97.3%), but prompt-injection detection (59.5%) and harmful-request refusal (70.0%) are not reliable enough to be a sole safety layer.
  • Calibration is aggregate. Overall ECE is low, but a few tasks (for example counsel/critique_quality, prompt_injections/detect, the table-count policy task) are noticeably miscalibrated; check calibration on your own distribution before tuning defer_threshold.
  • English only, and questions are truncated to 96 tokens and each option to 24 tokens by default (option_tokens can raise the per-option budget).
  • Long-context scores are synthetic. The 100% long-context results are in-distribution checks, not evidence of general long-document reasoning.

Files

File Description
model.safetensors fp16 weights
falcondec_config.json Head configuration, special-token ids, defer threshold and calibration temperatures
falcondec_modeling.py Model definition, loader (load_falcondec) and inference API (decide, score_items)
falcondec_report.json Full training and evaluation report
encoder/, tokenizer/ Ettin encoder config and tokenizer
benchmarks/triage_support_email_test.jsonl 101-email support-triage benchmark with reference answers

Citation

@misc{falconsai_lightdec_v2_2026,
  title        = {LightDec_V2: A Long-Context, Calibrated Decision Model},
  author       = {Falcons.ai},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/Falconsai/LightDec_V2}}
}

Built on the Ettin encoder from JHU CLSP (jhu-clsp/ettin-encoder-150m).

This card is generated from the surgical record itself; the package's lineage.intoto.jsonl is the signed source of truth (verify it free at the Surgeon's public verifier or with the bundled verify_attestation.py).

Architecture

  • Identification: NLP · Small Language Model (SLM) (98% confidence)
  • Source format: safetensors · Intended task: not declared
  • config.json: synthesized from the anatomy (no source config.json)
  • Source license: apache-2.0
  • Lineage chain: 1 surgery (no prior attestation reachable) · Falconsai/LightDec_V2
  • Post-surgery totals: 159,654,157 parameters · 168 tensors
  • Compute estimate: 15.02439 GFLOPs (comparison metric, not a measurement)

Provenance & operations

  • Parents: Falconsai/LightDec_V2/model.safetensors
  • Operations performed: load×1
  • Weight merges recorded: 0
  • Quantized tensors (F32→F16): 0

Surgery Log (ordered)

  1. load — hub:Falconsai/LightDec_V2/model.safetensors (319.3 MB, safetensors)

Validation

  • Tissue imaging: not run
  • Structural integrity is testable offline via the packaged load_and_test.py.

Compliance note

The signed attestation + this card together document model composition, modification history, and validation evidence — the record structure technical-documentation obligations (e.g. EU AI Act Annex IV) ask for. This is evidence, not legal advice.


Operated with Model Surgeon — verify this package at https://surgeon.falcons.ai/verify © 2026 FALCONS.AI — Model Surgeon record format. The model weights remain their owner's.

Downloads last month
-
Safetensors
Model size
0.2B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support