LightDec / README.md
RealFalconsAI's picture
Update README.md
6f5f341 verified
|
Raw History Blame Contribute Delete
56.9 kB
---
license: apache-2.0
language:
- en
library_name: transformers
base_model: jhu-clsp/ettin-encoder-150m
pipeline_tag: zero-shot-classification
tags:
- decision-model
- system-one
- falcondec
- lightdec
- calibrated-decisions
- multiple-choice
- intent-classification
- customer-support
- natural-language-inference
- code
- guardrails
- agents
- selective-prediction
- falconsai
- model-surgeon
- attested-lineage
---
> Source model card: `Falconsai/LightDec` @ `main`, carried verbatim below. Its licence is the repository's. The Model Surgeon record follows it.
# Falconsai/LightDec
**[View in Model Surgeon](https://surgeon.falcons.ai/?hub=Falconsai/LightDec)**
**A lightweight, single-pass, typed, calibrated decision model for agentic systems.** Give it a **state** (text, code or JSON), one or more **typed questions** (`choice`, `noul` yes/no, `score` ordinal) and a closed set of options. It returns a calibrated probability for every option, from one encoder pass per question.
LightDec is the FalconDec architecture trained with **FalconDec notebook V2** on the **`standard`** data preset. That run adds agent-specific decisions (AgentTrek, Counsel, HotpotQA) to a balanced mix of 58 test tasks across 9 domains. This page is both the **model card** and the **developer guide**: how to load it, test it, and use it as a decision component in an agentic system.
> **At a glance.** Test accuracy **0.725** (micro and task-macro) on 17,498 decisions from 58 tasks, with ECE **0.025**. At a 0.70 confidence threshold, LightDec answers **56%** of decisions at **89.6%** accuracy and defers the rest. Weights: **319 MB** fp16, **161 MB** int8. It is strongest on support routing, code understanding, intents, guardrails and agent-step checks, and weakest on multi-step arithmetic, date and table reasoning, and very wide label sets. Evaluate it on your own traffic (§6.4) before acting on its answers.
## Contents
1. [Model summary](#1-model-summary)
2. [Results](#2-results)
3. [Intended and out-of-scope uses](#3-intended-and-out-of-scope-uses)
4. [The Hub repository](#4-the-hub-repository)
5. [Install, load and read a result](#5-install-load-and-read-a-result)
6. [Testing the model](#6-testing-the-model)
7. [Using it in an agentic system](#7-using-it-in-an-agentic-system)
8. [Tuning for your domain](#8-tuning-for-your-domain)
9. [Architecture](#9-architecture)
10. [Training data](#10-training-data)
11. [Training procedure and calibration](#11-training-procedure-and-calibration)
12. [Operational notes](#12-operational-notes)
13. [Bias, risks and limitations](#13-bias-risks-and-limitations)
14. [Versioning and lineage](#14-versioning-and-lineage)
15. [API reference](#15-api-reference)
16. [Citation and references](#16-citation-and-references)
---
## 1. Model summary
| | |
|---|---|
| **Model** | LightDec: the FalconDec architecture, notebook V2, `standard` preset. The checkpoint's own config reports `FalconDec` version `1.0.0` |
| **Task** | Closed-set decisions: given a state, a question and 2–N options, return a calibrated probability per option |
| **Question types** | `choice` (pick one), `noul` (yes/no, returns P(true)), `score` (ordinal rubric, returns the expected level) |
| **Backbone** | [`jhu-clsp/ettin-encoder-150m`](https://huggingface.co/jhu-clsp/ettin-encoder-150m) (ModernBERT-style encoder), fully fine-tuned |
| **Decision head** | Option-marker scoring plus a permutation-equivariant set-transformer head (§9) |
| **Parameters** | ≈160M (Hub reports 0.2B) |
| **Weights** | fp16 **319 MB** (`model.safetensors`) · per-channel int8 **161 MB** (`compact-int8/model_int8.safetensors`) |
| **Context** | 512 tokens; automatically 2,048 for questions with more than 24 options; a tournament above 96 options |
| **Calibration** | One temperature per (question type × option-count bucket), stored in the checkpoint and applied automatically |
| **Inference cost** | One encoder pass per question. The same architecture (`Falconsai/proof_v3`) measured 15.7 ms p50 for one question on a GPU; LightDec's own latency is not in its report (measure with §6.3) |
| **Output** | Probabilities, the chosen option, confidence, a `defer` flag, `p_true` (noul) and `expected_level` (score) |
| **Language** | English, plus code in Python, Java, JavaScript, PHP, Ruby, Go and C |
| **Custom code** | `falcondec_modeling.py` ships with the weights and holds the model and all inference logic. Load it with `importlib` (§5); `AutoModel.from_pretrained` alone won't build the decision head |
| **License** | Apache-2.0 for the weights and code. Check each training dataset's license before redistributing derived data |
The model has no generative component. It can only rank the options you give it, so it cannot produce text outside that set.
---
## 2. Results
All numbers come from this checkpoint's `falcondec_report.json`: one run, `standard` preset, 2 epochs, seed 42, `MODE="scratch"`, notebook V2.
### 2.1 At a glance
| Metric | Value |
|---|---|
| Test decisions / tasks | 17,498 / 58 (10 of them held out) |
| Test accuracy, micro | **0.725** |
| Test accuracy, task-macro | **0.725** |
| Held-out tasks, task-macro (10 tasks never trained on) | **0.567** |
| Expected calibration error (ECE, 15 bins) | **0.025** |
| Negative log-likelihood / Brier score | 0.652 / 0.358 |
| Area under the risk–coverage curve (AURC, lower is better) | **0.097** |
| Ordinal (`score`) mean absolute error, in levels | 0.572 |
| Coverage / accuracy at confidence ≥ 0.70 | **56.4% / 0.896** |
| Weights | fp16 319 MB · int8 161 MB |
**Selective prediction is the headline.** Calibration is good (ECE 0.025), so the confidence score is a reliable gate. Acting only on decisions with confidence ≥ 0.70 covers 56% of traffic at 89.6% accuracy, against 72.5% accuracy when answering everything. That is the property an agent loop needs: answer the easy majority locally, and hand the rest to an LLM or a human.
### 2.2 Per domain
| Domain | Test decisions | Task-macro accuracy |
|---|---|---|
| support | 600 | **0.997** |
| code | 2,350 | **0.864** |
| intents | 1,200 | **0.773** |
| guardrails | 1,094 | **0.768** |
| agentic | 717 | **0.754** |
| workflows | 2,000 | **0.689** |
| policy | 2,700 | **0.669** |
| reasoning | 5,637 | **0.667** |
| classification | 1,200 | **0.584** |
### 2.3 Per task
*Held-out tasks were never used for training, calibration or model selection. "proof_v2 (card)" lists proof_v2's published score for the same source and task (different samples; indicative only).*
| Domain | Task | n | Chance | **Accuracy** | ECE | proof_v2 (card) |
|---|---|---|---|---|---|---|
| support | `bitext/route` | 300 | 0.200 | **1.000** | 0.002 | 0.958 |
| support | `bitext/category` | 300 | 0.172 | **0.993** | 0.008 | |
| code | `codexglue/lang_id` | 300 | 0.235 | **1.000** | 0.004 | 0.997 |
| code | `mbpp/solution` | 300 | 0.250 | **0.987** | 0.014 | 0.992 |
| code | `codexglue/code_to_doc` | 300 | 0.256 | **0.977** | 0.016 | 0.969 |
| code | `codexglue/doc_to_code` | 300 | 0.274 | **0.973** | 0.022 | 0.961 |
| code | `codexglue/func_name` | 275 | 0.263 | **0.938** | 0.021 | 0.901 |
| code | `bigclonebench/clone` | 300 | 0.500 | **0.863** | 0.098 | 0.383 |
| code | `humaneval/completion` *(held out)* | 119 | 0.394 | **0.756** | 0.152 | 0.575 |
| code | `mbpp/bugspot` | 156 | 0.413 | **0.731** | 0.086 | 0.475 |
| code | `devign/vulnerability` | 300 | 0.500 | **0.553** | 0.026 | 0.542 |
| intents | `banking77/intent` *(held out)* | 300 | 0.317 | **0.923** | 0.037 | 0.883 |
| intents | `massive_en/intent` | 300 | 0.122 | **0.907** | 0.046 | |
| intents | `clinc150/intent` | 300 | 0.122 | **0.793** | 0.081 | 0.850 |
| intents | `banking77/intent_77` *(held out)* | 300 | 0.013 | **0.470** | 0.200 | |
| guardrails | `jailbreak/detect` | 262 | 0.500 | **0.966** | 0.018 | |
| guardrails | `civil_comments/toxic` | 300 | 0.500 | **0.807** | 0.051 | |
| guardrails | `agentharm/refuse` *(held out)* | 416 | 0.500 | **0.654** | 0.178 | |
| guardrails | `prompt_injections/detect` *(held out)* | 116 | 0.500 | **0.647** | 0.272 | |
| agentic | `hotpotqa/retrieve` | 298 | 0.168 | **0.842** | 0.078 | |
| agentic | `hotpotqa/comparison_yes_no` | 17 | 0.500 | **0.824** | 0.185 | |
| agentic | `counsel/step_has_error` | 201 | 0.500 | **0.791** | 0.093 | |
| agentic | `counsel/critique_quality` | 201 | 0.333 | **0.557** | 0.094 | |
| workflows | `typed_decisions/customer_service` | 500 | 0.280 | **0.720** | 0.099 | |
| workflows | `typed_decisions/security_incidents` | 500 | 0.340 | **0.712** | 0.140 | |
| workflows | `typed_decisions/agent_trace_observability` | 500 | 0.300 | **0.696** | 0.106 | |
| workflows | `typed_decisions/invoice_processing` | 500 | 0.350 | **0.628** | 0.101 | |
| policy | `policy/access_control_transfer` | 300 | 0.333 | **1.000** | 0.001 | |
| policy | `policy/return_window_transfer` | 300 | 0.333 | **1.000** | 0.019 | |
| policy | `policy/free_shipping_transfer` | 300 | 0.500 | **0.850** | 0.067 | |
| policy | `policy/sla_urgency_transfer` | 300 | 0.250 | **0.713** | 0.214 | |
| policy | `policy/refund_approval_transfer` | 300 | 0.333 | **0.710** | 0.219 | |
| policy | `policy/count_threshold_transfer` | 300 | 0.179 | **0.523** | 0.085 | |
| policy | `policy/invoice_total_transfer` | 300 | 0.500 | **0.523** | 0.020 | |
| policy | `policy/invoice_overdue_transfer` | 300 | 0.500 | **0.407** | 0.364 | |
| policy | `policy/table_extreme_transfer` | 300 | 0.240 | **0.290** | 0.036 | |
| reasoning | `qasc/mcq` | 300 | 0.125 | **0.983** | 0.007 | |
| reasoning | `snli/must_be_true` | 300 | 0.333 | **0.970** | 0.029 | 0.908 |
| reasoning | `snli/contradicts` | 300 | 0.333 | **0.967** | 0.035 | 0.892 |
| reasoning | `scitail/support` | 300 | 0.500 | **0.957** | 0.030 | |
| reasoning | `sciq/mcq` | 300 | 0.250 | **0.950** | 0.023 | 0.692 |
| reasoning | `snli/nli` | 594 | 0.333 | **0.837** | 0.035 | |
| reasoning | `mnli/claim` | 300 | 0.333 | **0.807** | 0.072 | 0.492 |
| reasoning | `boolq/yes_no` | 300 | 0.500 | **0.783** | 0.071 | 0.717 |
| reasoning | `gsm8k/math` | 300 | 0.250 | **0.637** | 0.050 | 0.275 |
| reasoning | `commonsense_qa/mcq` | 296 | 0.200 | **0.611** | 0.058 | 0.442 |
| reasoning | `arc_easy/mcq` *(held out)* | 300 | 0.250 | **0.553** | 0.057 | 0.425 |
| reasoning | `openbookqa/mcq` | 300 | 0.250 | **0.550** | 0.068 | 0.292 |
| reasoning | `winogrande/blank` | 300 | 0.500 | **0.540** | 0.089 | |
| reasoning | `arc_challenge/mcq` *(held out)* | 300 | 0.250 | **0.423** | 0.085 | 0.308 |
| reasoning | `anli/nli` | 300 | 0.333 | **0.393** | 0.185 | |
| reasoning | `mmlu/mcq` *(held out)* | 300 | 0.250 | **0.393** | 0.094 | |
| reasoning | `hellaswag/continuation` | 300 | 0.250 | **0.383** | 0.171 | |
| reasoning | `aqua_rat/math` | 247 | 0.200 | **0.259** | 0.071 | |
| classification | `ag_news/topic` | 300 | 0.250 | **0.847** | 0.051 | |
| classification | `yelp/score` | 300 | 0.200 | **0.640** | 0.069 | |
| classification | `emotion/6way` *(held out)* | 300 | 0.167 | **0.480** | 0.049 | |
| classification | `sst5/score` *(held out)* | 300 | 0.200 | **0.370** | 0.086 | |
### 2.4 Comparison with proof_v2 (indicative)
On the 22 tasks that both this report and proof_v2's model card cover, LightDec's task-macro accuracy is **0.807 vs 0.679**, and it scores higher on **20 of 22**. The largest gains are on the tasks proof_v2 reported as weak:
| Task | proof_v2 (card) | LightDec |
|---|---|---|
| BigCloneBench clone detection | 0.383 | **0.863** |
| GSM8K (4-option numeric) | 0.275 | **0.637** |
| MultiNLI claim | 0.492 | **0.807** |
| OpenBookQA | 0.292 | **0.550** |
| MBPP bug spotting | 0.475 | **0.731** |
| HumanEval completion *(held out)* | 0.575 | **0.756** |
| ARC-Challenge *(held out)* | 0.308 | **0.423** |
LightDec is lower on CLINC150 (0.793 vs 0.850) and MBPP task→solution (0.987 vs 0.992).
These are **different test samples and, for some tasks, different question formats**. For example, LightDec's MultiNLI task is three-way NLI, and its bug-spotting mutants are verified to fail the unit tests. The like-for-like head-to-head, which runs proof_v2 on identical decisions (notebook cell 23), **did not run** for this checkpoint (`head_to_head: null`).
### 2.5 Comparison with Laya and TypeSafe Jev (indicative)
| Benchmark | LightDec | Laya | TypeSafe Jev 1.13.0 |
|---|---|---|---|
| typed-decisions test (2,000 decisions, 4 workflows) | 0.689 | 0.766 (fine-tuned) · 0.362 (zero-shot) | 0.727 |
| AG News (4 labels) | 0.847 | 0.950 | 0.910 |
| DAIR Emotion, 6 labels *(held out for LightDec)* | 0.480 | 0.595 | 0.480 |
| Banking77, all 77 labels in one question *(held out)* | 0.470 | 0.425 | 0.870 (72 labels) |
| SST-5 (ordinal) *(held out)* | 0.370 | 0.372 | — |
Laya's numbers are from its own benchmark report; Jev's are third-party published. LightDec was trained on the typed-decisions training split, like the fine-tuned Laya checkpoint. **On typed-decisions LightDec trails both** (the teacher-agreement ceiling is 0.735 and the majority-class baseline 0.461). It matches Jev on Emotion, edges Laya on all-77 Banking77, and trails both on AG News. LightDec is 2.6× smaller than Laya's 421M English checkpoint.
---
## 3. Intended and out-of-scope uses
### Intended
| Use | Measured evidence |
|---|---|
| **Support and ticket routing** | Bitext route 1.000, Bitext category 0.993; Banking77 (held out, 2–5 options) 0.923; MASSIVE 0.907 |
| **Code understanding against a menu** | Language ID 1.000; code↔description 0.973–0.977; task→solution 0.987; function naming 0.938 |
| **Guardrails** | Jailbreak detection 0.966; toxicity 0.807. Held out: prompt-injection 0.647, AgentHarm refusal 0.654, so recalibrate and test on your own traffic |
| **Agent loops** | Retrieval routing (HotpotQA) 0.842; "does this agent step contain an error?" (Counsel) 0.791 |
| **Statement verification** | SNLI must-be-true / contradicts 0.970 / 0.967; SciTail 0.957; SciQ 0.950 |
| **Selective automation** | 89.6% accuracy on the 56% of decisions with confidence ≥ 0.70 |
### Out of scope
- **Multi-step arithmetic and quantitative reasoning.** AQuA 0.259 (chance 0.20); counting and summing policies 0.52; comparing values in a table 0.290. GSM8K reaches 0.637 only because it is posed as 4-option multiple choice with near-miss distractors. Route real math to an LLM or code.
- **Date reasoning in unfamiliar formats.** The invoice-overdue transfer test (ISO dates, whereas training used "Month DD, YYYY") scores **0.407, below chance, with ECE 0.364**. It is confidently wrong there. Normalise dates before asking, or compute them in code.
- **Very wide label sets.** All 77 Banking77 intents in one question score 0.470. Pre-filter to a shortlist of about 20 options (§7.1).
- **Hard commonsense and exam knowledge**: HellaSwag 0.383, ANLI 0.393, MMLU 0.393.
- **Code security and correctness gating**: Devign 0.553 is near chance. Don't use it to approve code.
- **Open-ended questions**: it always picks one of your options. Add "None of the above" when appropriate; it was trained with that option.
- **Non-English text**, and **high-stakes decisions without human oversight**.
---
## 4. The Hub repository
| File | Contents |
|---|---|
| `model.safetensors` | fp16 weights (319 MB): encoder, decision head and the temperature buffer |
| `falcondec_config.json` | Layout (sequence lengths, option budgets), special-token ids, temperatures, defer threshold, version, lineage |
| `encoder/` | Backbone configuration (the encoder is rebuilt from this, then the weights are loaded) |
| `tokenizer/` | Tokenizer files |
| `falcondec_modeling.py` | `FalconDec`, `load_falcondec`, `decide`, `score_items`, `save_falcondec` and the int8 codec |
| `falcondec_report.json` | Training configuration, data counts, history, temperatures and all test results |
| `compact-int8/` | The same model with per-channel int8 weights (`model_int8.safetensors`, 161 MB); a complete, self-contained directory with its own config, tokenizer and modeling file |
| `README.md` | This card |
**Pin a revision in production.** The repo can change, so pass a commit hash when you load.
---
## 5. Install, load and read a result
```bash
pip install torch "transformers>=4.48" safetensors huggingface_hub numpy
```
### 5.1 Load
Save this helper as `lightdec.py` next to your code. Every example below uses it.
```python
# lightdec.py
import importlib.util, json, shutil
from pathlib import Path
from huggingface_hub import snapshot_download
def load_lightdec(repo="Falconsai/LightDec", revision=None, variant="fp16", device=None, dtype=None):
"""Returns (fdm, model, tokenizer). variant: "fp16" (319 MB) or "int8" (161 MB, dequantised on load)."""
path = Path(repo) if Path(repo).exists() else Path(snapshot_download(repo, revision=revision))
if variant == "int8":
path = path / "compact-int8"
fc = json.loads((path / "falcondec_config.json").read_text(encoding="utf-8"))
expected = fc.get("weights", "model.safetensors")
if not (path / expected).exists(): # e.g. a renamed weight file in a processed copy
cands = sorted(path.glob("*.safetensors"))
if not cands:
raise FileNotFoundError(f"no .safetensors weights in {path}")
local = Path("lightdec_local") / variant
shutil.copytree(path, local, dirs_exist_ok=True)
shutil.copy(cands[0], local / expected)
path = local
spec = importlib.util.spec_from_file_location("falcondec_modeling", str(path / "falcondec_modeling.py"))
fdm = importlib.util.module_from_spec(spec)
spec.loader.exec_module(fdm)
model, tok = fdm.load_falcondec(str(path), device=device, dtype=dtype) # cuda if available, else cpu
return fdm, model, tok
```
```python
from lightdec import load_lightdec
fdm, model, tok = load_lightdec() # or load_lightdec(revision="<commit>", variant="int8")
print(model.fcfg["name"], model.fcfg["version"], round(model.num_parameters() / 1e6, 1), "M params")
```
If loading prints `[FalconDec] load warning: missing=… unexpected=…`, the weights didn't match the architecture. Treat that as a failed load (§6.1 checks for it).
### 5.2 Decide
`decide()` takes one state and any number of typed questions: either a list, or a Jev/Laya-style dict keyed by name.
```python
state = {"from": "user@acme.com", "subject": "Duplicate charge on invoice #4411",
"body": "We were billed twice for March. Please refund the duplicate today or we will cancel our plan."}
out = fdm.decide(model, tok, state, {
"department": {"type": "choice", "instructions": "Which department should handle this request?",
"criteria": {"billing": "invoices, payments, refunds", "technical": "bugs, outages",
"sales": "pricing, contracts", "other": "everything else"}},
"urgency": {"type": "score", "instructions": "How urgent is this request?",
"criteria": ["not urgent", "soon", "critical deadline or blocking issue"]},
"churn_risk": {"type": "noul", "instructions": "Does the user threaten to cancel or leave?"},
})
a = out["answers"]
print(a["department"]["choice"], round(a["department"]["confidence"], 3), a["department"]["defer"])
print("urgency level", round(a["urgency"]["expected_level"], 2), "of", 2)
print("P(churn)", round(a["churn_risk"]["p_true"], 3))
```
Plain options work too: `{"question": "Which team?", "options": ["Accounts", "Billing", "Shipping"]}`.
### 5.3 Reading a result
Each item in `out["results"]` (and `out["answers"][key]`) contains:
| Field | Meaning |
|---|---|
| `key` | The question's name (dict input) or `None` |
| `type` | `choice`, `noul` or `score` |
| `choice` | The chosen key: the criteria key, the option text, `True`/`False` for `noul`, or the level index for `score` |
| `choice_text` | The option text the model saw |
| `confidence` | Calibrated probability of `choice` |
| `probs` | The full distribution, keyed by `str(key)` |
| `defer` | `True` when `confidence` is below the defer threshold (default 0.70, stored in the config): don't act on it (§7.3) |
| `p_true` | `noul` only: calibrated P(yes) |
| `expected_level` | `score` only: probability-weighted level (0 … k−1); better than the argmax for ordinal rubrics (test MAE 0.57 levels) |
How to read them:
- **Low confidence, spread probabilities**: the state doesn't support any option clearly. Defer, or add "None of the above".
- **`noul` near 0.5**: genuinely ambiguous. Ask for more information rather than guessing.
- **`score`**: use `expected_level` for thresholds ("escalate if ≥ 1.5") rather than `choice`.
---
## 6. Testing the model
Tests 6.1–6.3 need no labelled data, so run them in CI whenever you change the revision. Test 6.4 is the one that tells you whether to ship.
### 6.1 Integrity and determinism
Save as `check_lightdec.py` and run `python check_lightdec.py [revision]`.
```python
import contextlib, io, sys
import numpy as np
from lightdec import load_lightdec
rev = sys.argv[1] if len(sys.argv) > 1 else None
log = io.StringIO()
with contextlib.redirect_stdout(log):
fdm, model, tok = load_lightdec(revision=rev)
assert "load warning" not in log.getvalue(), log.getvalue()
fc = model.fcfg
assert fc["name"] == "FalconDec", fc["name"] # LightDec checkpoints use the FalconDec architecture name
T = model.temperature.float().cpu().numpy()
assert T.shape == (3, 4) and (T > 0).all(), T
q = {"team": {"question": "Which team?", "options": ["recover password", "shipping", "invoicing"]}}
a = fdm.decide(model, tok, "I forgot my password and can't sign in.", q)["answers"]["team"]
b = fdm.decide(model, tok, "I forgot my password and can't sign in.", q)["answers"]["team"]
assert all(abs(a["probs"][k] - b["probs"][k]) < 1e-4 for k in a["probs"]), "non-deterministic"
assert abs(sum(a["probs"].values()) - 1) < 1e-3
print(f"OK LightDec (FalconDec v{fc['version']}, notebook {fc.get('notebook_version')}) choice={a['choice']} "
f"conf={a['confidence']:.3f} defer_threshold={fc.get('defer_threshold')}")
```
The stored temperatures should read approximately `[[1.707, 1.352, 1.466, 1.349], [1.402 ×4], [1.453 ×4]]` (§11.2).
### 6.2 Behavioural tests (pytest)
Save as `tests/test_lightdec.py` and run `pytest -q`. Set `LIGHTDEC_REVISION` to test a pinned commit.
```python
import os, random
import pytest
from lightdec import load_lightdec
@pytest.fixture(scope="session")
def fd():
return load_lightdec(revision=os.environ.get("LIGHTDEC_REVISION"))
def ask(fd, state, question, options, **kw):
fdm, model, tok = fd
return fdm.decide(model, tok, state, [dict(question=question, options=options, **kw)])["results"][0]
def test_probabilities_are_valid(fd):
r = ask(fd, "The build failed on main.", "What next?", ["Revert", "Ignore", "Retry"])
assert all(0 <= p <= 1 for p in r["probs"].values()) and abs(sum(r["probs"].values()) - 1) < 1e-3
def test_typed_outputs(fd):
fdm, model, tok = fd
out = fdm.decide(model, tok, "I was charged twice. Refund me or I'm leaving.", {
"refund": {"type": "noul", "instructions": "Does the user ask for a refund?"},
"urgency": {"type": "score", "instructions": "How urgent?", "criteria": ["low", "medium", "high"]}})["answers"]
assert 0 <= out["refund"]["p_true"] <= 1 and out["refund"]["choice"] in (True, False)
assert 0 <= out["urgency"]["expected_level"] <= 2
def test_support_routing(fd):
r = ask(fd, "I forgot my password and the reset email never arrived.", "Which team should handle this?",
["recover password", "billing and payment", "delivery information"])
assert r["choice"] == "recover password"
def test_fanout_matches_single_questions(fd):
# Batching changes padding; under bf16 that moves probabilities slightly, never the substance.
fdm, model, tok = fd
state = "I forgot my password and can't sign in."
qs = [{"question": "Team?", "options": ["recover password", "shipping", "invoicing"]},
{"question": "Urgent?", "options": ["Yes", "No"]}]
together = fdm.decide(model, tok, state, qs)["results"]
for q, t in zip(qs, together):
alone = fdm.decide(model, tok, state, [q])["results"][0]
assert all(abs(alone["probs"][k] - t["probs"][k]) < 2e-2 for k in alone["probs"])
def test_option_order_is_mostly_irrelevant(fd):
# The head is order-equivariant, but the encoder sees positions; training reshuffled options every epoch.
state, q = "Where is my parcel? It's three days late.", "What should support do?"
opts = ["Give the delivery status", "Start a refund", "Book an appointment"]
base = ask(fd, state, q, opts)["choice"]
same = sum(ask(fd, state, q, random.Random(s).sample(opts, len(opts)))["choice"] == base for s in range(5))
assert same >= 4
def test_many_options_use_the_tournament(fd):
opts = [f"topic number {i}" for i in range(119)] + ["reset my password"]
r = ask(fd, "I can't log in, I need to reset my password.", "What does the user want?", opts)
assert len(r["probs"]) == 120 and abs(sum(r["probs"].values()) - 1) < 1e-3
def test_int8_agrees_with_fp16(fd):
fdm8, m8, tok8 = load_lightdec(revision=os.environ.get("LIGHTDEC_REVISION"), variant="int8")
fdm, model, tok = fd
items = [dict(state=s, question="Which team?", options=["billing", "shipping", "accounts", "technical"])
for s in ["I was double charged", "Where is my parcel?", "Change my email", "The app crashes on start",
"Refund the duplicate payment", "Package never arrived", "Reset my login", "Error 500 on checkout"]]
a = [p.argmax() for p in fdm.score_items(model, tok, items)]
b = [p.argmax() for p in fdm8.score_items(m8, tok8, items)]
assert sum(x == y for x, y in zip(a, b)) >= len(items) - 1
```
### 6.3 Latency
```python
import time, numpy as np, torch
from lightdec import load_lightdec
fdm, model, tok = load_lightdec()
q = [{"question": "Route?", "options": ["billing and payment", "shipping", "recover password"]}]
for _ in range(5):
fdm.decide(model, tok, "I was charged twice.", q)
t = []
for _ in range(100):
if torch.cuda.is_available(): torch.cuda.synchronize()
t0 = time.perf_counter(); fdm.decide(model, tok, "I was charged twice.", q)
if torch.cuda.is_available(): torch.cuda.synchronize()
t.append((time.perf_counter() - t0) * 1000)
print(f"p50 {np.percentile(t, 50):.1f} ms p95 {np.percentile(t, 95):.1f} ms on {model.device}")
```
For CPU serving, load the int8 variant with `device="cpu", dtype=torch.float32`, and optionally apply `torch.ao.quantization.quantize_dynamic(model, {torch.nn.Linear}, dtype=torch.qint8)` for int8 matrix multiplies.
### 6.4 Accuracy on your own labelled data
Write 50–500 decisions that look like your real traffic, one JSON object per line. `expected` may be a letter, a 0-based index or the option text; `type` is optional.
```json
{"id": "t1", "tag": "support", "state": "…", "question": "…", "options": ["…", "…"], "expected": "B", "type": "choice"}
```
```python
import json, string
import numpy as np
from lightdec import load_lightdec
fdm, model, tok = load_lightdec()
rows = [json.loads(l) for l in open("my_eval.jsonl", encoding="utf-8") if l.strip()]
def idx(v, opts):
if isinstance(v, int): return v
v = str(v).strip()
if len(v) == 1 and v.upper() in string.ascii_uppercase[:len(opts)]: return string.ascii_uppercase.index(v.upper())
return [o.lower() for o in opts].index(v.lower())
items = [dict(state=r["state"], question=r.get("question", ""), options=r["options"], type=r.get("type", "choice"))
for r in rows]
probs = fdm.score_items(model, tok, items, batch_size=64)
gold = np.array([idx(r["expected"], r["options"]) for r in rows])
pred = np.array([p.argmax() for p in probs]); conf = np.array([p.max() for p in probs]); ok = pred == gold
def ece(c, k, bins=15):
e = 0.0
for lo in np.linspace(0, 1, bins, endpoint=False):
m = (c > lo) & (c <= lo + 1 / bins)
if m.any(): e += m.mean() * abs(c[m].mean() - k[m].mean())
return e
print(f"accuracy {ok.mean():.3f} | ECE {ece(conf, ok):.3f}")
for tag in sorted({r.get("tag", "all") for r in rows}):
m = np.array([r.get("tag", "all") == tag for r in rows]); print(f" {tag:12s} n={m.sum():4d} acc={ok[m].mean():.3f}")
for t in (0.5, 0.6, 0.7, 0.8, 0.9):
m = conf >= t
print(f" act if conf >= {t}: answers {m.mean():6.1%}, accuracy when answering {ok[m].mean() if m.any() else float('nan'):.3f}")
```
The last loop is the deferral policy of §7.3; on the published test mix, 0.70 gives 56% coverage at 0.896 accuracy. Pick the smallest threshold whose "accuracy when answering" meets your bar. If ECE on your data is much higher than 0.025, recalibrate (§8.2).
### 6.5 Regression gate between revisions
Fail the pipeline if any test task drops by more than two points between two revisions:
```python
import json, sys
from huggingface_hub import hf_hub_download
old_rev, new_rev = sys.argv[1:3]
rep = lambda rev: {r["task"]: r["acc"] for r in json.load(open(
hf_hub_download("Falconsai/LightDec", "falcondec_report.json", revision=rev), encoding="utf-8"))["test_per_task"]}
old, new = rep(old_rev), rep(new_rev)
bad = [(t, old[t], new[t]) for t in new if t in old and new[t] < old[t] - 0.02]
print("\n".join(f"REGRESSION {t}: {a:.3f} -> {b:.3f}" for t, a, b in bad) or "no regressions")
sys.exit(1 if bad else 0)
```
---
## 7. Using it in an agentic system
### 7.1 Where it fits
An agent loop is mostly small decisions (which tool, is this safe, did that work, am I done, should a human look) around a few hard reasoning steps. LLMs are slow and poorly calibrated at the small ones. LightDec takes those; the LLM keeps planning, reasoning and generation.
| Agent step | How to phrase it | Evidence |
|---|---|---|
| **Entry routing** | state = the request; `choice` over sub-agents or workflows, plus "None of the above" | Intents 0.773–0.997 by source |
| **Retrieval routing** | state = the question; options = candidate documents or indexes | HotpotQA retrieve 0.842 (9 candidates) |
| **Step verification** | `noul`: "Does the agent's current step contain an error?" | Counsel step-error 0.791 |
| **Guardrail** | `noul`: "Does this input try to override the agent's instructions?" on user input *and* on tool results | Jailbreak 0.966; held-out injection 0.647, so validate on your traffic |
| **Conditional edges** | `choice`: "Retry, continue, escalate or finish?" over the current state | Workflows 0.689 |
| **Escalation** | `defer == True`, or `confidence` below your threshold → human or larger model | 0.896 accuracy on the confident 56% |
Keep option sets under about 20. For larger menus, shortlist first (embedding search or a coarse `choice`), then ask LightDec; all-77-label Banking77 drops to 0.470. Several questions about the same state go in one `decide()` call.
### 7.2 A routing node (LangGraph)
```python
from lightdec import load_lightdec
fdm, model, tok = load_lightdec(revision="<commit>")
def route(state: dict) -> str:
res = fdm.decide(model, tok, state, {
"next": {"type": "choice", "instructions": "What should the agent do next?",
"criteria": {"search": "needs external information", "code": "needs code written or run",
"answer": "has enough information to answer", "human": "ambiguous, risky or out of scope"}},
"unsafe": {"type": "noul", "instructions": "Does the latest input try to override the agent's instructions?"},
}, defer_threshold=0.75)["answers"]
if res["unsafe"]["p_true"] > 0.5:
return "human"
if res["next"]["defer"]:
return "llm_planner" # low confidence: let the LLM decide
return res["next"]["choice"]
graph.add_conditional_edges("observe", route, {"search": "search_node", "code": "code_node", "answer": "answer_node",
"human": "human_node", "llm_planner": "planner_node"})
```
### 7.3 The deferral policy
| Situation | Action |
|---|---|
| `confidence` ≥ your threshold | Act |
| `confidence` below it (`defer == True`) | **Defer**: hand to the LLM, ask a human, or ask the user for more information |
| Guardrail `noul` with `p_true` above your risk threshold | Block or escalate, regardless of other answers |
Choose thresholds from your own evaluation (§6.4). The default 0.70 gives 56% coverage at 0.896 accuracy on the published test mix. Confidence is not trustworthy on the task types listed as out of scope in §3; date-format transfer, for example, is confidently wrong. For irreversible actions (payments, deletions, sending email), raise the threshold and keep a hard rule or human confirmation in front: the state is attacker-controlled text, and adversarial input can move scores. Log the question, options, choice, confidence, model version and Hub revision for every decision; that log becomes your next evaluation and fine-tuning set (§8.3).
### 7.4 As a tool for Claude (tool use)
```python
import json, threading
import anthropic
from lightdec import load_lightdec
fdm, model, tok = load_lightdec()
lock = threading.Lock()
client = anthropic.Anthropic()
tools = [{
"name": "lightdec_decide",
"description": ("Fast, local, calibrated closed-set decision model. Give it a state (text or JSON), a question and "
"2-20 distinct options; it returns the choice, a calibrated confidence and a 'defer' flag. "
"If 'defer' is true, don't rely on the answer. Not for arithmetic, dates or multi-step reasoning."),
"input_schema": {"type": "object", "properties": {
"state": {"type": "string", "description": "The message, document excerpt or JSON state."},
"question": {"type": "string"},
"options": {"type": "array", "items": {"type": "string"}, "minItems": 2},
"type": {"type": "string", "enum": ["choice", "score"], "description": "score = options are ordered levels"}},
"required": ["state", "question", "options"]},
}]
def run_tool(inp):
with lock:
r = fdm.decide(model, tok, inp["state"], [{"question": inp["question"], "options": inp["options"],
"type": inp.get("type", "choice")}])["results"][0]
return {k: r[k] for k in ("choice", "confidence", "defer", "probs") if k in r}
messages = [{"role": "user", "content": "Triage: 'I forgot my password and the reset email never arrived.' "
"Teams: Accounts, Billing, Shipping."}]
while True:
resp = client.messages.create(model="claude-sonnet-5", max_tokens=1024, tools=tools, messages=messages)
if resp.stop_reason != "tool_use":
print("".join(b.text for b in resp.content if b.type == "text"))
break
messages.append({"role": "assistant", "content": resp.content})
results = []
for block in resp.content:
if block.type == "tool_use" and block.name == "lightdec_decide":
try:
results.append({"type": "tool_result", "tool_use_id": block.id, "content": json.dumps(run_tool(block.input))})
except Exception as exc:
results.append({"type": "tool_result", "tool_use_id": block.id, "content": str(exc), "is_error": True})
messages.append({"role": "user", "content": results})
```
### 7.5 As an MCP server
`pip install mcp`, then save `lightdec_mcp.py` next to `lightdec.py`:
```python
import os, threading
from mcp.server.fastmcp import FastMCP
from lightdec import load_lightdec
fdm, model, tok = load_lightdec(revision=os.environ.get("LIGHTDEC_REVISION"),
variant=os.environ.get("LIGHTDEC_VARIANT", "fp16"))
MIN_CONF = float(os.environ.get("LIGHTDEC_MIN_CONF", "0.7"))
lock = threading.Lock()
mcp = FastMCP("lightdec")
@mcp.tool()
def decide(state: str, question: str, options: list[str], type: str = "choice") -> dict:
"""Choose one of 2-20 distinct options for a question about a state. type="score" means ordered levels.
Returns the choice, a calibrated confidence and 'defer' (true = not reliable enough to act on)."""
with lock:
r = fdm.decide(model, tok, state, [{"question": question, "options": options, "type": type}],
defer_threshold=MIN_CONF)["results"][0]
return {k: r[k] for k in ("choice", "confidence", "defer", "probs", "expected_level") if k in r}
@mcp.tool()
def decide_many(state: str, questions: dict) -> dict:
"""Several typed questions about one state: {name: {"type": "choice"|"noul"|"score",
"instructions": str, "criteria": {key: description} | [levels]}}."""
with lock:
return fdm.decide(model, tok, state, questions, defer_threshold=MIN_CONF)["answers"]
if __name__ == "__main__":
mcp.run()
```
Register it in Claude Desktop's `claude_desktop_config.json`:
```json
{
"mcpServers": {
"lightdec": {
"command": "C:\\path\\to\\python.exe",
"args": ["C:\\path\\to\\lightdec_mcp.py"],
"env": { "LIGHTDEC_REVISION": "<commit>", "LIGHTDEC_VARIANT": "int8" }
}
}
}
```
### 7.6 As an HTTP microservice
`pip install fastapi uvicorn`, then save `serve_lightdec.py`:
```python
import threading
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel
from lightdec import load_lightdec
fdm, model, tok = load_lightdec()
lock = threading.Lock()
app = FastAPI(title="LightDec")
class Decide(BaseModel):
state: str | dict
questions: dict | list
defer_threshold: float = 0.7
@app.get("/health")
def health():
return {"ok": True, "model": "LightDec", "version": model.fcfg.get("version")}
@app.post("/v1/decide")
def decide(req: Decide):
try:
with lock:
return fdm.decide(model, tok, req.state, req.questions, defer_threshold=req.defer_threshold)
except (ValueError, KeyError, TypeError) as exc:
raise HTTPException(422, str(exc))
```
Run it with `uvicorn serve_lightdec:app --port 9904 --workers 1`. Each worker holds its own copy of the model; scale out with more processes.
### 7.7 Without an LLM: a support intake step
```python
from lightdec import load_lightdec
fdm, model, tok = load_lightdec()
QUEUES = {"billing": "charges, invoices, payments, refunds", "returns": "returning or exchanging items",
"delivery": "shipping status, late or missing parcels", "accounts": "login, password, profile",
"human": "complaints or requests to speak to a person"}
def intake(message: str) -> dict:
a = fdm.decide(model, tok, message, {
"queue": {"type": "choice", "instructions": "Which team should handle this message?", "criteria": QUEUES},
"urgent": {"type": "noul", "instructions": "Does this need a reply within the hour?"},
})["answers"]
if a["queue"]["defer"]:
return {"action": "human_review", "suggestion": a["queue"]["choice"],
"reason": f"low confidence ({a['queue']['confidence']:.0%})"}
return {"action": "enqueue", "queue": a["queue"]["choice"],
"priority": "high" if a["urgent"]["p_true"] >= 0.5 else "normal",
"evidence": {"confidence": round(a["queue"]["confidence"], 3), "model": "LightDec"}}
```
---
## 8. Tuning for your domain
### 8.1 Change the deferral threshold
Pass `defer_threshold=` to `decide()`, or set `model.fcfg["defer_threshold"]`. This changes nothing in the model and is usually enough.
### 8.2 Recalibrate on your data
If ECE on your traffic (§6.4) is noticeably worse than 0.025, fit one extra temperature on top of the stored ones and save a recalibrated copy:
```python
import numpy as np, torch
# probs, gold: from §6.4 (probabilities already include the stored temperatures)
logp = [np.log(np.clip(p, 1e-12, 1)) for p in probs]
def nll(s):
return -np.mean([(lp / s)[g] - np.log(np.exp(lp / s).sum()) for lp, g in zip(logp, gold)])
s = min(np.linspace(0.5, 3.0, 51), key=nll)
print("extra temperature", s)
with torch.no_grad():
model.temperature.mul_(float(s))
fdm.save_falcondec(model, tok, "lightdec_recalibrated") # add int8=True for the compact variant
```
Fit on one labelled set and measure on another.
### 8.3 Fine-tune on your own decisions
Use the FalconDec training notebook (V2):
1. Export your logged and corrected decisions as JSONL: `{"state", "question", "options", "answer", "type"?, "task"?}`.
2. In cell 2, set `MODE="finetune"`, `FINETUNE_FROM="Falconsai/LightDec"`, `CUSTOM_DATA_JSONL="your_file.jsonl"`, and a preset.
3. Run the notebook. The version bumps automatically (1.0.0 → 1.0.1), and the report records the lineage.
Your tasks are up-weighted (`CUSTOM_WEIGHT`), while the public tasks keep the model general. Adding decisions with ISO-format dates and numeric tables is the most direct fix for the transfer weaknesses in §3. Check the regression gate (§6.5) before publishing.
### 8.4 Save and publish
```python
fdm.save_falcondec(model, tok, "lightdec_out") # fp16
fdm.save_falcondec(model, tok, "lightdec_out/compact-int8", int8=True)
from huggingface_hub import HfApi # needs a write token: huggingface-cli login
HfApi().upload_folder(folder_path="lightdec_out", repo_id="Falconsai/LightDec", commit_message="LightDec v1.0.x")
```
---
## 9. Architecture
```
[CLS] question [SEP] [MASK] option₁ [MASK] option₂ … [MASK] optionₖ [SEP] state [SEP]
│
Ettin-150M encoder (22 layers, hidden 768)
│
hidden state at each [MASK] + CLS context + question-type embedding
│
set transformer: 2 layers, 8 heads, no positional encoding → options attend to each other, order-equivariant
│
MLP → one logit per option → ÷ temperature[type, option-count bucket] → softmax
```
- **Option markers** (the approach Laya uses): every option is read at its own `[MASK]` token, so all options are scored in **one** pass.
- **Question first, state last**: when the input is too long, the tail of the state is truncated, never the options.
- **Adaptive option budget**: up to 24 tokens per option within a 192-token head budget that grows with the option count, and 2,048-token sequences above 24 options, so labels stay distinct. Above 96 options, `decide()` runs a tournament.
- **Typed primitives**: `noul` is rendered as a neutral two-option Yes/No choice; `score` keeps its level order and reports an expected level.
- **Calibration** lives in the model: a 3 × 4 temperature table (question type × option-count bucket: ≤2, 3–5, 6–12, >12).
- **int8 storage**: per-output-channel symmetric int8 for every weight matrix, fp16 elsewhere, dequantised on load.
---
## 10. Training data
**155,747 training, 13,148 validation and 17,498 test decisions** from 58 tasks. No source failed to load (`skipped_builders` is empty). Every example is a **decision**: `state`, `question`, `options`, `answer`, plus type, task and domain.
| Domain | Sources | Decisions built |
|---|---|---|
| Support | Bitext customer support | Intent routing and category routing; 30% of messages wrapped as JSON program state |
| Intents | CLINC150, MASSIVE (en), Banking77 *(held out)* | Intents as runtime-defined options, 2–48 per question; Banking77 also as one 77-option question |
| Code | CodeXGLUE code-to-text (6 languages), Devign, BigCloneBench, MBPP, HumanEval *(held out)* | Language ID, code↔description, function naming, vulnerability, clones, task→solution; bug spotting against single-fault mutants **verified to fail the unit tests** |
| Guardrails | Jailbreak classification, Civil Comments; deepset prompt-injections and AgentHarm *(held out)* | `noul` detection and refusal |
| Agentic | Counsel (human meta-evaluations of agent-step critiques), HotpotQA, AgentTrek | Step-error and critique-quality, retrieval routing and comparison yes/no; AgentTrek next-action type and finish-now. AgentTrek loaded without errors, but none of its decisions appear in the test split, so its contribution isn't measured |
| Workflows | LocalLLaMA/typed-decisions | `choice`/`noul`/`score` with the teacher's soft probabilities; test = the benchmark's 2,000-decision test split |
| Policy | Synthetic, executable (in-notebook) | Return windows, approval tiers, AND/OR eligibility, overdue invoices, table look-ups, counting, SLA urgency, access control, invoice totals; the test split uses **transfer** wording, currencies and date formats |
| Reasoning | ARC-Easy/Challenge and MMLU *(held out)*, OpenBookQA, SciQ, CommonsenseQA, QASC, HellaSwag, WinoGrande, MMLU auxiliary-train, BoolQ, GSM8K, AQuA-RAT, SNLI, MultiNLI, ANLI, SciTail | Multiple choice, yes/no, NLI, numeric answers with near-miss distractors |
| Classification | AG News, Yelp (ordinal); DAIR Emotion and SST-5 *(held out)* | Topic; 5-level `score` sentiment |
**Augmentation**: options are reshuffled every epoch (ordinal levels keep their order). In 8% of choice questions the gold answer is removed and "None of the above" becomes correct; in another 4% it is added as a distractor. **Balance**: tasks are sampled with p ∝ n^0.5 each epoch. **Leak guard**: training decisions whose (state, question) appears in validation or test were removed. Mind2Web was excluded (opt-in in the notebook). Held-out sources were never used for training, calibration or model selection.
Check each dataset's card for its license before redistributing derived data. AgentHarm is used only as a held-out evaluation, in line with its intended use.
---
## 11. Training procedure and calibration
### 11.1 Setup
| Setting | Value |
|---|---|
| Mode / preset | `scratch` from the pretrained backbone / `standard` |
| Data caps | ≤4,000 train, ≤300 validation, ≤300 test decisions per source split |
| Epochs | 2 (best: epoch 2) |
| Objective | Strictly proper scoring rules: log score + 0.5 × spherical score, + 1.0 × ranked probability score for `score` questions; soft targets (50/50 with the hard label) where a teacher distribution exists. These are the RLCD rewards, optimised with exact gradients |
| Optional RLCD stage | Off |
| Optimiser | AdamW (β 0.9/0.98, weight decay 0.01), encoder LR 4e-5 with layer-wise decay 0.9, head LR 3e-4, 6% warm-up, cosine decay, gradient clipping 1.0 |
| Batching | Token-budget batches (16,384 tokens, ≤32 decisions), length-bucketed |
| Weights kept | EMA of the weights (decay 0.999), best validation task-macro accuracy |
| Precision | bf16 autocast, TF32 matmuls; attention `auto` (FlashAttention-2 if installed, else SDPA) |
| Sequence | 512 tokens; 2,048 above 24 options; question ≤96 tokens; ≤24 tokens per option (more for code options) |
| Hardware / time | NVIDIA GeForce RTX 5090 Laptop GPU · **46 minutes** |
| Software | Python 3.14.4 · PyTorch 2.11.0+cu128 · transformers 5.17.0 |
| Seed | 42 |
| Epoch | Train loss | Train acc | Val macro | Val micro | Val NLL |
|---|---|---|---|---|---|
| 1 | 0.868 | 0.654 | 0.731 | 0.738 | 0.565 |
| 2 | 0.540 | 0.802 | **0.763** | 0.775 | 0.520 |
Validation was still improving at epoch 2, so a longer schedule (the `full` preset) is likely to help.
### 11.2 Fitted temperatures
| Question type | k ≤ 2 | k 3–5 | k 6–12 | k > 12 |
|---|---|---|---|---|
| choice | 1.707 | 1.352 | 1.466 | 1.349 |
| noul | 1.402 | 1.402 | 1.402 | 1.402 |
| score | 1.453 | 1.453 | 1.453 | 1.453 |
All temperatures are above 1, so the raw model was over-confident, as Laya's checkpoints are. `noul` and `score` use a single per-type temperature because their questions fall into one option-count bucket (2 options for `noul`, mostly 3–5 levels for `score`).
---
## 12. Operational notes
- **Concurrency.** `decide()` isn't internally locked. Serialise calls with a lock per process (as in §7.4–7.6) and scale out with processes.
- **Hardware.** On a GPU, expect tens of milliseconds per call; measure yours with §6.3. On CPU, use the int8 variant (loaded as fp32) and optionally dynamic int8 quantisation.
- **Input length.** 512 tokens by default. The question and options come first, so an over-long state loses its *end*. Put the decisive information early, or summarise.
- **Determinism.** Repeated identical calls give identical results on the same hardware and library versions. Under bf16 on GPU, batching different questions together changes padding and can move probabilities slightly (about 1e-2); near-ties can flip. Decisions with a clear margin don't change.
- **Traceability.** Log the model version, the Hub revision and the variant (fp16/int8) with every decision.
- **Offline use.** After the first download, set `HF_HUB_OFFLINE=1`, or save a local copy and load it by path.
---
## 13. Bias, risks and limitations
- **Quantitative and date reasoning.** Arithmetic (AQuA 0.259), table comparisons (0.290) and counting or summing (0.523) are near chance. Overdue-invoice checks with ISO dates score 0.407 with ECE 0.364, meaning confidently wrong.
- **Wide option sets.** Accuracy falls to 0.470 with all 77 Banking77 intents in one question; shortlist first.
- **Generalisation gap.** Held-out tasks average 0.567 against 0.725 overall. Expect lower accuracy on traffic unlike the training mix, and measure it (§6.4).
- **Guardrail transfer.** In-distribution jailbreak detection is 0.966, but held-out prompt-injection (0.647) and AgentHarm refusal (0.654) are much lower, with ECE 0.272 and 0.178. Don't rely on it as the only safety layer.
- **No completed head-to-head with proof_v2.** The comparisons in §2.4–2.5 use published numbers on different samples.
- **Closed world.** The model always picks one of your options. Add "None of the above" when appropriate.
- **Distribution-dependent calibration.** ECE was measured on this test mix; re-check it on your traffic and recalibrate if needed (§8.2).
- **Adversarial input.** The state is untrusted text. The model can't be instructed like an LLM, but crafted input can shift its scores. Don't make it the only safeguard before irreversible actions.
- **Data provenance.** Training data is English, largely crowd-sourced, templated, synthetic or scraped from public code and web tasks. Biases in these sources and in the backbone's pre-training can carry into decisions. Automated routing can systematically misroute users whose phrasing differs from the training data (dialects, non-native speakers, assistive phrasing); monitor misroutes by group where possible.
- **Oversight.** Not for high-stakes decisions without human review.
---
## 14. Versioning and lineage
| | proof_V_1 | proof_v2 | proof_v3 | **LightDec** |
|---|---|---|---|---|
| Model | falconsproof v1 | falconsproof v2.0.0 | FalconDec v1.0.0 (notebook V1) | **FalconDec v1.0.0 (notebook V2)** |
| Backbone | DistilBERT, 128 tokens | ModernBERT-base, 384 tokens | Ettin-150M, 512 / 2,048 | Ettin-150M, 512 / 2,048 |
| Preset / epochs | — | small / 1 per stage | small / 1 | **standard / 2** |
| Training decisions | — | 18,795 | ≤500 per task | **155,747** |
| Agentic data | — | — | — | AgentTrek, Counsel, HotpotQA |
| Test accuracy | — | 0.813 (23-task, code-heavy mix) | 0.518 micro / 0.535 macro | **0.725 / 0.725** (58-task mix) |
| ECE | — | 0.008 | 0.014 | 0.025 |
| Weights | — | ≈596 MB fp32 | 319 MB fp16 | 319 MB fp16 · 161 MB int8 |
LightDec is a fresh `scratch` run from the pretrained Ettin backbone (lineage: `jhu-clsp/ettin-encoder-150m`). It doesn't inherit proof_v3's or proof_v2's weights. Fine-tuning LightDec with the notebook bumps the patch version (1.0.0 → 1.0.1).
**Changelog.** LightDec 1.0.0: first release. Notebook V2, `standard` preset, 2 epochs, seed 42.
---
## 15. API reference
All functions live in `falcondec_modeling.py`.
| Function | Description |
|---|---|
| `load_falcondec(path, device=None, dtype=None, attn_implementation="sdpa")` | Loads a FalconDec directory (fp16 or int8) or Hub repo id. Returns `(model, tokenizer)`; cuda if available |
| `decide(model, tok, state, questions, defer_threshold=None, batch_size=32)` | Typed questions about one state. `questions` is a list or a `{key: question}` dict. Returns `{"results": [...], "answers": {key: result}}` |
| `score_items(model, tok, items, batch_size=32)` | Batch scoring. `items = [{"state", "question", "options", "type"?, "option_tokens"?, "seq_len"?}]`. Returns a calibrated probability array per item |
| `save_falcondec(model, tok, out_dir, int8=False, extra_files=None)` | Writes a self-contained directory (copies the modeling file) |
| `quantize_int8(state_dict)` / `dequantize_int8(state_dict)` | The per-channel int8 codec used for `compact-int8/` |
| `assemble(...)`, `collate_features(...)` | Low-level sequence building and batching |
**Question fields**: `type` (`choice` / `noul` / `score`, default `choice`); `question` or `instructions`; `options` (list) or `criteria` (dict `{key: description}` for choice, list of levels for score); `labels` (`{"true": …, "false": …}` wording for noul); `option_tokens` and `seq_len` (optional per-question budgets).
**Model attributes**: `model.fcfg` (the live config: layout, `special`, `defer_threshold`, `version`, `lineage`); `model.temperature` (3 × 4 tensor); `model.num_parameters()`; `model.device`.
---
## 16. Citation and references
```bibtex
@misc{falconsai_lightdec_2026,
title = {LightDec: a lightweight, single-pass, typed, calibrated decision model for agentic systems},
author = {{Falconsai}},
year = {2026},
howpublished = {\url{https://huggingface.co/Falconsai/LightDec}},
note = {FalconDec architecture, Ettin-150M backbone; successor to Falconsai/proof_v3}
}
```
**Methods.** Warner et al. (2024), *ModernBERT*. Weller et al. (2025), *Ettin* encoders. Gneiting and Raftery (2007), *Strictly Proper Scoring Rules*. Guo et al. (2017), *On Calibration of Modern Neural Networks*. Geifman and El-Yaniv (2017), *Selective Classification*. Zaheer et al. (2017), *Deep Sets*; Lee et al. (2019), *Set Transformer*. Williams (1992), REINFORCE; Shao et al. (2024), GRPO. Hinton, Vinyals and Dean (2015), *Distillation*. Related decision models: Laya (Convai Innovations), TypeSafe Jev, Together Tev1.
**Data.** ARC, OpenBookQA, SciQ, CommonsenseQA, QASC, HellaSwag, WinoGrande, MMLU, BoolQ, GSM8K, AQuA-RAT, SNLI, MultiNLI, ANLI, SciTail, CLINC150, MASSIVE, Bitext, Banking77, AG News, Yelp, DAIR Emotion, SST-5, jailbreak-classification, Civil Comments, deepset prompt-injections, AgentHarm, CodeXGLUE, MBPP, HumanEval, LocalLLaMA/typed-decisions, AgentTrek, Counsel, HotpotQA.
Report issues, misroutes or evaluation results through the Community tab of this repository.
---
This card is generated from the surgical record itself; the package's
`lineage.intoto.jsonl` is the signed source of truth (verify it free at
the Surgeon's public verifier or with the bundled `verify_attestation.py`).
## Architecture
- Identification: **NLP · Small Language Model (SLM)** (98% confidence)
- Source format: `safetensors` · Intended task: not declared
- `config.json`: synthesized from the anatomy (no source config.json); model_type omitted — no architecture name in the source (QA-F-126)
- Source license: apache-2.0
- Lineage chain: 1 surgery (no prior attestation reachable) · Falconsai/LightDec
- Post-surgery totals: 159,654,157 parameters ·
168 tensors
- Compute estimate: 15.02439 GFLOPs (comparison
metric, not a measurement)
## Provenance & operations
- Parents: Falconsai/LightDec/model.safetensors
- Operations performed: forensics×1, hub_push×1, imaging×3, load×1, test×1
- Weight merges recorded: 0
- Quantized tensors (F32→F16): 0
## Surgery Log (ordered)
1. **load** — hub:Falconsai/LightDec/model.safetensors (319.3 MB, safetensors)
2. **hub_push** — Falconsai/LightDec ← 8 file(s) (attestation included)
3. **test** — PASS
4. **forensics** — clean — no anomalies
5. **imaging** — synthetic 16 token ids → [1, 2304] [WEAK]
6. **imaging** — synthetic 16 token ids → [1, 2304] [WEAK]
7. **imaging** — 41 token ids (hash-mapped, vocab 50368) → [1, 2304] [WEAK]
## Validation
- Tissue imaging: WEAK
- Structural integrity is testable offline via the packaged
`load_and_test.py`.
## Compliance note
The signed attestation + this card together document model composition,
modification history, and validation evidence — the record structure
technical-documentation obligations (e.g. EU AI Act Annex IV) ask for.
This is evidence, not legal advice.
---
*Operated with Model Surgeon — verify this package at https://surgeon.falcons.ai/verify*
*© 2026 FALCONS.AI — Model Surgeon record format. The model weights remain their owner's.*