File size: 18,849 Bytes
d9615c8
 
66ee87e
 
 
d9615c8
66ee87e
 
21c5089
d9615c8
 
 
66ee87e
 
 
 
6012dcc
66ee87e
 
6012dcc
66ee87e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
6012dcc
66ee87e
 
 
 
 
 
 
 
 
 
 
 
43c3b3c
66ee87e
43c3b3c
66ee87e
 
 
 
 
 
 
 
 
 
 
 
d8c255d
66ee87e
 
 
6012dcc
66ee87e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
6012dcc
66ee87e
 
 
 
 
 
 
 
 
d8c255d
66ee87e
 
 
 
 
 
996a2db
66ee87e
 
 
 
 
d8c255d
66ee87e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
21c5089
66ee87e
 
 
 
 
 
 
 
 
 
 
 
996a2db
d8c255d
66ee87e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
---
title: DecisionLab
emoji: ⚖️
colorFrom: red
colorTo: gray
sdk: gradio
sdk_version: 6.28.0
python_version: "3.12"
app_file: app.py
pinned: false
---

# DecisionLab: your decision models vs Laya

A self-hosted web lab that puts decision models on the same decisions and compares their answers, confidence and speed:

- **Four models from the Hugging Face Hub**, in this order:
  - **LightDec_Arthur**: [Falconsai/LightDec_Arthur](https://huggingface.co/Falconsai/LightDec_Arthur), the byte-level Arthur model (about 12M parameters). Only its `config.json` and `model.safetensors` are downloaded; the code that runs it ships with DecisionLab, copied verbatim from the arthur_v0_8_0 notebook.
  - **LightDec_V2**: [Falconsai/LightDec_V2](https://huggingface.co/Falconsai/LightDec_V2), FalconDec on the Ettin-150M encoder (about 160M parameters). Its `falcondec_modeling.py` runs only if its sha256 is on the allowlist; otherwise its card shows the hash and how to trust it.
  - **Enterprise Reflux Laya V2.1**: [yasserrmd/enterprise-reflux-laya-v21](https://huggingface.co/yasserrmd/enterprise-reflux-laya-v21), a Laya fine-tune for ranking enterprise actions (ModernBERT-large, like Laya), loaded with the same `laya` runtime. Its model card marks it a research prototype, not for production.
  - **Laya**: [convaiinnovations/laya](https://huggingface.co/convaiinnovations/laya), by Convai Innovations, ModernBERT-large, 421M parameters, Apache-2.0.
- **Every model in the `models` folder**, added after those and marked "(local)":
  - **LightDec (FalconDec) models**: a folder with a `falcondec_config.json`, named from that config.
  - **Arthur models**: a folder with Arthur's `config.json` and `model.safetensors`, named "Arthur <tier>". A folder trained with a different input layout shows an error card naming the mismatch.

The container is named `DecisionLab` and serves the UI and API on **port 9910**.

## Run it

Put the models you want to compare in the `models` folder next to the project, one folder per model (LightDec or Arthur). In PowerShell, from the `DecisionLab` folder:

```powershell
New-Item -ItemType Directory -Force models | Out-Null
Expand-Archive -Path "$HOME\Downloads\LightDec_V2_Long-v1_0_0.zip" -DestinationPath models -Force
Expand-Archive -Path "$HOME\Downloads\LightDec-v1_0_2.zip" -DestinationPath models -Force
Expand-Archive -Path "$HOME\Downloads\arthur-base.zip" -DestinationPath models -Force
Get-ChildItem models -Directory | Where-Object { (Test-Path (Join-Path $_.FullName "falcondec_config.json")) -or ((Test-Path (Join-Path $_.FullName "config.json")) -and (Test-Path (Join-Path $_.FullName "model.safetensors"))) } | Select-Object Name
```

The last command lists exactly the folders DecisionLab will pick up.

The page and the API need no token. The lab listens on this computer only (`127.0.0.1`). Setting `DLAB_BIND=0.0.0.0` in `.env` opens it to other machines, and then anyone who can reach it can use the page and the API, so only do that on a network you trust.

```powershell
docker compose up -d --build
# open http://localhost:9910
```

Or without Compose:

```powershell
docker build -t decisionlab .
docker run -d --name DecisionLab -p 9910:9910 -v decisionlab-hf:/data/hf -v "${PWD}\models:/models:ro" decisionlab
```

The first start downloads the four Hub models (about 2 GB) into the `decisionlab-hf` volume. The page shows each model's loading progress, and later starts load from the cache in seconds.

### NVIDIA GPU

```bash
docker build --build-arg TORCH_INDEX_URL=https://download.pytorch.org/whl/cu128 -t decisionlab:gpu .
docker run -d --name DecisionLab --gpus all -p 9910:9910 -v decisionlab-hf:/data/hf -v "${PWD}\models:/models:ro" decisionlab:gpu
```

With Compose, set the same build arg and uncomment the `deploy` block in `docker-compose.yml`. The NVIDIA Container Toolkit must be installed on the host. On CPU, expect a few hundred milliseconds to a few seconds per call; on a GPU, tens of milliseconds.

## Hugging Face Space

This folder is also a Gradio Space, set up for ZeroGPU hardware (it also runs on a CPU Space). `app.py` starts the Space through Gradio's own `launch()` (ZeroGPU needs it) and puts the lab's routes in front of Gradio's, so the page at `/` and the API at `/api/*` are the container's. Gradio's `decide` endpoint stays available for `gradio_client`, running the same decision code.

On ZeroGPU a GPU is attached for each decision. The models load before the Space starts serving (ZeroGPU prepares them at startup). Each model's time is measured on the GPU, and the round trip also includes the GPU hand-off. Each decision uses your ZeroGPU quota, so a full scoreboard run (one decision per demo) uses more of it.

There is no token on the Space: on a public Space anyone can use the page and the API. Make the Space private on Hugging Face if that is not what you want.

Push it (PowerShell or bash; needs git and git-lfs, and a Hugging Face write token when git asks for a password):

```bash
# create an empty Space first on huggingface.co (SDK: Gradio), then:
git clone https://huggingface.co/spaces/<your-user>/<your-space>
cd <your-space>
git lfs install
# copy the contents of this DecisionLab folder into it (not the folder itself), then:
git add .
git commit -m "DecisionLab 2.3.1"
git push
```

The Space downloads the four Hub models on first start. Model folders you put in `models/` are loaded too; their weights go through Git LFS (`.gitattributes`).

Call it from Python:

```python
from gradio_client import Client
client = Client("<your-user>/<your-space>")
result = client.predict("I was charged twice. Refund me or I cancel.",
                        '{"churn": {"type": "noul", "instructions": "The customer threatens to leave"}}',
                        "", api_name="/decide")
```

The plain HTTP API works on the Space too, at `https://<your-user>-<your-space>.hf.space/api/...`.

## Configuration

| Variable | Default | Meaning |
|---|---|---|
| `DEVICE` | `auto` | `auto` (GPU if available), `cuda` or `cpu` |
| `DLAB_BIND` | `127.0.0.1` | Host address that publishes port 9910 (Compose, from `.env`) |
| `TRUSTED_MODELING_SHA256` | – | Extra trusted hashes of `falcondec_modeling.py`, comma-separated |
| `MAX_BODY_BYTES` | `262144` | Largest request body the API accepts |
| `MAX_PENDING_DECIDES` | `4` | `/api/decide` requests in flight before the API answers 429 |
| `MODELS_DIR` | `/models` | Folder inside the container that is scanned for LightDec and Arthur models |
| `LIGHTDEC_VARIANT` | `fp16` | `fp16` (each model folder itself) or `int8` (its `compact-int8` subfolder, dequantised on load), for every discovered model |
| `LIGHTDEC_ARTHUR_REPO` | `Falconsai/LightDec_Arthur` | Hub repo shown as LightDec_Arthur |
| `LIGHTDEC_V2_REPO` | `Falconsai/LightDec_V2` | Hub repo shown as LightDec_V2 |
| `REFLUX_LAYA_REPO` | `yasserrmd/enterprise-reflux-laya-v21` | Hub repo shown as Enterprise Reflux Laya V2.1 (a Laya checkpoint) |
| `LAYA_REPO` | `convaiinnovations/laya` | Laya checkpoint loaded with `laya.load()` |
| `LAYA_FALLBACK_REPO` | – | Optional second Laya repo, tried if the first fails |
| `LOAD_ORDER` | comparison order | Which models to load at start, in order, by key (see `/api/status`) |
| `MAX_OPTIONS` | `40` | Largest option set the API accepts |
| `HF_TOKEN` | – | Only for gated or private repos |

## Using the lab

1. **Setup** shows each model's status, size, device, load time and how it defines confidence.
2. **Decision**: pick one of the 47 demos (tabs group them into triage, routing and planning, loop control, guardrails, and agent security) or write your own state and questions. Nothing is executed; every demo is a decision only. Questions use the Laya/Jev JSON shape:
   ```json
   {"team":    {"type": "choice", "instructions": "Which team?", "criteria": {"bug": "Something is broken", "sales": "Pricing"}},
    "urgency": {"type": "score",  "instructions": "How urgent?", "criteria": ["Can wait", "Today", "Right now"]},
    "angry":   {"type": "noul",   "instructions": "The customer sounds angry"}}
   ```
3. **Verdicts**: one probability bar per model for every option, each model's answer, confidence, an *act* or *defer* badge against your threshold, and how many models agree.
4. **Scoreboard** runs all 47 demos (88 questions) or one group; only the 54 questions with reference answers count toward matches and the Agentic Use Score. It highlights the best model on each metric (reference matches, median time, answers acted on, correct when acting, confident mistakes, mistakes deferred) and ranks every model by the **Agentic Use Score** below, then breaks matches and confident mistakes down by group.

**Confidence.** LightDec and Arthur models report the probability of their top option (Arthur's is temperature-calibrated per question type and option count). Laya reports 1 − normalised entropy for choice and score questions, and the larger of P(yes) and P(no) for yes/no questions. The lab computes both definitions for every model; pick one with the *Confidence means* selector.

## Demos

47 demos in 5 groups. The tab numbers match the table. Demos 25–47 are the **Agent security** group: the states and questions of the Athr_Agent_Sec nano demo set (`app/demos_nano.json`), with no reference answers, so they show each model's answers, confidence and agreement but are not scored. A scoreboard run of only these demos shows no Agentic Use Score (there is nothing to score it against), just what each model would act on, its speed and the agreement.

| # | Group | Demo | What it tests |
|---|---|---|---|
| 1 | Triage | Support ticket | Exact state and questions from the layaForWeb README |
| 2 | Triage | Lead scoring | From the article, rebuilt |
| 3 | Triage | Patient message | From the article, rebuilt; emergency symptoms |
| 4 | Triage | Delivery exception | From the article, rebuilt |
| 5 | Triage | Product review | From the article, rebuilt; mixed review with a defect |
| 6 | Triage | Human handoff | Third request for a person |
| 7 | Routing & planning | Tool router | First tool for a multi-step request; flags data leaving the company |
| 8 | Routing & planning | Retrieval router | Which knowledge source a RAG agent should search |
| 9 | Routing & planning | Model and tool router | Exact math goes to code, not a language model |
| 10 | Routing & planning | Clarify or proceed | Ask for missing booking details instead of guessing |
| 11 | Routing & planning | Policy check against state | Apply a written approval policy to a structured request |
| 12 | Routing & planning | Plan review | Migration before backup, no rollback |
| 13 | Loop control | Stop condition | Is the research task really complete? |
| 14 | Loop control | Stuck in a loop | Four identical failing calls; fix the arguments |
| 15 | Loop control | Tool error triage | A 429 rate limit; back off and retry |
| 16 | Loop control | Step verifier | Refund sent to the wrong payment method |
| 17 | Loop control | Grounding check | The RAG draft contradicts its source |
| 18 | Guardrails | Agent about to delete a table | From the article, rebuilt |
| 19 | Guardrails | Injection in a tool result | Indirect prompt injection in fetched content |
| 20 | Guardrails | Code change gate | Pull request touching auth |
| 21 | Guardrails | Payment approval | Changed bank details and an urgent wire |
| 22 | Guardrails | Outbound data check | Social security numbers in an email to a partner |
| 23 | Guardrails | Memory write | Save the preference, never the password |
| 24 | Guardrails | Authorization scope | A user asks to change someone else's account |

Reference answers are a careful human reading, not ground truth. Edit groups, demos, references and stakes in `app/demos.py`; the UI builds its tabs and counts from `GROUPS` and `DEMOS`, so nothing else needs to change when you add a demo.

## The Agentic Use Score

When an agent acts on a model's answer, the costly mistakes are the confident ones it acts on. A deferred mistake costs a quick human review; a deferred correct answer costs a little time. The score (0–100) combines the scoreboard's metrics with that weighting:

| Component | Max points | What it measures |
|---|---|---|
| Right when it acts | 35 | Correct answers among those that clear the threshold, with a small prior: (correct + 1) / (acted + 2) |
| Flags its own mistakes | 25 | Share of wrong answers that fell below the threshold and would go to a person |
| Overall accuracy | 15 | Answers that match the reference |
| Handles on its own | 15 | Answers that clear the threshold |
| Speed | 10 | 1.0 at ≤ 50 ms median, 0 at ≥ 1 s, log scale in between |

Each component earns its max points times the model's rate on it (right 60% of the time when acting = 21 of 35 points), and the points add up to the score. High-stakes questions count twice in every component except speed. The score is computed in the browser, so it updates when you change the threshold or confidence measure.

## API

The API needs no token. Requests are limited to 256 KB, 20 questions, 2,000 characters of question text and 500 characters per option; at most 4 decisions run at once (more get 429).

| Method | Path | Body / result |
|---|---|---|
| `GET` | `/api/health` | `{"ok": true}` (used by the container health check) |
| `GET` | `/api/status` | Load status, device and details for each model, `order` (comparison order) and `env.models_dir` |
| `GET` | `/api/demos` | All demos, or one group with `?group=guardrails` |
| `GET` | `/api/groups` | Group id, family, label, description and demo/question counts |
| `POST` | `/api/decide` | `{"state": str or object, "questions": {...}, "models": [keys from /api/status]}` (every model if omitted; an unknown key is 422) → per model: `answers` (normalised) and `ms` |
| `POST` | `/api/reload/{key}` | Retry loading a model; keys are the model folder names in lower case with `_` separators, plus `laya` |

Each normalised answer contains `choice`, `probs` (per option), `top_prob`, `entropy_conf`, `laya_conf`, plus `expected_level` for score questions and `p_true` for yes/no questions.

```powershell
$body = @{
  state     = "I was charged twice for March. Refund the duplicate or I cancel."
  questions = @{
    dept  = @{ type = "choice"; instructions = "Which team?"; criteria = @{ billing = "payments"; tech = "bugs" } }
    churn = @{ type = "noul"; instructions = "The customer threatens to leave" }
  }
} | ConvertTo-Json -Depth 5
Invoke-RestMethod -Method Post -Uri http://localhost:9910/api/decide -ContentType "application/json" `
  -Body $body | ConvertTo-Json -Depth 6
```

## This application is now more secure with the following updates

- Execution of untrusted model code from the models folder was identified and fixed: only known FalconDec modeling files are run.
- Unlimited request sizes were identified and fixed: bodies, questions and option text are capped.
- Unbounded queuing of decision requests was identified and fixed: excess requests get 429 instead of starving the server.
- A model-reload race was identified and fixed: a model can only load once at a time.
- Missing browser security headers were identified and fixed: strict Content Security Policy, no framing, no sniffing, no referrer, no caching of API responses.
- Public API documentation endpoints were identified and fixed: `/docs`, `/redoc` and `/openapi.json` are off.
- Silently ignored and duplicated model keys were identified and fixed: unknown keys are rejected and duplicates run once.
- Publishing the port on every network interface was identified and fixed: the lab listens on this computer only unless you choose otherwise.
- An unencoded model key in a request URL was identified and fixed.
- Unpinned dependencies were identified and fixed: every Python package, torch included, is installed at an exact version and the build fails if the set is inconsistent.

## Layout

```
DecisionLab/
├── Dockerfile
├── docker-compose.yml
├── constraints.txt    # every other package, exact versions (from pip freeze)
├── .env.example       # copy to .env for local settings (.env is never in the image)
├── app.py             # Hugging Face Space entry point (Gradio SDK, ZeroGPU)
├── requirements.txt   # Space dependencies
├── requirements-container.txt  # container dependencies (with constraints.txt)
├── models/            # your LightDec and Arthur models, one folder each (not in the image)
├── app/
│   ├── main.py        # FastAPI app: UI + API on :9910
│   ├── models.py      # LightDec (local folder or Hub) and Laya backends, timing
│   ├── registry.py    # finds the models in MODELS_DIR, checks their modeling code (no torch)
│   ├── arthur.py      # Arthur network and inference, verbatim from arthur_v0_8_0.ipynb
│   ├── arthur_io.py   # Arthur input layout and calibration (no torch)
│   ├── scoring.py     # normalises every model answer into one comparable record (no torch)
│   ├── security.py    # body limit and security headers (plain ASGI)
│   ├── validation.py  # checks /api/decide question sets, sizes and model keys (no web framework)
│   ├── demos.py       # groups, the 47 demos (24 scored + 23 nano), reference answers and high-stakes flags
│   ├── demos_nano.json  # Agent security: states and questions only
│   └── static/        # index.html, app.css, app.js, logo.png
└── tests/             # unittest suite, copied into the image
```

## Tests

The suite uses Python's built-in `unittest`, so it needs no extra packages. It checks the demo set (unique ids, known groups, valid reference answers, stakes, every demo passes request validation), request validation, and answer normalisation. `tests/test_api.py` also checks that a bad `/api/decide` request returns 422; it needs the app's real imports (fastapi, pydantic, torch), so outside the container it is skipped with the reason printed.

After `docker compose up -d --build`, in PowerShell:

```powershell
docker exec DecisionLab python -m unittest discover -s tests -v
```

The last line should read `OK` with no skips. Tests do not load the models and do not call Hugging Face.

## Notes

- Models run one after the other under a lock, so timings never compete for the device. Run one uvicorn worker per container; each worker would hold its own copy of every model.
- LightDec is loaded with the `falcondec_modeling.py` shipped in its repo; Laya with the `laya` package (`laya.load`).
- Laya is Apache-2.0. The LightDec checkpoints tested here state apache-2.0 in their model cards; check any other model you add. Laya is by Convai Innovations; this lab is not affiliated with Convai Innovations, TypeSafe AI or the layaForWeb project.