Spaces:
Running on Zero
Running on Zero
File size: 18,849 Bytes
d9615c8 66ee87e d9615c8 66ee87e 21c5089 d9615c8 66ee87e 6012dcc 66ee87e 6012dcc 66ee87e 6012dcc 66ee87e 43c3b3c 66ee87e 43c3b3c 66ee87e d8c255d 66ee87e 6012dcc 66ee87e 6012dcc 66ee87e d8c255d 66ee87e 996a2db 66ee87e d8c255d 66ee87e 21c5089 66ee87e 996a2db d8c255d 66ee87e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 | ---
title: DecisionLab
emoji: ⚖️
colorFrom: red
colorTo: gray
sdk: gradio
sdk_version: 6.28.0
python_version: "3.12"
app_file: app.py
pinned: false
---
# DecisionLab: your decision models vs Laya
A self-hosted web lab that puts decision models on the same decisions and compares their answers, confidence and speed:
- **Four models from the Hugging Face Hub**, in this order:
- **LightDec_Arthur**: [Falconsai/LightDec_Arthur](https://huggingface.co/Falconsai/LightDec_Arthur), the byte-level Arthur model (about 12M parameters). Only its `config.json` and `model.safetensors` are downloaded; the code that runs it ships with DecisionLab, copied verbatim from the arthur_v0_8_0 notebook.
- **LightDec_V2**: [Falconsai/LightDec_V2](https://huggingface.co/Falconsai/LightDec_V2), FalconDec on the Ettin-150M encoder (about 160M parameters). Its `falcondec_modeling.py` runs only if its sha256 is on the allowlist; otherwise its card shows the hash and how to trust it.
- **Enterprise Reflux Laya V2.1**: [yasserrmd/enterprise-reflux-laya-v21](https://huggingface.co/yasserrmd/enterprise-reflux-laya-v21), a Laya fine-tune for ranking enterprise actions (ModernBERT-large, like Laya), loaded with the same `laya` runtime. Its model card marks it a research prototype, not for production.
- **Laya**: [convaiinnovations/laya](https://huggingface.co/convaiinnovations/laya), by Convai Innovations, ModernBERT-large, 421M parameters, Apache-2.0.
- **Every model in the `models` folder**, added after those and marked "(local)":
- **LightDec (FalconDec) models**: a folder with a `falcondec_config.json`, named from that config.
- **Arthur models**: a folder with Arthur's `config.json` and `model.safetensors`, named "Arthur <tier>". A folder trained with a different input layout shows an error card naming the mismatch.
The container is named `DecisionLab` and serves the UI and API on **port 9910**.
## Run it
Put the models you want to compare in the `models` folder next to the project, one folder per model (LightDec or Arthur). In PowerShell, from the `DecisionLab` folder:
```powershell
New-Item -ItemType Directory -Force models | Out-Null
Expand-Archive -Path "$HOME\Downloads\LightDec_V2_Long-v1_0_0.zip" -DestinationPath models -Force
Expand-Archive -Path "$HOME\Downloads\LightDec-v1_0_2.zip" -DestinationPath models -Force
Expand-Archive -Path "$HOME\Downloads\arthur-base.zip" -DestinationPath models -Force
Get-ChildItem models -Directory | Where-Object { (Test-Path (Join-Path $_.FullName "falcondec_config.json")) -or ((Test-Path (Join-Path $_.FullName "config.json")) -and (Test-Path (Join-Path $_.FullName "model.safetensors"))) } | Select-Object Name
```
The last command lists exactly the folders DecisionLab will pick up.
The page and the API need no token. The lab listens on this computer only (`127.0.0.1`). Setting `DLAB_BIND=0.0.0.0` in `.env` opens it to other machines, and then anyone who can reach it can use the page and the API, so only do that on a network you trust.
```powershell
docker compose up -d --build
# open http://localhost:9910
```
Or without Compose:
```powershell
docker build -t decisionlab .
docker run -d --name DecisionLab -p 9910:9910 -v decisionlab-hf:/data/hf -v "${PWD}\models:/models:ro" decisionlab
```
The first start downloads the four Hub models (about 2 GB) into the `decisionlab-hf` volume. The page shows each model's loading progress, and later starts load from the cache in seconds.
### NVIDIA GPU
```bash
docker build --build-arg TORCH_INDEX_URL=https://download.pytorch.org/whl/cu128 -t decisionlab:gpu .
docker run -d --name DecisionLab --gpus all -p 9910:9910 -v decisionlab-hf:/data/hf -v "${PWD}\models:/models:ro" decisionlab:gpu
```
With Compose, set the same build arg and uncomment the `deploy` block in `docker-compose.yml`. The NVIDIA Container Toolkit must be installed on the host. On CPU, expect a few hundred milliseconds to a few seconds per call; on a GPU, tens of milliseconds.
## Hugging Face Space
This folder is also a Gradio Space, set up for ZeroGPU hardware (it also runs on a CPU Space). `app.py` starts the Space through Gradio's own `launch()` (ZeroGPU needs it) and puts the lab's routes in front of Gradio's, so the page at `/` and the API at `/api/*` are the container's. Gradio's `decide` endpoint stays available for `gradio_client`, running the same decision code.
On ZeroGPU a GPU is attached for each decision. The models load before the Space starts serving (ZeroGPU prepares them at startup). Each model's time is measured on the GPU, and the round trip also includes the GPU hand-off. Each decision uses your ZeroGPU quota, so a full scoreboard run (one decision per demo) uses more of it.
There is no token on the Space: on a public Space anyone can use the page and the API. Make the Space private on Hugging Face if that is not what you want.
Push it (PowerShell or bash; needs git and git-lfs, and a Hugging Face write token when git asks for a password):
```bash
# create an empty Space first on huggingface.co (SDK: Gradio), then:
git clone https://huggingface.co/spaces/<your-user>/<your-space>
cd <your-space>
git lfs install
# copy the contents of this DecisionLab folder into it (not the folder itself), then:
git add .
git commit -m "DecisionLab 2.3.1"
git push
```
The Space downloads the four Hub models on first start. Model folders you put in `models/` are loaded too; their weights go through Git LFS (`.gitattributes`).
Call it from Python:
```python
from gradio_client import Client
client = Client("<your-user>/<your-space>")
result = client.predict("I was charged twice. Refund me or I cancel.",
'{"churn": {"type": "noul", "instructions": "The customer threatens to leave"}}',
"", api_name="/decide")
```
The plain HTTP API works on the Space too, at `https://<your-user>-<your-space>.hf.space/api/...`.
## Configuration
| Variable | Default | Meaning |
|---|---|---|
| `DEVICE` | `auto` | `auto` (GPU if available), `cuda` or `cpu` |
| `DLAB_BIND` | `127.0.0.1` | Host address that publishes port 9910 (Compose, from `.env`) |
| `TRUSTED_MODELING_SHA256` | – | Extra trusted hashes of `falcondec_modeling.py`, comma-separated |
| `MAX_BODY_BYTES` | `262144` | Largest request body the API accepts |
| `MAX_PENDING_DECIDES` | `4` | `/api/decide` requests in flight before the API answers 429 |
| `MODELS_DIR` | `/models` | Folder inside the container that is scanned for LightDec and Arthur models |
| `LIGHTDEC_VARIANT` | `fp16` | `fp16` (each model folder itself) or `int8` (its `compact-int8` subfolder, dequantised on load), for every discovered model |
| `LIGHTDEC_ARTHUR_REPO` | `Falconsai/LightDec_Arthur` | Hub repo shown as LightDec_Arthur |
| `LIGHTDEC_V2_REPO` | `Falconsai/LightDec_V2` | Hub repo shown as LightDec_V2 |
| `REFLUX_LAYA_REPO` | `yasserrmd/enterprise-reflux-laya-v21` | Hub repo shown as Enterprise Reflux Laya V2.1 (a Laya checkpoint) |
| `LAYA_REPO` | `convaiinnovations/laya` | Laya checkpoint loaded with `laya.load()` |
| `LAYA_FALLBACK_REPO` | – | Optional second Laya repo, tried if the first fails |
| `LOAD_ORDER` | comparison order | Which models to load at start, in order, by key (see `/api/status`) |
| `MAX_OPTIONS` | `40` | Largest option set the API accepts |
| `HF_TOKEN` | – | Only for gated or private repos |
## Using the lab
1. **Setup** shows each model's status, size, device, load time and how it defines confidence.
2. **Decision**: pick one of the 47 demos (tabs group them into triage, routing and planning, loop control, guardrails, and agent security) or write your own state and questions. Nothing is executed; every demo is a decision only. Questions use the Laya/Jev JSON shape:
```json
{"team": {"type": "choice", "instructions": "Which team?", "criteria": {"bug": "Something is broken", "sales": "Pricing"}},
"urgency": {"type": "score", "instructions": "How urgent?", "criteria": ["Can wait", "Today", "Right now"]},
"angry": {"type": "noul", "instructions": "The customer sounds angry"}}
```
3. **Verdicts**: one probability bar per model for every option, each model's answer, confidence, an *act* or *defer* badge against your threshold, and how many models agree.
4. **Scoreboard** runs all 47 demos (88 questions) or one group; only the 54 questions with reference answers count toward matches and the Agentic Use Score. It highlights the best model on each metric (reference matches, median time, answers acted on, correct when acting, confident mistakes, mistakes deferred) and ranks every model by the **Agentic Use Score** below, then breaks matches and confident mistakes down by group.
**Confidence.** LightDec and Arthur models report the probability of their top option (Arthur's is temperature-calibrated per question type and option count). Laya reports 1 − normalised entropy for choice and score questions, and the larger of P(yes) and P(no) for yes/no questions. The lab computes both definitions for every model; pick one with the *Confidence means* selector.
## Demos
47 demos in 5 groups. The tab numbers match the table. Demos 25–47 are the **Agent security** group: the states and questions of the Athr_Agent_Sec nano demo set (`app/demos_nano.json`), with no reference answers, so they show each model's answers, confidence and agreement but are not scored. A scoreboard run of only these demos shows no Agentic Use Score (there is nothing to score it against), just what each model would act on, its speed and the agreement.
| # | Group | Demo | What it tests |
|---|---|---|---|
| 1 | Triage | Support ticket | Exact state and questions from the layaForWeb README |
| 2 | Triage | Lead scoring | From the article, rebuilt |
| 3 | Triage | Patient message | From the article, rebuilt; emergency symptoms |
| 4 | Triage | Delivery exception | From the article, rebuilt |
| 5 | Triage | Product review | From the article, rebuilt; mixed review with a defect |
| 6 | Triage | Human handoff | Third request for a person |
| 7 | Routing & planning | Tool router | First tool for a multi-step request; flags data leaving the company |
| 8 | Routing & planning | Retrieval router | Which knowledge source a RAG agent should search |
| 9 | Routing & planning | Model and tool router | Exact math goes to code, not a language model |
| 10 | Routing & planning | Clarify or proceed | Ask for missing booking details instead of guessing |
| 11 | Routing & planning | Policy check against state | Apply a written approval policy to a structured request |
| 12 | Routing & planning | Plan review | Migration before backup, no rollback |
| 13 | Loop control | Stop condition | Is the research task really complete? |
| 14 | Loop control | Stuck in a loop | Four identical failing calls; fix the arguments |
| 15 | Loop control | Tool error triage | A 429 rate limit; back off and retry |
| 16 | Loop control | Step verifier | Refund sent to the wrong payment method |
| 17 | Loop control | Grounding check | The RAG draft contradicts its source |
| 18 | Guardrails | Agent about to delete a table | From the article, rebuilt |
| 19 | Guardrails | Injection in a tool result | Indirect prompt injection in fetched content |
| 20 | Guardrails | Code change gate | Pull request touching auth |
| 21 | Guardrails | Payment approval | Changed bank details and an urgent wire |
| 22 | Guardrails | Outbound data check | Social security numbers in an email to a partner |
| 23 | Guardrails | Memory write | Save the preference, never the password |
| 24 | Guardrails | Authorization scope | A user asks to change someone else's account |
Reference answers are a careful human reading, not ground truth. Edit groups, demos, references and stakes in `app/demos.py`; the UI builds its tabs and counts from `GROUPS` and `DEMOS`, so nothing else needs to change when you add a demo.
## The Agentic Use Score
When an agent acts on a model's answer, the costly mistakes are the confident ones it acts on. A deferred mistake costs a quick human review; a deferred correct answer costs a little time. The score (0–100) combines the scoreboard's metrics with that weighting:
| Component | Max points | What it measures |
|---|---|---|
| Right when it acts | 35 | Correct answers among those that clear the threshold, with a small prior: (correct + 1) / (acted + 2) |
| Flags its own mistakes | 25 | Share of wrong answers that fell below the threshold and would go to a person |
| Overall accuracy | 15 | Answers that match the reference |
| Handles on its own | 15 | Answers that clear the threshold |
| Speed | 10 | 1.0 at ≤ 50 ms median, 0 at ≥ 1 s, log scale in between |
Each component earns its max points times the model's rate on it (right 60% of the time when acting = 21 of 35 points), and the points add up to the score. High-stakes questions count twice in every component except speed. The score is computed in the browser, so it updates when you change the threshold or confidence measure.
## API
The API needs no token. Requests are limited to 256 KB, 20 questions, 2,000 characters of question text and 500 characters per option; at most 4 decisions run at once (more get 429).
| Method | Path | Body / result |
|---|---|---|
| `GET` | `/api/health` | `{"ok": true}` (used by the container health check) |
| `GET` | `/api/status` | Load status, device and details for each model, `order` (comparison order) and `env.models_dir` |
| `GET` | `/api/demos` | All demos, or one group with `?group=guardrails` |
| `GET` | `/api/groups` | Group id, family, label, description and demo/question counts |
| `POST` | `/api/decide` | `{"state": str or object, "questions": {...}, "models": [keys from /api/status]}` (every model if omitted; an unknown key is 422) → per model: `answers` (normalised) and `ms` |
| `POST` | `/api/reload/{key}` | Retry loading a model; keys are the model folder names in lower case with `_` separators, plus `laya` |
Each normalised answer contains `choice`, `probs` (per option), `top_prob`, `entropy_conf`, `laya_conf`, plus `expected_level` for score questions and `p_true` for yes/no questions.
```powershell
$body = @{
state = "I was charged twice for March. Refund the duplicate or I cancel."
questions = @{
dept = @{ type = "choice"; instructions = "Which team?"; criteria = @{ billing = "payments"; tech = "bugs" } }
churn = @{ type = "noul"; instructions = "The customer threatens to leave" }
}
} | ConvertTo-Json -Depth 5
Invoke-RestMethod -Method Post -Uri http://localhost:9910/api/decide -ContentType "application/json" `
-Body $body | ConvertTo-Json -Depth 6
```
## This application is now more secure with the following updates
- Execution of untrusted model code from the models folder was identified and fixed: only known FalconDec modeling files are run.
- Unlimited request sizes were identified and fixed: bodies, questions and option text are capped.
- Unbounded queuing of decision requests was identified and fixed: excess requests get 429 instead of starving the server.
- A model-reload race was identified and fixed: a model can only load once at a time.
- Missing browser security headers were identified and fixed: strict Content Security Policy, no framing, no sniffing, no referrer, no caching of API responses.
- Public API documentation endpoints were identified and fixed: `/docs`, `/redoc` and `/openapi.json` are off.
- Silently ignored and duplicated model keys were identified and fixed: unknown keys are rejected and duplicates run once.
- Publishing the port on every network interface was identified and fixed: the lab listens on this computer only unless you choose otherwise.
- An unencoded model key in a request URL was identified and fixed.
- Unpinned dependencies were identified and fixed: every Python package, torch included, is installed at an exact version and the build fails if the set is inconsistent.
## Layout
```
DecisionLab/
├── Dockerfile
├── docker-compose.yml
├── constraints.txt # every other package, exact versions (from pip freeze)
├── .env.example # copy to .env for local settings (.env is never in the image)
├── app.py # Hugging Face Space entry point (Gradio SDK, ZeroGPU)
├── requirements.txt # Space dependencies
├── requirements-container.txt # container dependencies (with constraints.txt)
├── models/ # your LightDec and Arthur models, one folder each (not in the image)
├── app/
│ ├── main.py # FastAPI app: UI + API on :9910
│ ├── models.py # LightDec (local folder or Hub) and Laya backends, timing
│ ├── registry.py # finds the models in MODELS_DIR, checks their modeling code (no torch)
│ ├── arthur.py # Arthur network and inference, verbatim from arthur_v0_8_0.ipynb
│ ├── arthur_io.py # Arthur input layout and calibration (no torch)
│ ├── scoring.py # normalises every model answer into one comparable record (no torch)
│ ├── security.py # body limit and security headers (plain ASGI)
│ ├── validation.py # checks /api/decide question sets, sizes and model keys (no web framework)
│ ├── demos.py # groups, the 47 demos (24 scored + 23 nano), reference answers and high-stakes flags
│ ├── demos_nano.json # Agent security: states and questions only
│ └── static/ # index.html, app.css, app.js, logo.png
└── tests/ # unittest suite, copied into the image
```
## Tests
The suite uses Python's built-in `unittest`, so it needs no extra packages. It checks the demo set (unique ids, known groups, valid reference answers, stakes, every demo passes request validation), request validation, and answer normalisation. `tests/test_api.py` also checks that a bad `/api/decide` request returns 422; it needs the app's real imports (fastapi, pydantic, torch), so outside the container it is skipped with the reason printed.
After `docker compose up -d --build`, in PowerShell:
```powershell
docker exec DecisionLab python -m unittest discover -s tests -v
```
The last line should read `OK` with no skips. Tests do not load the models and do not call Hugging Face.
## Notes
- Models run one after the other under a lock, so timings never compete for the device. Run one uvicorn worker per container; each worker would hold its own copy of every model.
- LightDec is loaded with the `falcondec_modeling.py` shipped in its repo; Laya with the `laya` package (`laya.load`).
- Laya is Apache-2.0. The LightDec checkpoints tested here state apache-2.0 in their model cards; check any other model you add. Laya is by Convai Innovations; this lab is not affiliated with Convai Innovations, TypeSafe AI or the layaForWeb project.
|