decisor-4b
Under Labs · Code and examples · Release announcement
decisor-4b — quero-quero — Technical Preview (v0.1.0).
We are developing a family of language models with a focus on Brazilian Portuguese and Brazil's legal, tax and fiscal contexts. decisor-4b is the first technical preview in this effort, built on the Qwen3.5-4B checkpoint.
Read this first. This is a technical preview (v0.1.0) for experimentation and integration, not a production-ready deployment. This release distributes the FP8 model for SGLang only, served with the patched SGLang image from the decisor repository. Option probabilities aren't calibrated estimates of correctness. Legal and tax outputs require review by a qualified professional against current authoritative sources.
The model was trained for typed decisions in English and Brazilian Portuguese, with additional emphasis on Brazilian legal and tax contexts. Training data and the training recipe aren't released. Training was designed to preserve instruction following while specializing the model for typed decisions. We evaluated free-form generation on the BF16 reference checkpoint with a small automated battery; we haven't established broader instruction-following quality or tested generation on FP8.
The model supports two workloads:
- Typed decisions: picks one of the supplied options for a state and question, and returns a probability distribution over all of them, using option scores from a single prefill pass.
- Instruction following: responds to instructions and generates text. The GitHub SDK and CLI in this release focus on typed decisions.
| Focus area | Scope |
|---|---|
| General-purpose decisions | Classification and selection across varied tasks, including English-language tasks |
| Brazilian Portuguese | Additional training emphasis on Brazilian Portuguese language and context |
| Brazilian legal and tax contexts | Additional domain-focused training for Brazilian legal, tax and fiscal content |
The focus areas overlap. They describe training emphasis on a single model, not separate models, heads or routing modes. Support for other languages comes from Qwen3.5-4B, and performance varies by language, domain and task.
Its nickname, quero-quero, comes from the southern lapwing, the state bird of Rio Grande do Sul, Brazil.
Code and integration
decisor on GitHub provides a minimal Python SDK, the patched SGLang image, and a CLI demo with examples in English and Brazilian Portuguese.
Follow the GitHub quickstart to download the engine image, start SGLang, and try the demo.
How it works
state + question + options → prefill → option scores → decision + probabilities
The SDK formats the request using the model's decision prompt format and reads option scores from the next-token logits produced during prefill. The current engine request emits one token, but the SDK uses the scores — not the generated text — to select an option.
Successful calls return one of the supplied option IDs and probabilities normalized over the supplied options. If your task allows an "insufficient information" outcome, include it among the options and state when it should apply.
Decision prompt format
The SDK renders each request into the model's decision prompt format. You can also use it directly, without the SDK. Rendered example:
You make decisions with exactly one letter. Read the state, apply the question, and pick from the listed options.
Semantics: a field absent from the state is unknown — it is neither true nor false. Use the insufficient-information option, when listed, only if a missing fact would change the decision; otherwise decide with the facts given.
STATE (JSON):
{"policy": "Refunds are allowed for unused items within 30 days of delivery.", "request": {"days_since_delivery": 12, "item_unused": true}}
Q: Is the refund allowed under this policy?
Options: (A) Refund allowed (B) Refund not allowed (C) Not enough information
A:
Options are assigned letters in order — (A), (B), (C) — and
the prompt ends with the answer marker \nA: (not option A).
The SDK sends the rendered prompt to SGLang's /generate endpoint as
a raw completion. It doesn't apply the chat template or produce a
thinking trace on this path.
For each option letter, the SDK checks three candidate forms: A,
" A" (with a leading space) and (A, retaining only valid
single-token continuations in the prompt context. It combines the
logprobs of eligible forms for each option with logsumexp, then
normalizes across the supplied options.
Using this prompt directly requires reproducing that token selection and scoring procedure. Sampling a letter doesn't reproduce the SDK's aggregated decision or probability distribution. The implementation is available in the decisor SDK.
Model details
| Field | Value |
|---|---|
| Model ID | decisor-4b |
| Version | 0.1.0 — technical preview |
| Publisher | Under Labs |
| Backbone | Fine-tuned from Qwen/Qwen3.5-4B, the original post-trained checkpoint |
| Format | compressed-tensors, FP8 dynamic (per-channel weights, per-token activations) — 128 tensors in FP8 (96 MLP + 32 full-attention); 427 tensors kept in BF16 (including Gated DeltaNet (GDN) layers, norms and embeddings). "FP8" doesn't mean 8-bit everywhere. |
| Weight size | 7.13 GB (7,125,997,144 bytes) |
| Evaluation context | We used an 8,192-token context for the FP8 evaluations below. |
Distributed format: FP8 only. This repository contains the model
weights, configuration, tokenizer and chat template. The v0.1.0
weights won't change; any change to the weights ships as a new
version. Pin revision="v0.1.0" to stay on this release.
Intended use
decisor-4b is intended for experimentation with classification and selection tasks where the application supplies the decision criteria and candidate outcomes.
Don't use this preview as the sole basis for consequential decisions about people, including credit, employment, benefits or sanctions. Legal and tax applications require review by a qualified professional and verification against current authoritative sources.
Usage
The supported deployment path for this release is the patched SGLang v0.5.20 image provided by the decisor repository.
The stock SGLang v0.5.20 loader doesn't correctly load this fused
checkpoint; the patched SGLang image includes the required loader fix
and the logprobs guard, which keeps the scheduler from crashing when
a token_ids_logprob request shares a batch with one that doesn't
ask for logprobs. Use it rather than the stock SGLang snippet from
Hugging Face's Use this model button. The published image is
underlabsai/decisor-sglang:0.5.20, digest
sha256:60aee212ef25b0d303213b1c144b32f01403191b923d93d09a610770f8a6ad60.
| GPU | Status |
|---|---|
| NVIDIA RTX 5090 (32 GB) | functional checks (demo and example requests) |
| NVIDIA RTX PRO 6000 Blackwell Server Edition (96 GB) | all reported measurements |
We haven't tested vLLM or direct Transformers inference for this release.
On first startup, the engine downloads the model weights from this Hugging Face repository and caches them for later runs.
Python
With the SDK installed from the repository's ./sdks/python
directory, the engine running, and SGLANG_API_KEY available in your
shell (standalone Python doesn't load .env automatically):
import os
from decisor import decide
result = decide(
"http://127.0.0.1:8768",
api_key=os.environ["SGLANG_API_KEY"],
state={
"policy": "Refunds are allowed for unused items within 30 days of delivery.",
"request": {"days_since_delivery": 12, "item_unused": True},
},
question="Is the refund allowed under this policy?",
options=[
{"id": "approved", "text": "Refund allowed"},
{"id": "rejected", "text": "Refund not allowed"},
{"id": "insufficient", "text": "Not enough information"},
],
)
print(result["decision"])
Example result object (probabilities rounded for display):
{
"decision": "approved",
"options": [
{"id": "approved", "prob": 1.0},
{"id": "rejected", "prob": 0.0},
{"id": "insufficient", "prob": 0.0}
]
}
Limits: the SDK accepts 2 to 18 options per request and rejects more; ties resolve to the first option. This is an SDK limit. We've evaluated FP8 with up to 10 options. Failures raise an exception (connection, HTTP status or timeout).
Evaluation
We ran all evaluations ourselves, under the conditions stated with each result. The BF16 reference checkpoint is our fine-tuned decisor model before FP8 quantization; we keep it internally and don't distribute it. Each result names the checkpoint it was measured on. We aren't releasing the exact adapted items and internal evaluation tools for tc193, the development set, the 1,621-question evaluation or the generation battery with this preview. Links to the source datasets provide context, but aren't enough to reproduce those evaluations exactly.
We start with decision quality on the FP8 model distributed here, followed by BF16 reference checkpoint results and FP8 throughput.
tc193
tc193 is our internal 193-item adaptation of LegalBench BR, licensed under CC BY-SA 4.0, used here for evaluation under our decision protocol (10 options per item, in a fixed order, scored by option letter). The adapted items aren't redistributed.
Seven items include the gold category name in the case text, which
may give away the answer. We kept all seven in the reported score;
one of them, lbbr-632, is still under review for exclusion. This is
a problem with the test items, not evidence of train/test
contamination, but it limits how much tc193 can tell you.
| Model | tc193 score | Runtime and conditions |
|---|---|---|
| decisor-4b FP8 (this release) | 142/193 (73.58%) | patched SGLang image (same patches as the image we distribute), NVIDIA RTX PRO 6000 Blackwell Server Edition (96 GB), concurrency 1 |
| decisor-4b BF16 reference checkpoint | 143/193 (74.09%) | same image, GPU, engine flags and concurrency as the FP8 row |
Qwen3.5-4B (commit 851bf6e8) |
74/193 (38.34%) | in-process Transformers harness, BF16, NVIDIA RTX PRO 6000 Blackwell Server Edition (96 GB), no thinking |
| Most-frequent-category baseline | 53/193 (27.46%) | always selects "Direito Tributário"; calculated from the answer key |
All three models answered the same 193 items, each with the same 10 options in the same order, using the same prompts and letter-based scoring. The runtime differs between the Qwen and decisor rows.
We trained decisor on this decision prompt format. The comparison shows how each model performs with that format. It doesn't show how Qwen would do with a prompt suited to it, and it doesn't separate the effect of training from the runtime difference.
Item by item, the BF16 reference checkpoint got 76 items right that Qwen missed and missed 7 that Qwen got right (net +69). For FP8, the counts are 75 and 7 (net +68). Between BF16 and FP8, three decisions changed: two from right to wrong and one from wrong to right.
"Direito Tributário" is the most frequent answer (53 of 193) and always appears as option J. Always choosing option A would score 33/193 (17.10%).
BF16–FP8 decision agreement
We compared the stored predictions from the BF16 and FP8 runs.
| Evaluation set | Same answer | Different answer |
|---|---|---|
| tc193 | 190/193 | 3/193 |
| Internal development set (320 questions) | 316/320 | 4/320 |
Within each set, both checkpoints ran on the same patched SGLang image, engine flags and concurrency. These counts compare the selected options without using the answer keys. Agreement shows how closely FP8 preserved the BF16 decisions; both checkpoints can agree on an incorrect answer.
BF16 reference checkpoint results
We ran the evaluations in this table on the BF16 reference checkpoint, except the Qwen3.5-4B row. "Adapted" means we used a subset or our own decision protocol; those rows aren't official full-benchmark scores.
| Benchmark / dataset | Correct | Accuracy | Evaluation setup |
|---|---|---|---|
| JevBench-231 | 177/231 | 76.62% | JevBench harness v1.4.1, frozen v1.2 protocol |
| Gevva0-650 | 649/650 | 99.85% | Public suite, revision 845e11cd |
| BFCL (adapted) | 127/130 | 97.69% | Adapted |
| iSarcasmEval | 126/135 | 93.33% | Adapted |
| FinEntity | 80/92 | 86.96% | Adapted |
| SATA-Bench | 130/154 | 84.42% | Adapted |
| NLI4CT | 39/50 | 78.00% | Adapted |
| BPoMP | 262/340 | 77.06% | Adapted |
| WinoGrande | 37/50 | 74.00% | Adapted |
| GSM8K (adapted) | 72/100 | 72.00% | Adapted |
| HellaSwag | 35/50 | 70.00% | Adapted |
| CLadder | 35/50 | 70.00% | Adapted |
| MuSR | 32/50 | 64.00% | Adapted |
| cfcolor | 32/50 | 64.00% | Adapted |
| Amazon ESCI | 32/50 | 64.00% | Adapted |
| ANLI | 29/50 | 58.00% | Adapted |
| Humicroedit | 29/50 | 58.00% | Adapted |
| VAST | 26/50 | 52.00% | Adapted |
| CRUXEval (adapted) | 25/50 | 50.00% | Adapted |
| GPQA Diamond | 23/50 | 46.00% | Adapted |
| ARC-Challenge | 14/15 | 93.33% | Adapted |
| ARC-Easy | 13/15 | 86.67% | Adapted |
| MMLU | 12/15 | 80.00% | Adapted |
| SGD/SGD-X | 11/15 | 73.33% | Adapted |
| SimpleBench | 0/10 | 0.00% | Adapted |
| Total — 23 adapted datasets | 1,221/1,621 | 75.32% | Adapted; up to 18 options per question |
| Qwen3.5-4B on the same 1,621 questions | 926/1,621 | 57.13% | Same questions, harness and protocol as the total above |
JevBench conditions: our decision prompt format and option scoring, thinking disabled, serial, one pass, no retries; reproduced twice with identical results. Self-measured — not a leaderboard entry.
Gevva0: ECE 0.0027, Brier 0.0036, OOD 120/120. The suite has no official leaderboard, and 649/650 is near its ceiling. The ECE and Brier scores describe calibration on this suite for the BF16 reference checkpoint; they don't establish calibration for the FP8 model or other tasks and datasets.
The 23 adapted datasets ran through the same in-process Transformers harness for both models, with our decision prompt format and scoring, without thinking. We trained decisor on this format and used the same format to evaluate Qwen3.5-4B. We haven't tested Qwen with a prompt optimized for it in these comparisons. The total gives each question equal weight, so larger subsets contribute more. We haven't evaluated the distributed FP8 model on this combined set.
We evaluated JevBench-231 and Gevva0-650 only on the BF16 reference checkpoint. The evaluation route we used couldn't load the FP8 model; we haven't rerun them through SGLang.
BFCL adapted measures selection under our decision protocol, not successful execution of real tools. The ARC, MMLU and SGD/SGD-X subsets contain only 15 questions each and SimpleBench contains 10 — don't generalize these scores to the full benchmarks.
Free-form generation
We evaluated the BF16 reference checkpoint on an internal free-form generation battery: 72 prompts scored automatically on a 108-point scale, with objective and keyword-based checks and partial credit.
| Model | Points out of 108 | Runtime |
|---|---|---|
| decisor-4b BF16 reference checkpoint | 73.0 | Separate SGLang setup; memory and context settings weren't recorded |
| Qwen3.5-4B | 67.5 | SGLang; 32,768-token context |
Both runs used non-thinking mode with the same frozen generation parameters (temperature 0). decisor scored 5.5 points higher on this battery, but the runtime settings weren't fully matched. The scores aren't a count of correct responses, and separate writing-rubric and model-judge reviews aren't included.
We haven't run this evaluation on the distributed FP8 model.
Decision throughput — FP8 (SGLang)
We measured throughput with a dedicated evaluation harness under controlled prefix-cache conditions. The repository disables the radix cache by default, so these numbers don't describe the default setup. We used the FP8 model on an NVIDIA RTX PRO 6000 Blackwell Server Edition (96 GB) with a patched SGLang image carrying the same patches as the image we distribute. This isn't an end-to-end benchmark of the current Python SDK or CLI.
The workload was the 193 tc193 prompts at concurrency 1, with cache conditions verified through per-request token accounting. The cold result is one exploratory repetition with an initially empty prefix cache. For the warm result, we report the median of three repetitions with a prewarmed prefix cache, without restarting the engine between them.
| Condition (FP8) | Valid decisions/s |
|---|---|
| cold, concurrency 1 (exploratory, 1 rep) | 33.7 |
| warm, concurrency 1 (median of 3) | 49.5 |
Valid-decision throughput excludes requests whose option scores tied after rounding to four decimal places.
Cold and warm describe the cache state during measurement.
Measurement protocol
Across the four measurement windows, 772 requests completed without failure or timeout. Of these, 768 counted toward valid-decision throughput; four were excluded under the tie rule above.
The engine ran with --mem-fraction-static 0.35 --max-running-requests 16 --context-length 8192 for these windows.
Peak VRAM usage across these windows was 37,135–37,249 MiB. This is configuration-specific memory usage, not a minimum VRAM requirement.
Limitations
What we tested
- FP8 inference with the patched SGLang image on an NVIDIA RTX 5090 (32 GB) and an RTX PRO 6000 Blackwell Server Edition (96 GB).
- SDK and CLI integration with the running FP8 model, including four example requests in English and Brazilian Portuguese that produced their expected decisions (functional checks, not a quality benchmark).
- FP8 decision quality on the 193-item tc193 adaptation, subject to the evaluation caveats above.
- BF16–FP8 decision agreement on tc193 and a 320-question development set.
- FP8 throughput on the RTX PRO 6000 under the specific harness, workload and cache conditions reported above.
Evaluation boundaries
- We haven't tested smaller GPUs. Weight-file size alone doesn't show runtime compatibility or sufficient VRAM.
- We measured JevBench-231, Gevva0-650, the 23-dataset table and the generation battery on the BF16 reference checkpoint.
- We haven't repeated our earlier end-to-end testing of the BF16 setup in full for FP8. The current integration checks don't replace it.
- Our FP8 evaluations used 10 options per item in tc193 and four in an internal development set of 320 questions; we report agreement, not accuracy, for that set. We haven't evaluated FP8 with more options.
- We used an 8,192-token context for the FP8 evaluations reported here. These evaluations don't establish quality at longer contexts.
- The tools we used to create the FP8 weights included a dependency combination outside their declared support ranges. This concerns the conversion environment, which is separate from the patched SGLang image used to serve the model.
- We haven't run a dedicated tax evaluation. The tc193 results cover legal tasks only and don't show how the model performs on tax content.
- We compared Qwen3.5-4B with decisor on typed decisions and free-form generation. The decision evaluations use our decision prompt format; we trained decisor on this format and used the same format to evaluate Qwen3.5-4B. We haven't tested Qwen with a prompt optimized for it in these comparisons. The 1,621-question comparison used the same evaluation harness for both models; the tc193 and generation comparisons used different runtime setups, described alongside their results.
Behavioral limitations
Option probabilities aren't calibrated. The model can be confidently wrong, and its decisions can change with wording, language and how the input is laid out.
In small exploratory tests, changing the field order or the policy wording changed some decisions. In one case, instructions hidden in the input data flipped a correct decision to a wrong one. These tests weren't systematic, so they don't tell you how often this happens in practice.
Treat instructions inside untrusted input as a prompt-injection risk. Keeping your policy separate from user content makes prompts clearer, but the model won't always respect that boundary.
Write options that are clearly distinct, say when the insufficient-information option applies, and test the model on your own representative and adversarial cases. Keeping the field order fixed doesn't make the model insensitive to it.
In the C=1 serving benchmark, FP8 scored 142/192 in the initially empty-cache window and 144/192 in each of the three prewarmed windows. One tied-score request was excluded from each FP8 window. The BF16 reference checkpoint scored 143/193 in all four windows, with no ties excluded. These counts use the serving benchmark's tie-exclusion rule. They are separate from the 142/193 FP8 quality result above and aren't enough to establish a general accuracy advantage over BF16.
Generated responses can include fabricated facts, legal citations or tax guidance. Check them against current authoritative sources.
License and attribution
| Component | License / attribution |
|---|---|
| Model weights | Apache-2.0. See LICENSE and NOTICE. |
| Python SDK and integration code | Apache-2.0. See the decisor repository. |
| Backbone | Fine-tuned from Qwen3.5-4B, licensed under the Apache License, Version 2.0. |
The patched SGLang image includes a fused-checkpoint loader fix and the logprobs guard from SGLang PR #35052, with an in-build regression test from PR #35852. Upstream license files are preserved.
Copyright 2026 Under Serviços de Internet Ltda.
Developed by Under Labs. Provided as-is, without a commitment to user support.
Citation
@misc{decisor4b2026,
title = {decisor-4b},
author = {{Under Labs}},
year = {2026},
url = {https://huggingface.co/underlabs/decisor-4b},
note = {Technical preview, version 0.1.0}
}
For the SDK and serving integration:
@software{decisor2026,
title = {decisor: Python SDK and SGLang runtime for decisor-4b},
author = {{Under Labs}},
year = {2026},
url = {https://github.com/underlabs-ai/decisor},
note = {Model: https://huggingface.co/underlabs/decisor-4b}
}
Contact
For non-sensitive bug reports and reproducible model-behavior issues, use GitHub Issues. Do not include API keys, personal data or confidential inputs.
To report a security vulnerability, use GitHub private vulnerability reporting. Do not open a public issue.
- Downloads last month
- 207


