tacet-2b-decision(双语 README / Bilingual README)

License / 许可:apache-2.0(代码与模型卡)· 数据:见「训练数据 / Training data」

tacet(拉丁语乐谱记号:该声部此处静默)——一个 2B 的 System One 决策模型。给它一个 state(工单 / 邮件 / 票据 / 检索文档 / 游戏局面)和一组带类型的 typed questions(choice / noul / score),它在单次前向内输出每个候选上的概率分布。它不生成文本——没有可解析、可幻觉的内容。

tacet (Latin musical notation: "this voice is silent here") — a 2B System One decision model. Given a state (a ticket, an email, an invoice, retrieval documents, a game position) and a set of typed questions (choice / noul / score), it outputs a probability distribution over the candidates in a single forward pass. It does not generate text — there is nothing parseable to hallucinate.

English

What this model is

  • Output is a distribution, not text: one forward pass answers all questions at once (~150 ms level on a consumer GPU);
  • The answer space is defined at request time: candidate sets, K, and criteria all live in the request — changing the schema requires no retraining;
  • Calibrated and abstaining: per-candidate calibrated probabilities (with temperature calibration) and an "no candidate satisfies the requirement / insufficient information" abstention signal.

Usage

One-liner: the prompt is the state + typed questions rendered in the block_instr layout, ending with Answer:; the model emits a single letter at the first generated position; the readout is a softmax over the logits of the 26 letter tokens at that position, within that question's K candidates. The whole call is a deterministic single forward pass — do not sample; temperature is meaningless.

Prompt construction (block_instr layout)

[STATE]
{state text}

[QUESTIONS]
[Q1] [TASK] Which department should handle this?
[CRITERIA]
[A] billing: invoices, payments, refunds
[B] technical: bugs, outages, system errors
[C] other: everything else
[OUTPUT] Give a probability for every option; the probabilities must sum to 1.
...
[OPTIONS] A B C <pad to 26>
Answer:

Key points: ① wrap the above in the base model's chat template as a single user message (the letter channel only works on the chat base); ② [OPTIONS] lists the letter slots bound to this question's candidates (padded to 26); ③ the trailing Answer: is the generation cue; ④ several questions can be sent in one request (each with its own [OPTIONS]) and are all answered by the same forward pass.

Option A: llama.cpp (GGUF, parity-verified)

llama-server -m tacet-2b-decision-Q4_K_M.gguf -c 65536 -ngl 99 --host 0.0.0.0 --port 8080
curl http://127.0.0.1:8080/completion -d '{
  "prompt": "<chat-templated prompt as above, ending with Answer:>",
  "n_predict": 1, "temperature": 0.0, "n_probs": 64
}'

The chat template is applied client-side (same as our internal path — see the Python example below). From the response's top-logprob list at the first position: for each letter slot take the larger logit of the two tokens (bare letter / leading-space variant), then softmax over that question's K candidates (in-question renormalization; letters not returned by the server count as near-zero probability). We verified GGUF-vs-original parity: 4/5 agreement on a 5-question spot check (the single miss was a near-tie, 0.407/0.593).

Option B: transformers (exact logits)

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("zyang94/tacet-2b-decision")
model = AutoModelForCausalLM.from_pretrained("zyang94/tacet-2b-decision", dtype="auto").eval()

messages = [{"role": "user", "content": prompt}]   # prompt as above, ending with Answer:
ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt")
with torch.no_grad():
    logits = model(ids).logits[0, -1]              # logits of the FIRST generated position

def letter_logit(c):
    v = tok.encode(c, add_special_tokens=False)
    s = tok.encode(" " + c, add_special_tokens=False)
    assert len(v) == len(s) == 1                   # letters must be single tokens — readout assumption
    return max(logits[v[0]].item(), logits[s[0]].item())

K = 3                                              # number of candidates in THIS question
probs = torch.softmax(torch.tensor([letter_logit(c) for c in "ABCDEFGHIJKLMNOPQRSTUVWXYZ"[:K]]), 0)

Readout rules: ① take the logits at the first generated position (this is not sampling); ② per letter, take the max logit of the two variants (bare / leading space); ③ softmax only over that question's K letters (not over all 26, and never over the whole vocabulary); ④ letters must be single tokens (they are for this model, but validate if you swap tokenizer/vocab). Finer protocol details and pitfalls: TRAINING_GUIDE.md §5.1/§8.

Python requests examples

llama.cpp path (client-side chat template + /completion readout):

import requests
from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("zyang94/tacet-2b-decision")
body = render_block_instr(state, question, candidates)   # block_instr text, ending with "Answer:"
prompt = tok.apply_chat_template(
    [{"role": "user", "content": body}],
    tokenize=False, add_generation_prompt=True,
)
r = requests.post("http://127.0.0.1:8080/completion", json={
    "prompt": prompt, "n_predict": 1, "temperature": 0.0, "n_probs": 64,
}, timeout=30)
top = r.json()["completion_probabilities"][0]["probs"]   # [{"tok_str": " A", "logprob": ...}, ...]
# → aggregate to the K candidates per the readout rules above (tok_str may be "A" or " A"; take the max of the two variants)

Jev-compatible wire API (ships with tacet_serve.py, see "Option C" below):

import requests

r = requests.post("http://127.0.0.1:8231/v1/systemone", json={
    "state": "Customer #4127 wrote: the invoice total is wrong, we were charged twice.",
    "model": "tacet-letter",
    "questions": {
        "routing": {
            "type": "choice",
            "instructions": "Which department should handle this?",
            "criteria": {
                "billing": "invoices, payments, refunds",
                "technical": "bugs, outages, system errors",
                "other": "everything else",
            },
        }
    },
}, timeout=60)
ans = r.json()["answers"]["routing"]
print(ans["choice"])          # argmax label (ties broken by earliest label order)
print(ans["probabilities"])   # {"billing": p, ...} keys exactly equal the label set, sums to 1

For noul questions: "criteria": {"false": "…", "true": "…"} → {"type": "noul", "noul": <numeric p(yes)>}; for score questions: criteria is a list → {"type": "score", "probabilities": {"0": p, "1": p, ...}}. state may be a string or any dict (dicts are JSON-serialized into the prompt).

Option C: Jev-compatible wire API (ships with tacet_serve.py)

A single-file, stdlib-only compatible service exposing the Jev-compatible wire (the typed-decision protocol: noul / choice / score): clients neither build prompts nor read logits themselves — send JSON, get distributions back. The default block_instr rendering and server-side readout are verbatim-consistent.

# protocol integration (no torch needed; uniform distribution)
python3 tacet_serve.py --backend dummy --port 8231
# production (transformers load, defaults to zyang94/tacet-2b-decision)
python3 tacet_serve.py --backend letter --layout block_instr --max-len 6144 --port 8231
Endpoint Purpose
POST /v1/systemone main endpoint: multiple questions per request, all answered in one forward pass
GET /health 200 {"ok": true} (orchestration readiness check)

Request/response shapes per question type (questions is a dict with arbitrary keys; multiple questions per request):

Type criteria sent Response
noul {"false": desc, "true": desc} {"type": "noul", "noul": <numeric p(yes)>}
choice {label: desc, ...} (label set = criteria keys) {"type": "choice", "choice": <argmax label>, "probabilities": {<label>: p, ...}}
score [tier desc, ...] (labels "0".."n-1") {"type": "score", "probabilities": {"0": p, ...}}

The server enforces three hard constraints (verbatim from the compatible ecosystem): ① the noul value must be a number (bools are rejected); ② probabilities keys must exactly equal the label set (no more, no fewer), values ∈ [0,1], sum = 1 (strict tolerance 1e-3); ③ the wire request does not send labels — choice labels = criteria keys, score = criteria indices, noul is fixed to ["no", "yes"].

Evaluation

On a public typed-decision evaluation set of 231 questions (fully independent of the training distribution):

Metric Value
Top-1 accuracy (argmax) 0.7273
Per question type: choice / noul / score —
ECE (10 bins, pre-calibration) 0.14

Abstention (sentinel-candidate mechanism): on a 261-question test set containing "no correct answer" labels, no-correct-answer recall / false-alarm-on-solvable curves are reported under a threshold sweep (this release does not include an abstention slot; abstention ships in later versions of the model family).

Honest limits

  • Single forward pass, no chain of thought: the model is trained to output a distribution in one forward pass — it is limited on questions requiring multi-step serial computation (long policy reasoning, multi-hop, date arithmetic). This is a form boundary, not undertraining.
  • 26-option cap: at most 26 candidates per choice question (+1 abstention sentinel). For larger option spaces, use a two-level / factorized decomposition.
  • Context window 6144 tokens: states beyond this are truncated.
  • Language: English (the base model's multilingual ability is not aligned in this version).
  • Probability calibration: the shape of the letter-readout distribution is the model's own output; re-measure ECE on your own data before using it for threshold decisions.

Training data

Training data is 100% programmatically generated with verifiable ground truth — zero human annotation, zero LLM-teacher outputs: game domains (maze/snake BFS ground truth), retrieval domain (teacher scores over public upstream Apache-2.0 data), synthetic UI (deterministic behavior ground truth), judge/routing (option-count histogram mirroring), long text (long policies / tickets). No conversation-model outputs were used as training targets.

Hardware & authorship

  • Training hardware: 2× NVIDIA GeForce RTX 2080 Ti 22GB (consumer GPUs) — the full pipeline (LoRA pretraining, weight averaging, quantization verification) ran on these two cards; inference is ~150 ms level per card (10 candidates).
  • Authorship: GLM-5.3-flash (Zhipu AI) is a co-first author of this project, responsible for training-recipe iteration, experiment diagnosis and engineering implementation; Zhengxing Yang is the corresponding author, responsible for direction and experiment decisions. The work is human–AI collaborative with roughly equal effort; judgment and direction rest with the human.

Citation

@misc{tacet2026,
  title  = {tacet-2b-decision: a 2B System One decision model with calibrated, abstaining outputs},
  author = {GLM-5.3-flash (co-first author, engineering and recipe iteration) and Zhengxing Yang (corresponding author, direction and decisions)},
  year   = {2026},
  url    = {https://huggingface.co/zyang94/tacet-2b-decision}
}

中文

这是什么

  • 输出是分布,不是文字:同一次前向回答所有问题(~150ms 级,消费级 GPU);
  • 答案空间在请求时定义:候选集 / K 值 / 判据全部写在请求里,换 schema 无需重新训练;
  • 可校准 + 可弃权:支持给每个候选输出校准后的概率(含温度校准),以及「没有候选满足要求 / 信息不足」的弃权信号。

使用 / How to call

原理一句话:prompt = state + typed questions(block_instr 布局,末尾 Answer: 收尾);模型在首个生成位上输出一个字母;读出 = 取该位置上 26 个字母 token 的 logits,对该题的候选数 K 做题内 softmax。整个调用是确定性单次前向——不要采样,温度无意义。

One-liner: the prompt is the state + typed questions rendered in the block_instr layout, ending with Answer:; the model emits a single letter at the first generated position; the readout is a softmax over the logits of the 26 letter tokens at that position, within that question's K candidates. The whole call is a deterministic single forward pass — do not sample; temperature is meaningless.

Prompt 构造(block_instr 布局)

[STATE]
{state text}

[QUESTIONS]
[Q1] [TASK] Which department should handle this?
[CRITERIA]
[A] billing: invoices, payments, refunds
[B] technical: bugs, outages, system errors
[C] other: everything else
[OUTPUT] Give a probability for every option; the probabilities must sum to 1.
...
[OPTIONS] A B C <pad to 26>
Answer:

要点:① 用底座的 chat 模板以单条 user 消息包裹上述内容(字母通道只在 chat 基座上通);② [OPTIONS] 列出该题候选绑定的字母槽(补齐到 26);③ 末尾的 Answer: 是生成提示;④ 多个问题可写在一次请求里(每题自己的 [OPTIONS]),同一次前向全部作答。

方式 A:llama.cpp(GGUF,实测对拍)

llama-server -m tacet-2b-decision-Q4_K_M.gguf -c 65536 -ngl 99 --host 0.0.0.0 --port 8080
curl http://127.0.0.1:8080/completion -d '{
  "prompt": "<chat-templated prompt as above, ending with Answer:>",
  "n_predict": 1, "temperature": 0.0, "n_probs": 64
}'

chat 模板在客户端套(与我们的内部路径一致,见下方 Python 示例)。从响应的 top-logprobs 列表取首个位置:对每个字母槽取「裸字母 / 带前置空格」两个 token 中 logit 较大者,再在该题 K 个候选上 softmax(题内重归一化;未上榜的字母视为极小值)。我们实测 GGUF 与原模型对拍 5 题中 4/5 一致(唯一 miss 是 0.407/0.593 的近似平局)。

方式 B:transformers(精确 logits)

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("zyang94/tacet-2b-decision")
model = AutoModelForCausalLM.from_pretrained("zyang94/tacet-2b-decision", dtype="auto").eval()

messages = [{"role": "user", "content": prompt}]   # prompt as above, ending with Answer:
ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt")
with torch.no_grad():
    logits = model(ids).logits[0, -1]              # logits of the FIRST generated position

def letter_logit(c):
    v = tok.encode(c, add_special_tokens=False)
    s = tok.encode(" " + c, add_special_tokens=False)
    assert len(v) == len(s) == 1                   # letters must be single tokens — readout assumption
    return max(logits[v[0]].item(), logits[s[0]].item())

K = 3                                              # number of candidates in THIS question
probs = torch.softmax(torch.tensor([letter_logit(c) for c in "ABCDEFGHIJKLMNOPQRSTUVWXYZ"[:K]]), 0)

读出铁律:① 取的是首个生成位的 logits(不是采样);② 每个字母取「裸 / 空格前缀」两个 token 变体的最大 logit;③ softmax 只在该题的 K 个字母上做(不要对全 26 个、更不要对全词表归一化);④ 字母必须是单 token(本模型全为单 token,但自定义字表时须校验)。更细的协议与坑见 TRAINING_GUIDE.md §5.1/§8。

Python requests 示例

llama.cpp 路径(客户端套模板 + /completion 读出):

import requests
from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("zyang94/tacet-2b-decision")
body = render_block_instr(state, question, candidates)   # block_instr 文本,末尾 Answer:
prompt = tok.apply_chat_template(
    [{"role": "user", "content": body}],
    tokenize=False, add_generation_prompt=True,
)
r = requests.post("http://127.0.0.1:8080/completion", json={
    "prompt": prompt, "n_predict": 1, "temperature": 0.0, "n_probs": 64,
}, timeout=30)
top = r.json()["completion_probabilities"][0]["probs"]   # [{"tok_str": " A", "logprob": ...}, ...]
# → 按上方读出铁律聚合到 K 个候选(tok_str 可能是 "A" 或 " A",两个变体取 max)

Jev-compatible wire API(随附 tacet_serve.py,见下方「方式 C」):

import requests

r = requests.post("http://127.0.0.1:8231/v1/systemone", json={
    "state": "Customer #4127 wrote: the invoice total is wrong, we were charged twice.",
    "model": "tacet-letter",
    "questions": {
        "routing": {
            "type": "choice",
            "instructions": "Which department should handle this?",
            "criteria": {
                "billing": "invoices, payments, refunds",
                "technical": "bugs, outages, system errors",
                "other": "everything else",
            },
        }
    },
}, timeout=60)
ans = r.json()["answers"]["routing"]
print(ans["choice"])          # argmax 标签(并列时取标签顺序在前的)
print(ans["probabilities"])   # {"billing": p, ...} 键恰等于标签集,和=1

noul 题:"criteria": {"false": "…", "true": "…"} → 应答 {"type": "noul", "noul": <p_yes 数字>};score 题:criteria 为列表 → 应答 {"type": "score", "probabilities": {"0": p, "1": p, ...}}。state 可以是字符串或任意 dict(dict 会被 JSON 序列化后进 prompt)。

方式 C:Jev-compatible wire API(随附 tacet_serve.py)

单文件、纯 stdlib 的兼容服务:暴露 Jev-compatible wire(typed-decision 协议:noul / choice / score 三题型),客户端不用自己拼 prompt、不用自己读 logits——发 JSON 拿分布。默认 block_instr 口径与服务端读出逐字同源。

# 协议联调(不需要 torch,均匀分布)
python3 tacet_serve.py --backend dummy --port 8231
# 生产(transformers 加载,默认 zyang94/tacet-2b-decision)
python3 tacet_serve.py --backend letter --layout block_instr --max-len 6144 --port 8231
端点 说明
POST /v1/systemone 主端点:一次请求多道题,同一次前向全部作答
GET /health 200 {"ok": true}(编排脚本等就绪用)

三种题型的请求与应答形状(questions 为任意键名的 dict,一次可带多题):

题型 criteria 传入 应答
noul {"false": 描述, "true": 描述} {"type": "noul", "noul": <数字 p(yes)>}
choice {标签: 描述, ...}(标签集 = criteria 的键) {"type": "choice", "choice": <argmax 标签>, "probabilities": {<标签>: p, ...}}
score [档位描述, ...](标签为 "0".."n-1") {"type": "score", "probabilities": {"0": p, ...}}

服务端三条硬约束(与兼容生态逐字一致):① noul 的值必须是数字(bool 拒收);② probabilities 的键恰好等于标签集(不多不少)、值 ∈ [0,1]、和 = 1(严格容差 1e-3);③ wire 请求**不送 labels**——choice 标签集 = criteria 的键、score = criteria 下标、noul 固定 ["no", "yes"]。

评测

在 231 题的公开 typed-decision 评测集上(与训练分布完全独立):

指标 数值
top-1 准确率(argmax 口径) 0.7273
按题型:choice / noul / score —
ECE(10 bin,校准前) 0.14

弃权能力(哨兵候选机制):在一组含「无正解」标注的 261 题测试上,无正解召回 / 有正解误报两指标随阈值扫描报告(该发布版本未包含弃权槽,弃权见模型家族后续版本)。

诚实声明

  • 单次前向,无思考链:本模型被训练为单 forward 直接输出分布——它在需要多步串行计算的题目上(长政策推理、多跳、日期运算)能力有限,这是形态边界而非训练不足。
  • 选项上限 26:choice 题最多 26 个候选(+1 个弃权哨兵)。更大的选项空间建议两级/因子化分解。
  • 上下文 6144 token:超过的 state 会被截断。
  • 语言:英文(底座多语言能力未在本版本对齐)。
  • 概率校准:letter 读出的分布形状为模型输出;建议在自己的数据上重测 ECE 后再用于阈值决策。

训练数据

训练数据 100% 程序化生成,带可验证真值,零人工标注、零 LLM 教师输出:游戏域(迷宫/蛇 BFS 真值)、检索域(公开上游 Apache-2.0 数据的 teacher 分数)、合成 UI(确定性行为真值)、judge/routing(选项数直方图镜像)、长文本(长政策/工单)。未使用任何对话模型输出作为训练目标。

训练硬件与作者栏

  • 训练硬件:2× NVIDIA GeForce RTX 2080 Ti 22GB(消费级显卡)——全套训练(预训练 LoRA、权重平均、量化校验)在这两张卡上完成;推理单卡 ~150ms 级(10 候选)。
  • 作者与致谢:GLM-5.3-flash(智谱 AI)为本项目共同一作(co-first author),负责训练配方迭代、实验诊断与工程实现;Zhengxing Yang 为通讯作者,负责方向决策与实验取舍。本文由人机协作完成,工作量大致对半,判断与方向归人。

引用

@misc{tacet2026,
  title  = {tacet-2b-decision: a 2B System One decision model with calibrated, abstaining outputs},
  author = {GLM-5.3-flash (co-first author, engineering and recipe iteration) and Zhengxing Yang (corresponding author, direction and decisions)},
  year   = {2026},
  url    = {https://huggingface.co/zyang94/tacet-2b-decision}
}

仓库内容 / What's in this repo

  • model.safetensors — F16 权重 / F16 weights
  • tacet-2b-decision-Q4_K_M.gguf — Q4_K_M 量化(llama.cpp)/ Q4_K_M quantization for llama.cpp
  • README.md — 本模型卡 / this model card
  • TRAINING_GUIDE.md / TRAINING_GUIDE_EN.md — 从零复现训练方法指引(中/英;只讲方法,不含代码)/ from-scratch training method guides (ZH/EN; method only, no code)
  • tacet_serve.py — Jev-compatible wire 推理服务(单文件纯 stdlib;POST /v1/systemone,noul/choice/score) / single-file Jev-compatible wire inference service (POST /v1/systemone, noul/choice/score)
  • 本仓库不含训练代码、数据生成器与数据本体;复现路径见两份指南 §4–§7 / this repo ships no training code, generators, or data; see §4–§7 of the guides for the reproduction path

所有数字均为我们自己测得的绝对数字,本卡不包含任何第三方对比。 All numbers above are our own absolute measurements; this card contains no third-party comparisons.

Downloads last month
49
Safetensors
Model size
2B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for zyang94/tacet-2b-decision

Finetuned
Qwen/Qwen3.5-2B
Quantized
(233)
this model