- claes preview
- Evaluation report
- Setup
- Danish test split (
ctx1024, 6,664 questions) - English regression check:
LocalLLaMA/typed-decisionstest (2,000 decisions) - MASSIVE intent (20 options, 500 utterances per language)
- Temperature fitting (validation, 5,537 questions)
- Do option descriptions help? (400 sampled test choice questions)
- What didn't work / limitations
- What to try next
- Setup
claes preview
The multilingual Laya decision model (mmBERT-base encoder, 322M
params) fine-tuned on syvai/danish-dynaword-laya (ctx1024 config): 337k typed questions over Danish Dynaword
texts. Laya scores each answer option of a choice/noul question at a [MASK] marker and returns calibrated
probabilities. It never generates text, so the output is always one of the options you supplied.
import laya
agent = laya.load("syvai/claes")
res = agent.predict(
"Forslaget om kontingentforhΓΈjelse blev forkastet med 41 stemmer mod 12.",
{"udfald": {"type": "choice", "instructions": "Udfaldet af afstemningen.",
"criteria": {"forkastet": "", "vedtaget": ""}},
"har_sted": {"type": "noul", "instructions": "Oplyser teksten fΓΈlgende: Stedet hvor mΓΈdet blev afholdt.",
"criteria": {"false": "nej", "true": "ja"}}})
print(res["answers"])
# {'udfald': {'type': 'choice', 'choice': 'forkastet', 'probabilities': {'forkastet': 0.64, 'vedtaget': 0.36}, ...},
# 'har_sted': {'type': 'noul', 'noul': 0.28, ...}}
Question phrasing that matches the training data works best:
- presence:
Oplyser teksten fΓΈlgende: <feltbeskrivelse>(noul, criteria{"false": "nej", "true": "ja"}) - boolean / enum field: the bare field description
- field on one item of a list:
Om "<record>" (<nΓΈglefelt>): <feltbeskrivelse>
Enum options can be bare keys; adding descriptions neither helps nor hurts after fine-tuning (see ablation below).
Files: model.safetensors, rl_agent_config.json (with max_len 1024, head_max_len 256 and the fitted
temperatures), tokenizer/, encoder/config.json, temperature_fit.json.
Evaluation report
Fine-tune of the multilingual Laya checkpoint (convaiinnovations/laya, subfolder multilingual/, mmBERT-base, 322M
params) on syvai/danish-dynaword-laya, config ctx1024 (states up to 1,024 tokens). The 16k default config was
not trained: the only available hardware was a Mac mini M4 with 16 GB, which runs out of memory at 8k tokens
(see "What didn't work").
All metrics below use the temperature-fitted config shipped in this repo.
Setup
| Hardware | Mac mini M4, 16 GB unified memory, PyTorch MPS, bf16 autocast, SDPA attention (flash-attn is CUDA-only) |
| Data | ctx1024 train, 2 % of documents held out as validation (grouped by doc_id): 331,858 train / 5,537 val questions; official test 6,664 questions |
| Recipe | Notebook train_ddp.py RLCD update: Gaussian noise on logits (Ο 0.4 β 0.1), group size 4, group-mean baseline, proper_reward (w_sph 0.75, w_rps 1.0) + soft cross-entropy |
| Hyperparameters | encoder LR 2.5e-5, head LR 1e-4, cosine to 1e-6, AdamW wd 0.01, 64 sequences per update, 4,096 padded tokens per micro-batch, grad-clip 1.0 |
| Epochs | 1 (4,126 updates); enevaeldens cases sampled at 0.4 of their natural rate |
| Time | 48.8 h at ~1,100 tokens/s |
| Selection | best validation Brier (was the final checkpoint) |
| Calibration | one temperature per (type, option-count) bucket fitted on validation: choice 1.76, noul 3.55; by options: choice:2 2.91, choice:3-5 2.00, choice:6-10 1.46, choice:11+ 1.43 |
Danish test split (ctx1024, 6,664 questions)
| Question kind | n | Acc before β after | Brier before β after | ECE before β after |
|---|---|---|---|---|
| All | 6664 | 0.397 β 0.743 | 0.871 β 0.342 | 0.318 β 0.029 |
| record_enum | 4008 | 0.296 β 0.614 | 0.939 β 0.503 | 0.296 β 0.039 |
| presence | 1282 | 0.753 β 0.937 | 0.405 β 0.116 | 0.155 β 0.059 |
| presence_negative | 1042 | 0.297 β 0.985 | 1.205 β 0.024 | 0.605 β 0.011 |
| record_bool | 177 | 0.390 β 0.718 | 1.130 β 0.350 | 0.560 β 0.087 |
| bool | 108 | 0.778 β 0.954 | 0.393 β 0.116 | 0.192 β 0.105 |
| enum | 47 | 0.617 β 0.745 | 0.561 β 0.331 | 0.186 β 0.158 |
Length buckets: every ctx1024 test state is β€ 1k tokens, so the 1kβ8k and 8kβ16k buckets are empty. The 8kβ16k
regime was not evaluated (the default config was never run).
Selected sources (accuracy before β after): enevaeldens 0.48 β 0.86, kb 0.39 β 0.68, wiki 0.54 β 0.86, health 0.32 β 0.81, cellar 0.43 β 0.85, ai-aktindsigt 0.49 β 0.88, municipality 0.43 β 0.84, hest 0.32 β 0.63, opensub 0.32 β 0.50.
Validation trajectory (3,000-question subset, raw logits): baseline Brier 0.830 β 0.377 (0.1 epoch) β 0.313 (25 %) β 0.284 (50 %) β 0.272 (75 %) β 0.271 (100 %). Gains flattened in the last quarter.
English regression check: LocalLLaMA/typed-decisions test (2,000 decisions)
| Acc before β after | Brier before β after | ECE before β after | |
|---|---|---|---|
| All | 0.344 β 0.384 | 0.466 β 0.267 | 0.324 β 0.090 |
| choice | 0.275 β 0.372 | 0.497 β 0.261 | 0.373 β 0.062 |
| noul | 0.507 β 0.503 | 0.576 β 0.285 | 0.378 β 0.140 |
| score | 0.275 β 0.302 | 0.359 β 0.259 | 0.250 β 0.118 |
No regression. Note the multilingual base is near chance on this benchmark zero-shot (upstream's BENCHMARKS.md says the same), so it is a weak English check; MASSIVE English is the better signal.
MASSIVE intent (20 options, 500 utterances per language)
| Language | Acc before β after | Brier before β after | ECE before β after |
|---|---|---|---|
| da | 0.464 β 0.518 | 0.860 β 0.612 | 0.370 β 0.057 |
| en | 0.640 β 0.688 | 0.570 β 0.424 | 0.248 β 0.052 |
| de | 0.470 β 0.516 | 0.859 β 0.632 | 0.351 β 0.046 |
| sv | 0.464 β 0.460 | 0.845 β 0.662 | 0.357 β 0.059 |
Danish, English and German improve; Swedish accuracy is flat (β0.4 pt, within noise; the 75 % checkpoint was 2.6 pt below base) while its Brier and ECE improve.
Temperature fitting (validation, 5,537 questions)
| Bucket | ECE before β after | Brier before β after |
|---|---|---|
| All | 0.067 β 0.038 | 0.279 β 0.272 |
| noul:2 | 0.038 β 0.036 | 0.094 β 0.095 |
| choice:2 | 0.115 β 0.036 | 0.291 β 0.257 |
| choice:3-5 | 0.106 β 0.046 | 0.415 β 0.398 |
| choice:6-10 | 0.083 β 0.067 | 0.543 β 0.539 |
| choice:11+ | 0.144 β 0.115 | 0.619 β 0.607 |
All fitted temperatures are > 1, i.e. the raw model is overconfident, as expected from one-hot targets.
Do option descriptions help? (400 sampled test choice questions)
Enum options in the dataset are bare keys with empty descriptions. Three renderings of the same questions:
| Option rendering | Base acc / Brier | 0.1-epoch acc / Brier |
|---|---|---|
bare keys (ikke_angivet) |
0.247 / 0.977 | 0.513 / 0.619 |
humanized keys (ikke_angivet: ikke angivet) |
0.312 / 0.911 | 0.508 / 0.624 |
| LLM-written descriptions (schema-only) | 0.273 / 0.934 | 0.498 / 0.636 |
Descriptions help the untrained model a little; after fine-tuning on bare keys the differences vanish (Β±2 pt noise). Whether training with descriptions raises the ceiling was not tested.
What didn't work / limitations
- 16k context not trained. MPS on 16 GB fits ~4k-token sequences for training (fwd+bwd at 4k: 8.7 s, 10 GB;
8k: OOM). Throughput at 1k was ~1.1k tok/s, so the
defaultepoch (1.49B tokens) would take > 1 month. Needs an A100/H100. Everything here isctx1024; states longer than 1,024 tokens are silently truncated at inference. - First run swapped. 8,192-token micro-batches hit a 19 GB footprint and throughput fell from 1,170 to 520 tok/s
over 1.3 h; restarted at 4,096 tokens with periodic
torch.mps.empty_cache()(13β15 GB, stable). - record_enum is the weak kind (0.61). Records identified by weak keys (e.g.
"55" (co2_quota_price_2020)) and evidence cut off by the 1k window are the visible failure modes; some gold labels are judgment calls. - Overconfidence, corrected by temperature (noul T = 3.55). Label smoothing was implemented (
--label-smoothing) but not run. - Dataset oddity: many test-split
workflowvalues are document IDs (e.g.1334,AE003077) rather than source names; per-source breakdowns for those rows are meaningless.
What to try next
- Train the
default(16k) config on a GPU with flash attention; curriculum from this checkpoint. - Second epoch or a
ctx1024βdefaultcurriculum: validation was still improving slowly at the end of epoch 1. - Label smoothing 0.05β0.1 (one flag) to reduce the fitted temperatures.
- Mix a few percent of MASSIVE/typed-decisions data into training to protect other languages (Swedish).
- Add short enum descriptions to the training schemas and re-run the description ablation.
- Teacher self-agreement on the test split to get a ceiling for the LLM-made labels.
- Downloads last month
- 9
Model tree for syvai/claes-preview
Base model
convaiinnovations/laya