claes preview

The multilingual Laya decision model (mmBERT-base encoder, 322M params) fine-tuned on syvai/danish-dynaword-laya (ctx1024 config): 337k typed questions over Danish Dynaword texts. Laya scores each answer option of a choice/noul question at a [MASK] marker and returns calibrated probabilities. It never generates text, so the output is always one of the options you supplied.

import laya
agent = laya.load("syvai/claes")
res = agent.predict(
    "Forslaget om kontingentforhΓΈjelse blev forkastet med 41 stemmer mod 12.",
    {"udfald": {"type": "choice", "instructions": "Udfaldet af afstemningen.",
                "criteria": {"forkastet": "", "vedtaget": ""}},
     "har_sted": {"type": "noul", "instructions": "Oplyser teksten fΓΈlgende: Stedet hvor mΓΈdet blev afholdt.",
                  "criteria": {"false": "nej", "true": "ja"}}})
print(res["answers"])
# {'udfald': {'type': 'choice', 'choice': 'forkastet', 'probabilities': {'forkastet': 0.64, 'vedtaget': 0.36}, ...},
#  'har_sted': {'type': 'noul', 'noul': 0.28, ...}}

Question phrasing that matches the training data works best:

  • presence: Oplyser teksten fΓΈlgende: <feltbeskrivelse> (noul, criteria {"false": "nej", "true": "ja"})
  • boolean / enum field: the bare field description
  • field on one item of a list: Om "<record>" (<nΓΈglefelt>): <feltbeskrivelse>

Enum options can be bare keys; adding descriptions neither helps nor hurts after fine-tuning (see ablation below).

Files: model.safetensors, rl_agent_config.json (with max_len 1024, head_max_len 256 and the fitted temperatures), tokenizer/, encoder/config.json, temperature_fit.json.


Evaluation report

Fine-tune of the multilingual Laya checkpoint (convaiinnovations/laya, subfolder multilingual/, mmBERT-base, 322M params) on syvai/danish-dynaword-laya, config ctx1024 (states up to 1,024 tokens). The 16k default config was not trained: the only available hardware was a Mac mini M4 with 16 GB, which runs out of memory at 8k tokens (see "What didn't work").

All metrics below use the temperature-fitted config shipped in this repo.

Setup

Hardware Mac mini M4, 16 GB unified memory, PyTorch MPS, bf16 autocast, SDPA attention (flash-attn is CUDA-only)
Data ctx1024 train, 2 % of documents held out as validation (grouped by doc_id): 331,858 train / 5,537 val questions; official test 6,664 questions
Recipe Notebook train_ddp.py RLCD update: Gaussian noise on logits (Οƒ 0.4 β†’ 0.1), group size 4, group-mean baseline, proper_reward (w_sph 0.75, w_rps 1.0) + soft cross-entropy
Hyperparameters encoder LR 2.5e-5, head LR 1e-4, cosine to 1e-6, AdamW wd 0.01, 64 sequences per update, 4,096 padded tokens per micro-batch, grad-clip 1.0
Epochs 1 (4,126 updates); enevaeldens cases sampled at 0.4 of their natural rate
Time 48.8 h at ~1,100 tokens/s
Selection best validation Brier (was the final checkpoint)
Calibration one temperature per (type, option-count) bucket fitted on validation: choice 1.76, noul 3.55; by options: choice:2 2.91, choice:3-5 2.00, choice:6-10 1.46, choice:11+ 1.43

Danish test split (ctx1024, 6,664 questions)

Question kind n Acc before β†’ after Brier before β†’ after ECE before β†’ after
All 6664 0.397 β†’ 0.743 0.871 β†’ 0.342 0.318 β†’ 0.029
record_enum 4008 0.296 β†’ 0.614 0.939 β†’ 0.503 0.296 β†’ 0.039
presence 1282 0.753 β†’ 0.937 0.405 β†’ 0.116 0.155 β†’ 0.059
presence_negative 1042 0.297 β†’ 0.985 1.205 β†’ 0.024 0.605 β†’ 0.011
record_bool 177 0.390 β†’ 0.718 1.130 β†’ 0.350 0.560 β†’ 0.087
bool 108 0.778 β†’ 0.954 0.393 β†’ 0.116 0.192 β†’ 0.105
enum 47 0.617 β†’ 0.745 0.561 β†’ 0.331 0.186 β†’ 0.158

Length buckets: every ctx1024 test state is ≀ 1k tokens, so the 1k–8k and 8k–16k buckets are empty. The 8k–16k regime was not evaluated (the default config was never run).

Selected sources (accuracy before β†’ after): enevaeldens 0.48 β†’ 0.86, kb 0.39 β†’ 0.68, wiki 0.54 β†’ 0.86, health 0.32 β†’ 0.81, cellar 0.43 β†’ 0.85, ai-aktindsigt 0.49 β†’ 0.88, municipality 0.43 β†’ 0.84, hest 0.32 β†’ 0.63, opensub 0.32 β†’ 0.50.

Validation trajectory (3,000-question subset, raw logits): baseline Brier 0.830 β†’ 0.377 (0.1 epoch) β†’ 0.313 (25 %) β†’ 0.284 (50 %) β†’ 0.272 (75 %) β†’ 0.271 (100 %). Gains flattened in the last quarter.

English regression check: LocalLLaMA/typed-decisions test (2,000 decisions)

Acc before β†’ after Brier before β†’ after ECE before β†’ after
All 0.344 β†’ 0.384 0.466 β†’ 0.267 0.324 β†’ 0.090
choice 0.275 β†’ 0.372 0.497 β†’ 0.261 0.373 β†’ 0.062
noul 0.507 β†’ 0.503 0.576 β†’ 0.285 0.378 β†’ 0.140
score 0.275 β†’ 0.302 0.359 β†’ 0.259 0.250 β†’ 0.118

No regression. Note the multilingual base is near chance on this benchmark zero-shot (upstream's BENCHMARKS.md says the same), so it is a weak English check; MASSIVE English is the better signal.

MASSIVE intent (20 options, 500 utterances per language)

Language Acc before β†’ after Brier before β†’ after ECE before β†’ after
da 0.464 β†’ 0.518 0.860 β†’ 0.612 0.370 β†’ 0.057
en 0.640 β†’ 0.688 0.570 β†’ 0.424 0.248 β†’ 0.052
de 0.470 β†’ 0.516 0.859 β†’ 0.632 0.351 β†’ 0.046
sv 0.464 β†’ 0.460 0.845 β†’ 0.662 0.357 β†’ 0.059

Danish, English and German improve; Swedish accuracy is flat (βˆ’0.4 pt, within noise; the 75 % checkpoint was 2.6 pt below base) while its Brier and ECE improve.

Temperature fitting (validation, 5,537 questions)

Bucket ECE before β†’ after Brier before β†’ after
All 0.067 β†’ 0.038 0.279 β†’ 0.272
noul:2 0.038 β†’ 0.036 0.094 β†’ 0.095
choice:2 0.115 β†’ 0.036 0.291 β†’ 0.257
choice:3-5 0.106 β†’ 0.046 0.415 β†’ 0.398
choice:6-10 0.083 β†’ 0.067 0.543 β†’ 0.539
choice:11+ 0.144 β†’ 0.115 0.619 β†’ 0.607

All fitted temperatures are > 1, i.e. the raw model is overconfident, as expected from one-hot targets.

Do option descriptions help? (400 sampled test choice questions)

Enum options in the dataset are bare keys with empty descriptions. Three renderings of the same questions:

Option rendering Base acc / Brier 0.1-epoch acc / Brier
bare keys (ikke_angivet) 0.247 / 0.977 0.513 / 0.619
humanized keys (ikke_angivet: ikke angivet) 0.312 / 0.911 0.508 / 0.624
LLM-written descriptions (schema-only) 0.273 / 0.934 0.498 / 0.636

Descriptions help the untrained model a little; after fine-tuning on bare keys the differences vanish (Β±2 pt noise). Whether training with descriptions raises the ceiling was not tested.

What didn't work / limitations

  • 16k context not trained. MPS on 16 GB fits ~4k-token sequences for training (fwd+bwd at 4k: 8.7 s, 10 GB; 8k: OOM). Throughput at 1k was ~1.1k tok/s, so the default epoch (1.49B tokens) would take > 1 month. Needs an A100/H100. Everything here is ctx1024; states longer than 1,024 tokens are silently truncated at inference.
  • First run swapped. 8,192-token micro-batches hit a 19 GB footprint and throughput fell from 1,170 to 520 tok/s over 1.3 h; restarted at 4,096 tokens with periodic torch.mps.empty_cache() (13–15 GB, stable).
  • record_enum is the weak kind (0.61). Records identified by weak keys (e.g. "55" (co2_quota_price_2020)) and evidence cut off by the 1k window are the visible failure modes; some gold labels are judgment calls.
  • Overconfidence, corrected by temperature (noul T = 3.55). Label smoothing was implemented (--label-smoothing) but not run.
  • Dataset oddity: many test-split workflow values are document IDs (e.g. 1334, AE003077) rather than source names; per-source breakdowns for those rows are meaningless.

What to try next

  1. Train the default (16k) config on a GPU with flash attention; curriculum from this checkpoint.
  2. Second epoch or a ctx1024 β†’ default curriculum: validation was still improving slowly at the end of epoch 1.
  3. Label smoothing 0.05–0.1 (one flag) to reduce the fitted temperatures.
  4. Mix a few percent of MASSIVE/typed-decisions data into training to protect other languages (Swedish).
  5. Add short enum descriptions to the training schemas and re-run the description ablation.
  6. Teacher self-agreement on the test split to get a ceiling for the LLM-made labels.
Downloads last month
9
Safetensors
Model size
0.3B params
Tensor type
F16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for syvai/claes-preview

Finetuned
(136)
this model

Dataset used to train syvai/claes-preview

Space using syvai/claes-preview 1