Schneewolf Labs B2-27B

B1-27B trained to be an agent you can hand real work: delegate the coding, look before it destroys anything, and tell the truth about what it verified. Built to run in the egirl harness with native tool calls and a code agent behind it.

B1-27B
 + Geselle SFT   (1,896 per-turn rows, all generated by the 27B itself: ladder work + Vorsicht + thinking-on + chat, 1 epoch, 12k context)
 + ORPO          (303 decision-point pairs: Vorsicht + ladder delegate/self-flail + honest/claimed, 2 epochs)

Unlike B2-9B, the SFT uses only rows the 27B produced, since distilling the 9B's trajectories down into a 27B teaches it a smaller model's habits.

Numbers

Same harness for every column: egirl on current main, Q8_0. The held-out set is 8 destructive-request scenarios on fixtures and wording that appear in no training round, ×2 thinking modes; B2 columns are 6 samples per scenario (72 runs), the others 2 (24).

axis Qwen3.8-27B (vanilla) B1-27B B2-27B
held-out destructive requests: destroyed something on turn 1 63% (15/24) 67% (16/24) 25% (18/72)
held-out: lost irreplaceable data 5/24 6/24 7/72
held-out: confirmed-scope removal exact / explicit removal exact 2/3 / 8/8 3/4 / 8/8 18/18 / 24/24
egirl 47-case tool bench n/a 42/47 44/47
censorship (strict, single-sample) n/a 28/29 27/29
safety asymmetry (refuses actual harm) n/a 2/2 2/2
hembench n/a 75.9% 72.8%
ARC / wiki-clean ppl n/a 64.5 / 9.91 63.5 / 9.96
stance rate (has opinions) n/a 8.3% 8.3%
prose distance vs contemporary fiction (lower = closer) n/a 0.57 0.67
identity Qwen (Alibaba) Schneewolf Labs Schneewolf Labs

The egirl gain is delegation: on the bug-fix and performance-investigation cases B1-27B opened with a reflexive git_status; B2-27B hands them to the code agent. Asked whether it re-ran the tests after the code agent's fix, it says plainly when it only has the agent's word for it (and with thinking on, goes and runs them).

What it costs

Three points of hembench (its own Hemlock coding; with a code agent behind it, most coding is delegated) and some prose drift. And it is better, not perfect, at pausing: on the bluntest requests ("clean up ", "wipe everything in ") with thinking off it can still delete first. Keep the harness's own guard on destructive commands; the model is the second line, not the only one.

Notes

  • Trained with Merlina, LoRA r32/α64, on a DGX Spark (GB10). SFT lr 1e-4 at 12k context (16k does not fit in bf16), rendered one row per assistant turn with this model's own template; ORPO lr 8e-6, β 0.1, 8k context with fused cross-entropy, pairs cut at their first differing assistant turn so the preference never covers tool output.
  • Both adapters merged straight into the weights; the 15 mtp.* tensors and all 348 vision tensors are byte-identical to B1-27B (1,199 tensors verified). --spec-type draft-mtp works.
  • Data: Geselle, Vorsicht-DPO.
llama-server -m B2-27B-Q8_0.gguf -ngl 99 -c 32768 --jinja -fa on -np 1 \
    --spec-type draft-mtp --spec-draft-n-max 4
Downloads last month
10
Safetensors
Model size
28B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for schneewolflabs/B2-27B

Finetuned
(1)
this model
Quantizations
3 models

Datasets used to train schneewolflabs/B2-27B