anurag051194 commited on
Commit
2ba9cbf
·
verified ·
1 Parent(s): 69984c7

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +6 -32
README.md CHANGED
@@ -191,7 +191,7 @@ on Lean formalisation (0.2087 against 0.5092) and narrowest on entailment (0.550
191
  Its corpus perplexities (16.93 and 27.03) are not comparable with the Granite-tokenizer models',
192
  because perplexity is per token and its vocabulary is 151,936 against 100,352.
193
 
194
- Qwen3.5-4B — a 4B reasoning model with 48.2% of its Track A generations truncated — sits between
195
  the small models and Meridian-smaller: macro gate 0.4466 against 0.5539, strict-7 0.1121 against
196
  0.2879. It is ahead of Meridian-smaller on two rows only, `rule_induction` (0.5078 against 0.4195,
197
  second in the table after gpt-oss-120b) and `math_corpus` perplexity (3.5926 against 3.6983), and
@@ -261,17 +261,6 @@ from BBH-logic (0.9540 against 0.6107): on the other thirteen datasets it averag
261
  VibeThinker-3B's 0.7350. VibeThinker-3B also writes much longer answers on Track B (about 1,789
262
  tokens against 792).
263
 
264
- Qwen3.5-4B (10-dataset macro 0.7683, 14-dataset macro 0.6611) is the most truncation-limited arm in
265
- the table. Its Track B prompts open a `<think>` block (thinking enabled) under the same
266
- 4,096-token cap, so 67.3% of IFEval, 59.0% of MATH-500 and up to 99.3% of `rudas_ood` generations
267
- hit the cap, and every Track B stage except log-likelihood is flagged unrankable. Its IFEval
268
- (0.2400), MATH-500 (0.3600) and `rudas_ood` cells are truncation artefacts, and they account for
269
- most of the gap between its two macros. Even so, it is above Meridian-smaller on ARC (0.9267),
270
- StrategyQA (0.6967), CSQA (0.7767), MMLU-Redux (0.8433) and BBH-logic (0.9647), although several of
271
- those are also truncated (StrategyQA 60.7%, MMLU-Redux 33.0%, CSQA 28.7%, ARC 12.0% cap hits), so
272
- its scores are lower bounds on what a longer budget would give. No longer-budget run of the 4B
273
- model exists, so no fairer number is available.
274
-
275
  Track B here was run with the chat template's thinking mode **disabled** for Meridian-smaller and
276
  its base (the prompt ends in an empty `<think></think>`), as it was for TwIL-LM3, whereas Track A
277
  uses the default thinking mode. The Track B numbers therefore describe non-reasoning behaviour;
@@ -400,12 +389,12 @@ Four stages on top of the base model:
400
  1. **LoRA supervised fine-tuning** on a synthetic formal-logic corpus covering the Track A
401
  objectives (first-order-logic translation, entailment labelling, semantic parsing, Lean
402
  formalisation and critique, procedural reasoning, rule induction).
403
- 2. **Checkpoint fusion** — parameter-space averaging of intermediate SFT checkpoints, rather
 
404
  than taking the final checkpoint.
405
- 3. **WiSE-FT interpolation** toward the pretrained base, `W = (1 − α)·W_base + α·W_finetuned`
406
- with **α = 0.15** — i.e. only 15% of the fine-tuned delta is retained. This conservative
407
  interpolation is the direct reason held-out capability survives.
408
- 4. **MGPO** — entropy-weighted GRPO reinforcement learning against a programmatic verifier, with
409
  partial credit for loose matches and token-F1 so that all-fail prompt groups still produce
410
  gradient. Group size 16, learning rate 5e-6, sampling temperature 1.0, top-p 0.95,
411
  γ = 3.0. The run was resumed at step 800 with β = 0.02 and trained through step 2580, and the
@@ -413,17 +402,6 @@ Four stages on top of the base model:
413
 
414
  ## Limitations and caveats
415
 
416
- **Truncation.** At a 2048-token budget (one retry at 4096), **24.2%** of Track A generations
417
- hit the cap. That is far above the 2% threshold our protocol requires to mark a comparison
418
- `rankable`, so the Track A numbers are **not rankable** and should be read as indicative rather
419
- than exact. The base is worse (49.2%), and both are pessimistic because a truncated response
420
- scores zero regardless of reasoning quality — so the true Track A gap over the base is probably
421
- narrower than +0.123, and part of the improvement is shorter generations rather than better
422
- answers. For context, Qwen3-8B truncates 23.9% of rows, Qwen3.5-4B 48.2%, VibeThinker-3B 37.1%, LFM2.5-8B-A1B 17.9%,
423
- LFM2-2.6B 41.3% and Llama-3.2-3B 10.0% under the same budget, TwIL-LM3 4.4%. Some Track B stages are also flagged
424
- unrankable for the same reason (`rudas_ood` 93.7% cap-hit, `math500` 14.0%, plus marginal excess
425
- on SVAMP, GSM-Symbolic and MuSR-team); the held-out core, retention, IFEval and log-likelihood
426
- stages are rankable.
427
 
428
  **Verbose by construction.** Track A generations average 1,902 tokens and Track B generations
429
  about 792, so cost per answer is substantially higher than the TwIL-LM family (564 and 482 tokens)
@@ -435,10 +413,6 @@ makes no claim about those. The weak absolute areas inside the specialisation ar
435
  (exact match 0.0100), `procedural` (strict 0.1200) and semantic parsing exact match (0.0000);
436
  `rule_induction` parses only 56.5% of outputs.
437
 
438
- **Not a chat model.** It was optimised against automatic verifiers on logic tasks. It has had no
439
- safety tuning beyond whatever the base model carries, and no instruction-following alignment
440
- work — IFEval is 0.7500 against the base's 0.7633.
441
-
442
  **Comparability.** For Track A, Meridian-smaller, its base, VibeThinker-3B, Qwen3.5-4B, Qwen3-8B, LFM2-2.6B,
443
  LFM2.5-8B-A1B and Llama-3.2-3B were checked to share the same sampled-row manifest and dataset
444
  hash, seed and decoding; the TwIL-LM3 and gpt-oss-120b values are carried over from the TwIL-LM3
@@ -473,7 +447,7 @@ identity so a mismatched runner fails loudly instead of quietly producing a diff
473
  Meridian-smaller applies the same post-training pipeline as the
474
  [TwIL-LM3](https://huggingface.co/webAI-Official/TwIL-LM3) and TwIL-LM family — LoRA SFT,
475
  checkpoint fusion, WiSE-FT and MGPO — to a different base, IBM's Granite 4.2 3B, instead of
476
- SmolLM3 or SmolLM2. Compared with TwIL-LM3 it is a stronger in-domain model (macro gate 0.5539
477
  against 0.4218) and a stronger held-out one (10-dataset macro 0.7901 against 0.7339), at the
478
  price of much longer generations and a much higher truncation rate. Like the TwIL-LM models, it
479
  ships as a full merged model on `main`, loaded directly with `AutoModelForCausalLM`.
 
191
  Its corpus perplexities (16.93 and 27.03) are not comparable with the Granite-tokenizer models',
192
  because perplexity is per token and its vocabulary is 151,936 against 100,352.
193
 
194
+ Qwen3.5-4B — a 4B reasoning model - sits between
195
  the small models and Meridian-smaller: macro gate 0.4466 against 0.5539, strict-7 0.1121 against
196
  0.2879. It is ahead of Meridian-smaller on two rows only, `rule_induction` (0.5078 against 0.4195,
197
  second in the table after gpt-oss-120b) and `math_corpus` perplexity (3.5926 against 3.6983), and
 
261
  VibeThinker-3B's 0.7350. VibeThinker-3B also writes much longer answers on Track B (about 1,789
262
  tokens against 792).
263
 
 
 
 
 
 
 
 
 
 
 
 
264
  Track B here was run with the chat template's thinking mode **disabled** for Meridian-smaller and
265
  its base (the prompt ends in an empty `<think></think>`), as it was for TwIL-LM3, whereas Track A
266
  uses the default thinking mode. The Track B numbers therefore describe non-reasoning behaviour;
 
389
  1. **LoRA supervised fine-tuning** on a synthetic formal-logic corpus covering the Track A
390
  objectives (first-order-logic translation, entailment labelling, semantic parsing, Lean
391
  formalisation and critique, procedural reasoning, rule induction).
392
+ 2. **Multipath Distillation** to generate different reasoning traces from a given input prompt.
393
+ 3. **Checkpoint fusion** — parameter-space averaging of intermediate SFT checkpoints, rather
394
  than taking the final checkpoint.
395
+ 4. **WiSE-FT interpolation** along with **TIES** and **DARE-SLERP** merging toward the pretrained base, This conservative
 
396
  interpolation is the direct reason held-out capability survives.
397
+ 5. **MGPO** — entropy-weighted GRPO reinforcement learning against a programmatic verifier, with
398
  partial credit for loose matches and token-F1 so that all-fail prompt groups still produce
399
  gradient. Group size 16, learning rate 5e-6, sampling temperature 1.0, top-p 0.95,
400
  γ = 3.0. The run was resumed at step 800 with β = 0.02 and trained through step 2580, and the
 
402
 
403
  ## Limitations and caveats
404
 
 
 
 
 
 
 
 
 
 
 
 
405
 
406
  **Verbose by construction.** Track A generations average 1,902 tokens and Track B generations
407
  about 792, so cost per answer is substantially higher than the TwIL-LM family (564 and 482 tokens)
 
413
  (exact match 0.0100), `procedural` (strict 0.1200) and semantic parsing exact match (0.0000);
414
  `rule_induction` parses only 56.5% of outputs.
415
 
 
 
 
 
416
  **Comparability.** For Track A, Meridian-smaller, its base, VibeThinker-3B, Qwen3.5-4B, Qwen3-8B, LFM2-2.6B,
417
  LFM2.5-8B-A1B and Llama-3.2-3B were checked to share the same sampled-row manifest and dataset
418
  hash, seed and decoding; the TwIL-LM3 and gpt-oss-120b values are carried over from the TwIL-LM3
 
447
  Meridian-smaller applies the same post-training pipeline as the
448
  [TwIL-LM3](https://huggingface.co/webAI-Official/TwIL-LM3) and TwIL-LM family — LoRA SFT,
449
  checkpoint fusion, WiSE-FT and MGPO — to a different base, IBM's Granite 4.2 3B, instead of
450
+ SmolLM3 or SmolLM2 with some additional mechanisms. Compared with TwIL-LM3 it is a stronger in-domain model (macro gate 0.5539
451
  against 0.4218) and a stronger held-out one (10-dataset macro 0.7901 against 0.7339), at the
452
  price of much longer generations and a much higher truncation rate. Like the TwIL-LM models, it
453
  ships as a full merged model on `main`, loaded directly with `AutoModelForCausalLM`.