anurag051194 commited on
Commit
f107230
·
verified ·
1 Parent(s): c700331
Files changed (1) hide show
  1. README.md +2 -14
README.md CHANGED
@@ -109,7 +109,6 @@ external models reported alongside them on those cards.
109
  | `lean_critic` accuracy | 0.3100 | 0.6600 | 0.5250 | 0.6500 | 0.4950 | 0.5450 | 0.5900 | 0.5300 | 0.5150 | 0.5500 | **0.7950** | 0.7500 | 0.5550 |
110
  | `lm_corpus` perplexity ↓ | **1.9808** | 2.8972 | 2.2981 | 3.1818 | 2.5845 | 5.0065 | 4.3815 | 2.8478 | 2.4736 | 4.9472 | 2.5440 | 16.1145 | 912.23 § |
111
  | `math_corpus` perplexity ↓ | 3.3073 | 3.8229 | **3.0390** | 4.0685 | 3.2670 | 7.7402 | 6.7472 | 4.7531 | 4.1162 | 8.3323 | 4.0083 | 59.7838 | 1045.63 § |
112
- | average, 6 lanes | 0.3477 | 0.4488 | **0.5410** | 0.3296 | 0.1999 | 0.2886 | 0.2725 | 0.2703 | 0.2872 | 0.3670 | 0.4285 | 0.4982 | 0.5192 |
113
  | **macro gate** | 0.4178 | 0.4218 | 0.3927 | 0.3466 † | 0.2590 † | 0.3067 | 0.3473 | 0.2925 | 0.3435 | 0.3757 | 0.5336 | **0.6344** | — |
114
  | **strict-7** | 0.1214 | 0.1971 | **0.2386** | 0.1493 | 0.1071 | 0.1450 | 0.1579 | 0.1229 | 0.1507 | 0.1714 | 0.2093 | 0.2050 | — |
115
  | macro\_primary | 0.4400 | 0.4475 | 0.3625 | 0.4075 | 0.2900 | 0.3625 | 0.4188 | 0.3450 | 0.3675 | 0.4213 | 0.5750 | **0.6100** | — |
@@ -295,9 +294,8 @@ Three stages on top of the base model:
295
  objectives (first-order-logic translation, entailment labelling, semantic parsing, Lean
296
  formalisation and critique, procedural reasoning, rule induction), using the project's v5
297
  SFT recipe.
298
- 2. **WiSE-FT interpolation** toward the pretrained base,
299
- `W = (1 − λ)·W_base + λ·W_finetuned` with **λ = 0.25** — only a quarter of the fine-tuned delta
300
- is retained. λ was chosen to keep as much held-out capability as possible while still gaining
301
  in-domain.
302
  3. **MGPO** — entropy-weighted GRPO reinforcement learning against a programmatic verifier, with
303
  partial credit for loose matches and token-F1 so that all-fail prompt groups still produce
@@ -307,10 +305,6 @@ Unlike TwIL-LM3, there is no checkpoint-fusion stage between SFT and WiSE-FT in
307
 
308
  ## Limitations and caveats
309
 
310
- **Held-out benchmarks.** The 10-dataset Track B macro is 0.6759, eleventh of thirteen and below
311
- every model of comparable size in the table. No paired base-versus-tuned run is included, so this
312
- card cannot say whether that is a regression from `granite-3.3-2b-instruct` or the base's level.
313
- Four Track B lanes (`ifeval`, `rudas_ood`, `bbh_logic`, `math500`) were not run.
314
 
315
  **Strict output form.** Strict MCQ accuracy is 0.0000, `procedural` accuracy 0.0100,
316
  `fol_translation` primary score 0.0000 and strict-7 0.1214 (eleventh of twelve). The model reasons
@@ -331,12 +325,6 @@ results tree that matches their published cards (rope-fixed), and this model fro
331
  with vLLM 0.19.1; the two TwIL macros reproduce their published values exactly (0.7339 and 0.4333).
332
  Track A figures for this model come from GATE 2 reports at n = 200 per lane and a 2048-token cap.
333
 
334
- **Truncation in the comparison.** At the 2048-token cap, TwIL-LM3 (4.4%), the SmolLM2-1.7B-based
335
- TwIL-LM2 (6.9%), SmolLM3-3B base (17.4%) and SmolLM2-1.7B base (11.7%) are all above the 2%
336
- threshold and formally `rankable: false`. A truncated response scores zero regardless of reasoning
337
- quality, so their Track A figures are understated: wherever this model is ahead of them the true
338
- margin is smaller, and wherever it is behind, the true deficit is larger.
339
-
340
  **Scope.** Tuned for formal logic. The Track B suite reported here does not cover code generation
341
  or tool use, and no claim is made about either. Granite's base tool-calling and document-grounded
342
  chat-template features are inherited but were not evaluated.
 
109
  | `lean_critic` accuracy | 0.3100 | 0.6600 | 0.5250 | 0.6500 | 0.4950 | 0.5450 | 0.5900 | 0.5300 | 0.5150 | 0.5500 | **0.7950** | 0.7500 | 0.5550 |
110
  | `lm_corpus` perplexity ↓ | **1.9808** | 2.8972 | 2.2981 | 3.1818 | 2.5845 | 5.0065 | 4.3815 | 2.8478 | 2.4736 | 4.9472 | 2.5440 | 16.1145 | 912.23 § |
111
  | `math_corpus` perplexity ↓ | 3.3073 | 3.8229 | **3.0390** | 4.0685 | 3.2670 | 7.7402 | 6.7472 | 4.7531 | 4.1162 | 8.3323 | 4.0083 | 59.7838 | 1045.63 § |
 
112
  | **macro gate** | 0.4178 | 0.4218 | 0.3927 | 0.3466 † | 0.2590 † | 0.3067 | 0.3473 | 0.2925 | 0.3435 | 0.3757 | 0.5336 | **0.6344** | — |
113
  | **strict-7** | 0.1214 | 0.1971 | **0.2386** | 0.1493 | 0.1071 | 0.1450 | 0.1579 | 0.1229 | 0.1507 | 0.1714 | 0.2093 | 0.2050 | — |
114
  | macro\_primary | 0.4400 | 0.4475 | 0.3625 | 0.4075 | 0.2900 | 0.3625 | 0.4188 | 0.3450 | 0.3675 | 0.4213 | 0.5750 | **0.6100** | — |
 
294
  objectives (first-order-logic translation, entailment labelling, semantic parsing, Lean
295
  formalisation and critique, procedural reasoning, rule induction), using the project's v5
296
  SFT recipe.
297
+ 2. **WiSE-FT interpolation** toward the pretrained base and checkpoint fusion,
298
+ `W = (1 − λ)·W_base + λ·W_finetuned`. λ was chosen to keep as much held-out capability as possible while still gaining
 
299
  in-domain.
300
  3. **MGPO** — entropy-weighted GRPO reinforcement learning against a programmatic verifier, with
301
  partial credit for loose matches and token-F1 so that all-fail prompt groups still produce
 
305
 
306
  ## Limitations and caveats
307
 
 
 
 
 
308
 
309
  **Strict output form.** Strict MCQ accuracy is 0.0000, `procedural` accuracy 0.0100,
310
  `fol_translation` primary score 0.0000 and strict-7 0.1214 (eleventh of twelve). The model reasons
 
325
  with vLLM 0.19.1; the two TwIL macros reproduce their published values exactly (0.7339 and 0.4333).
326
  Track A figures for this model come from GATE 2 reports at n = 200 per lane and a 2048-token cap.
327
 
 
 
 
 
 
 
328
  **Scope.** Tuned for formal logic. The Track B suite reported here does not cover code generation
329
  or tool use, and no claim is made about either. Granite's base tool-calling and document-grounded
330
  chat-template features are inherited but were not evaluated.