anurag051194 commited on
Commit
5221069
·
verified ·
1 Parent(s): 8e59f4d

Add TwIL-LM2 (Granite-3.3-2B): model card, licences, config and tokenizer

Browse files
LICENSE.md ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ webAI Non-Commercial License ver. 1.0
2
+
3
+
4
+ 1. Definitions
5
+
6
+ “Licensor” means any person or entity that distributes its Work. “Work” means (a) the original work of authorship made available under this license, which may include software, documentation, or other files, and (b) any additions to or derivative works thereof that are made available under this license. The terms “reproduce,” “reproduction,” “derivative works,” and “distribution” have the meaning as provided under U.S. copyright law; provided, however, that for the purposes of this license, derivative works shall not include works that remain separable from, or merely link (or bind by name) to the interfaces of, the Work. Works are “made available” under this license by including in or with the Work either (a) a copyright notice referencing the applicability of this license to the Work, or (b) a copy of this license.
7
+
8
+
9
+ 2. License Grant
10
+
11
+ 2.1 Copyright Grant. Subject to the terms and conditions of this license, each Licensor grants to you a perpetual, worldwide, non-exclusive, royalty-free, copyright license to use, reproduce, prepare derivative works of, publicly display, publicly perform, sublicense and distribute its Work and any resulting derivative works in any form.
12
+
13
+
14
+ 3. Limitations
15
+
16
+ 3.1 Redistribution. You may reproduce or distribute the Work only if (a) you do so under this license, (b) you include a complete copy of this license with your distribution, and (c) you retain without modification any copyright, patent, trademark, or attribution notices that are present in the Work.
17
+
18
+ 3.2 Derivative Works. You may specify that additional or different terms apply to the use, reproduction, and distribution of your derivative works of the Work (“Your Terms”) only if (a) Your Terms provide that the use limitation in Section 3.3 applies to your derivative works, and (b) you identify the specific derivative works that are subject to Your Terms. Notwithstanding Your Terms, this license (including the redistribution requirements in Section 3.1) will continue to apply to the Work itself.
19
+
20
+ 3.3 Use Limitation. The Work and any derivative works thereof only may be used or intended for use non-commercially. As used herein, “non-commercially” means for non-commercial research and educational purposes only.
21
+
22
+ 3.4 Patent Claims. If you bring or threaten to bring a patent claim against any Licensor (including any claim, cross-claim or counterclaim in a lawsuit) to enforce any patents that you allege are infringed by any Work, then your rights under this license from such Licensor (including the grant in Section 2.1) will terminate immediately.
23
+
24
+ 3.5 Trademarks. This license does not grant any rights to use any Licensor's or its affiliates' names, logos, or trademarks, except as necessary to reproduce the notices described in this license. Nothing in this license shall be construed as permission to use the trade names, trademarks, service marks, or product names of webAI, Inc. or its affiliates to endorse, promote, or imply association with any derivative work, product, service, or entity, without prior written consent from webAI, Inc. Use of the Work does not imply endorsement by webAI, Inc. of any derivative work, product, service, or entity.
25
+
26
+ 3.6 Termination. If you violate any term of this license, then your rights under this license (including the grant in Section 2.1) will terminate immediately.
27
+
28
+
29
+ 4. Disclaimer of Warranty.
30
+
31
+ THE WORK IS PROVIDED “AS IS” WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, EITHER EXPRESS OR IMPLIED, INCLUDING WARRANTIES OR CONDITIONS OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE, TITLE OR NON-INFRINGEMENT. YOU BEAR THE RISK OF UNDERTAKING ANY ACTIVITIES UNDER THIS LICENSE.
32
+
33
+
34
+ 5. Limitation of Liability.
35
+
36
+ EXCEPT AS PROHIBITED BY APPLICABLE LAW, IN NO EVENT AND UNDER NO LEGAL THEORY, WHETHER IN TORT (INCLUDING NEGLIGENCE), CONTRACT, OR OTHERWISE SHALL ANY LICENSOR BE LIABLE TO YOU FOR DAMAGES, INCLUDING ANY DIRECT, INDIRECT, SPECIAL, INCIDENTAL, OR CONSEQUENTIAL DAMAGES ARISING OUT OF OR RELATED TO THIS LICENSE, THE USE OR INABILITY TO USE THE WORK (INCLUDING BUT NOT LIMITED TO LOSS OF GOODWILL, BUSINESS INTERRUPTION, LOST PROFITS OR DATA, COMPUTER FAILURE OR MALFUNCTION, OR ANY OTHER DAMAGES OR LOSSES), EVEN IF THE LICENSOR HAS BEEN ADVISED OF THE POSSIBILITY OF SUCH DAMAGES.
README.md ADDED
@@ -0,0 +1,372 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ library_name: transformers
5
+ pipeline_tag: text-generation
6
+ base_model: ibm-granite/granite-3.3-2b-instruct
7
+ license: other
8
+ license_name: webai-non-commercial-license-ver.-1.0
9
+ license_link: https://huggingface.co/webAI-Official/TwIL-LM2/blob/main/LICENSE.md
10
+ tags:
11
+ - granite
12
+ - formal-logic
13
+ - reasoning
14
+ - lora
15
+ - model-merging
16
+ - wise-ft
17
+ - reinforcement-learning
18
+ - grpo
19
+ - twil-lm
20
+ - gguf
21
+ ---
22
+
23
+ # TwIL-LM2
24
+
25
+ A 2.5B reasoning model for **formal logic** tasks, built from
26
+ [`ibm-granite/granite-3.3-2b-instruct`](https://huggingface.co/ibm-granite/granite-3.3-2b-instruct)
27
+ through LoRA supervised fine-tuning, WiSE-FT weight interpolation (λ = 0.25) and entropy-weighted
28
+ GRPO reinforcement learning (MGPO, step 1400).
29
+
30
+ It is the **balanced** point of its training arm. On the in-domain macro gate it scores **0.4178**,
31
+ within 0.004 of TwIL-LM3 (0.4218) at about 18% fewer parameters, and it keeps a 10-dataset
32
+ held-out macro of **0.6759** — second only to TwIL-LM3 (0.7339) among the models compared below.
33
+ It also decodes **1.57x faster than TwIL-LM3** under identical forced work.
34
+
35
+ The price of that balance is structured-output strictness. Its strict-7 score (0.1214) and strict
36
+ MCQ accuracy (0.0000) are well below the λ = 0.5 sibling from the same arm (0.2457 and 0.2850),
37
+ which gives back roughly 0.044 of held-out macro to get them. This release is the held-out-friendly
38
+ end of that trade; the other end was not published. See [Limitations](#limitations-and-caveats).
39
+
40
+ ## Highlights
41
+
42
+ * **Gate parity with TwIL-LM3 at a smaller size.** Macro gate 0.4178 against 0.4218, at 2.53B
43
+ against 3.08B parameters. Read the parity as approximate: TwIL-LM3's figure is understated by
44
+ truncation (4.4% of its Track A rows hit the token cap), so its true gate is somewhat higher.
45
+ * **Held-out capability is largely retained.** 10-dataset Track B macro 0.6759, against 0.6317 for
46
+ the λ = 0.5 sibling and 0.4333 for the SmolLM2-1.7B-based TwIL-LM2. The gap to the λ = 0.5 arm is
47
+ widest on the math sets: GSM8K 0.7867 against 0.6767 and GSM-Symbolic 0.7233 against 0.5033.
48
+ * **A cleaner measurement.** Only 0.9% of Track A generations hit the 2048-token cap, under the 2%
49
+ threshold our protocol requires to mark a comparison `rankable`. TwIL-LM3 (4.4%) and the
50
+ SmolLM2-1.7B-based model (6.9%) are both above it.
51
+ * **Fast.** 21,369 decode tokens/s on one H100 in a controlled bench — 1.57x TwIL-LM3 and 1.22x the
52
+ SmolLM2-1.7B-based model — because Granite's architecture decodes quickly, not because it answers
53
+ short.
54
+ * **Lowest `lm_corpus` perplexity in the comparison** (1.9808), and 3.3073 on `math_corpus`. Read
55
+ these with the tokenizer caveat under [Limitations](#limitations-and-caveats).
56
+ * **Runs anywhere.** 2.53B parameters in bf16 (4.72 GiB), with a Q4\_K\_M GGUF at 1.44 GiB for CPU.
57
+
58
+ It is not a general assistant, and it is not the strongest structured-output model here:
59
+ `lean_critic` accuracy is 0.3100 against 0.6600 for TwIL-LM3.
60
+
61
+ ## Model Details
62
+
63
+ | Property | Value |
64
+ | ------------------------- | ----------------------------------------------------------------------------------------------- |
65
+ | Model ID | `webAI-Official/TwIL-LM2` |
66
+ | Base model | [`ibm-granite/granite-3.3-2b-instruct`](https://huggingface.co/ibm-granite/granite-3.3-2b-instruct) |
67
+ | Total parameters | 2.53B (2,533,539,840; tied input/output embeddings) |
68
+ | Architecture | Granite decoder-only transformer; 40 layers, hidden size 2048, 32 attention heads, 8 KV heads |
69
+ | Input / output | Text / text |
70
+ | Language | English |
71
+ | Vocabulary size | 49,159 embedding rows |
72
+ | Context window | 131,072 tokens (inherited from the base; see note below) |
73
+ | Checkpoint precision | bfloat16 (4.72 GiB), plus Q4\_K\_M / Q5\_K\_M / Q6\_K / Q8\_0 / F16 GGUF builds |
74
+ | Post-training | LoRA SFT → WiSE-FT (λ = 0.25) → MGPO reinforcement learning (step 1400) |
75
+ | Chat template | Granite chat template (`<\|start_of_role\|>…<\|end_of_role\|>`), EOS `<\|end_of_text\|>` |
76
+ | Reasoning format | Answers directly under the default chat template; no `<think>` block was observed |
77
+ | Evaluated decoding | Greedy; 2048 new tokens (Track A), 4096 new tokens (Track B) |
78
+ | Specialisation | Formal logic: FOL translation, entailment, semantic parsing, Lean formalisation and critique |
79
+ | License | webAI Non-Commercial License ver. 1.0 |
80
+
81
+ The base model's 131,072-token context is carried through unchanged, but every score on this card
82
+ was measured with generation budgets of 2048 (Track A) or 4096 (Track B) tokens. Longer contexts are
83
+ inherited rather than validated here. Granite's optional `thinking=True` chat-template mode is also
84
+ inherited from the base and was not evaluated for this card.
85
+
86
+ ## Results
87
+
88
+ Every column below is a checkpoint we trained or released, scored through the same harness and
89
+ decoding settings described under [Evaluation protocol](#evaluation-protocol):
90
+
91
+ * **this model** — the λ = 0.25, MGPO step-1400 checkpoint in this repository;
92
+ * **λ = 0.25, step 300** — the next-best gate in the same arm;
93
+ * **λ = 0.5, step 800** — the arm's highest Track A gate, with weaker held-out results (not released);
94
+ * **TwIL-LM3** — [`webAI-Official/TwIL-LM3`](https://huggingface.co/webAI-Official/TwIL-LM3),
95
+ SmolLM3-3B, WiSE-FT λ = 0.25, MGPO step 2071;
96
+ * **TwIL-LM2 (SmolLM2-1.7B)** — the earlier, SmolLM2-1.7B-Instruct-based TwIL-LM2 (WiSE-FT
97
+ λ = 0.75, MGPO step 1680), using its published figures.
98
+
99
+ ### Track A — in-domain formal logic
100
+
101
+ | lane / metric | **This model** | λ=0.25 s300 | λ=0.5 s800 | TwIL-LM3 | TwIL-LM2 (SmolLM2-1.7B) |
102
+ | -------------------------- | -------------- | ----------- | ---------- | ---------- | ----------------------- |
103
+ | parameters | 2.53B | 2.53B | 2.53B | 3.08B | 1.7B |
104
+ | `lean_formalize` token-F1 | 0.5159 | 0.5455 | **0.7175** | 0.5869 | 0.6199 |
105
+ | `rule_induction` derivation | 0.3292 | 0.3471 | **0.5516** | 0.3192 | 0.5136 |
106
+ | `entailment_label` accuracy | 0.5300 | 0.5000 | **0.6250** | 0.5750 | 0.5850 |
107
+ | `mcq_answer` accuracy | 0.0000 | 0.0000 | **0.2850** | 0.1100 | 0.1600 |
108
+ | `semantic_parse` token-F1 | 0.4013 | 0.4171 | 0.6171 | 0.4416 | **0.8428** |
109
+ | `lean_critic` accuracy | 0.3100 | 0.3400 | 0.5350 | **0.6600** | 0.5250 |
110
+ | `procedural` accuracy | 0.0100 | 0.0150 | **0.0750** | 0.0300 | 0.0100 |
111
+ | `fol_translation` primary | 0.0000 | 0.0000 | 0.1000 | 0.0000 | **0.1150** |
112
+ | `lm_corpus` perplexity ↓ | **1.9808** | 1.9948 | 1.9982 | 2.8972 | 2.2981 |
113
+ | `math_corpus` perplexity ↓ | 3.3073 | 3.3731 | 3.2403 | 3.8229 | **3.0390** |
114
+ | average, 6 lanes | 0.3477 | 0.3583 | **0.5552** | 0.4488 | 0.5410 |
115
+ | **macro gate** | 0.4178 | 0.4114 | **0.4393** | 0.4218 | 0.3927 |
116
+ | **strict-7** | 0.1214 | 0.1221 | **0.2457** | 0.1971 | 0.2386 |
117
+ | macro\_primary | 0.4400 | 0.4275 | 0.4113 | **0.4475** | 0.3625 |
118
+ | mean generation length ↓ (tokens) | 517 | 505 | **213** | 443 | 460 |
119
+ | truncation at 2048 tokens | 0.9% | 0.6% | 0.5% | 4.4% | 6.9% |
120
+ | `rankable` (< 2% truncation) | yes | yes | yes | no | no |
121
+
122
+ `average, 6 lanes`, `macro gate`, `macro_primary` and `strict-7` are the harness aggregates defined
123
+ on the [TwIL-LM3 model card](https://huggingface.co/webAI-Official/TwIL-LM3); this card does not
124
+ redefine them. They are not interchangeable, and the ordering changes between them — this model
125
+ leads neither `macro_primary` nor `strict-7`.
126
+
127
+ **Reading it.** This is a mid-table formal-logic model whose main asset is held-out retention.
128
+ Against TwIL-LM3 it wins on `rule_induction` (0.3292 against 0.3192) and on both perplexities, and
129
+ loses on the rest, heavily on `lean_critic` (0.3100 against 0.6600). Its strict-7 (0.1214) is well
130
+ below TwIL-LM3's (0.1971), so the near-parity is on the gate, which credits loose matches on some
131
+ lanes, and not on strict scoring.
132
+
133
+ The λ = 0.5 sibling is the stronger *formal-logic* model — it leads nine of the fifteen scored rows
134
+ above — and the SmolLM2-1.7B-based model is the semantic-parsing specialist at 0.8428. Neither
135
+ keeps held-out capability the way this checkpoint does; see Track B.
136
+
137
+ Two things are worth stating plainly. Strict MCQ accuracy is **0.0000** for both λ = 0.25
138
+ checkpoints: they do not emit the exactly-requested answer form on that lane. And harness
139
+ answers-per-second (12.07 for this model, 17.57 for λ = 0.5) reflects how long each model chooses
140
+ to write, not how fast it decodes; use the controlled bench below for speed.
141
+
142
+ ### Track B — held-out benchmarks
143
+
144
+ Nothing in this suite was trained on. Ten chain-of-thought datasets, 300 randomly sampled examples
145
+ each, identical rows for every column.
146
+
147
+ | dataset | **This model** | λ=0.25 s300 | λ=0.5 s800 | TwIL-LM3 | TwIL-LM2 (SmolLM2-1.7B) |
148
+ | --------------------------- | -------------- | ----------- | ---------- | ---------- | ----------------------- |
149
+ | `gsm8k` | 0.7867 | 0.8000 | 0.6767 | **0.8733** | 0.4633 |
150
+ | `svamp` | 0.8200 | 0.8400 | 0.7333 | **0.8500** | 0.3833 |
151
+ | `gsm_symbolic` | 0.7233 | 0.7400 | 0.5033 | **0.7567** | 0.2600 |
152
+ | `arc_cot` | 0.7367 | 0.7300 | 0.7067 | **0.8467** | 0.5200 |
153
+ | `logicbench` | 0.6900 | 0.6933 | 0.6833 | **0.7167** | 0.5400 |
154
+ | `strategyqa` | 0.6933 | 0.6733 | **0.7233** | 0.6500 | 0.5900 |
155
+ | `drop` | 0.5833 | 0.6000 | 0.5700 | **0.7467** | 0.4367 |
156
+ | `csqa` | 0.6900 | 0.6767 | 0.7233 | **0.7367** | 0.4333 |
157
+ | `musr` (mean of 3) | 0.4957 | 0.4877 | 0.4904 | **0.4958** | 0.3131 |
158
+ | `mmlu_redux` | 0.5400 | 0.5167 | 0.5067 | **0.6667** | 0.3933 |
159
+ | **macro (10 CoT datasets)** | 0.6759 | 0.6758 | 0.6317 | **0.7339** | 0.4333 |
160
+
161
+ **TwIL-LM3 leads this table**, on nine of ten datasets and on the macro. This model is second on
162
+ the macro (and ties with the s300 checkpoint), ahead of the λ = 0.5 arm by 0.044 and of the
163
+ SmolLM2-1.7B-based model by 0.243. It is close to TwIL-LM3 on `musr` (0.4957 against 0.4958) and
164
+ `logicbench` (0.6900 against 0.7167), and furthest behind on `drop` (−0.163) and `mmlu_redux`
165
+ (−0.127).
166
+
167
+ The 14-dataset macro (adding `ifeval`, `rudas_ood`, `bbh_logic` and `math500`) was not run for the
168
+ Granite checkpoints in a directly comparable tree, so it is not reported here, and **no
169
+ instruction-following result is claimed**. It is also not known from these runs how this model
170
+ compares with its own base, `granite-3.3-2b-instruct`: no paired base-versus-tuned evaluation is
171
+ included in this card.
172
+
173
+ ### Speed
174
+
175
+ Harness tokens-per-second and answers-per-second are confounded by how much each model writes. For
176
+ an actual speed comparison, each model ran alone on one idle H100 (vLLM 0.19.1, torch 2.10.0+cu128,
177
+ transformers 5.15.0) over the same 128 prompts with `ignore_eos` and a hard 512-token cap, so every
178
+ model emitted exactly 65,536 output tokens.
179
+
180
+ | controlled decode bench | **This model** | TwIL-LM3 | TwIL-LM2 (SmolLM2-1.7B) |
181
+ | ------------------------------- | -------------- | -------- | ----------------------- |
182
+ | decode tokens/s | **21,369** | 13,623 | 17,542 |
183
+ | 512-token completions/s | **41.7** | 26.6 | 34.2 |
184
+ | decode wall seconds (65,536 tok) | **3.07** | 4.81 | 3.74 |
185
+ | relative to this model | 1.00x | 0.64x | 0.82x |
186
+
187
+ The three Granite checkpoints decode within 1.1% of one another (21,236 to 21,467 tokens/s), so
188
+ the speed advantage comes from the architecture, not from a particular checkpoint. The shared
189
+ prompt file is Granite-templated, so prefill differs slightly by tokenizer (30.7K tokens for
190
+ Granite and SmolLM3, 38.2K for SmolLM2); it is a small share of the forced output and, if anything,
191
+ slightly handicaps the SmolLM2-based model.
192
+
193
+ ### Choosing a checkpoint within the arm
194
+
195
+ All four evaluated λ = 0.25 checkpoints sit within 0.0006 of one another on Track B, so held-out
196
+ performance cannot separate them; the ordering comes from Track A and the probe.
197
+
198
+ | MGPO step | Track B R10 | Track A macro gate |
199
+ | --------- | ----------- | ------------------ |
200
+ | 300 | 0.6758 | 0.4114 |
201
+ | 1100 | 0.6764 | 0.4019 |
202
+ | **1400** | 0.6759 | **0.4178** |
203
+ | 1700 | 0.6761 | 0.4059 |
204
+
205
+ Step 1400 is the checkpoint chosen by probe Pass@1 and has the highest gate of the four.
206
+
207
+ ### Where it sits in the wider Granite family
208
+
209
+ For context, an internal, **unreleased** Granite 4.1 3B MGPO checkpoint (fusion + WiSE-FT λ = 0.5,
210
+ step 600), scored on the same Track A and Track B protocol, reaches a macro gate of 0.5105,
211
+ strict-7 of 0.2886 and Track B R10 of 0.7761 — 9.3, 16.7 and 10.0 percentage points above this
212
+ model respectively. This model is the smaller, faster option; it is not the strongest Granite
213
+ checkpoint we have trained.
214
+
215
+ ## Usage
216
+
217
+ ```python
218
+ import torch
219
+ from transformers import AutoModelForCausalLM, AutoTokenizer
220
+
221
+ model_id = "webAI-Official/TwIL-LM2"
222
+ tok = AutoTokenizer.from_pretrained(model_id)
223
+ model = AutoModelForCausalLM.from_pretrained(
224
+ model_id, torch_dtype=torch.bfloat16, device_map="auto"
225
+ )
226
+
227
+ messages = [{"role": "user", "content":
228
+ "Does 'All dogs are mammals. Rex is a dog.' entail 'Rex is a mammal'? "
229
+ "Answer entailment, contradiction, or neutral."}]
230
+ inputs = tok.apply_chat_template(
231
+ messages, add_generation_prompt=True,
232
+ return_tensors="pt", return_dict=True,
233
+ ).to(model.device)
234
+
235
+ out = model.generate(**inputs, max_new_tokens=2048, do_sample=False)
236
+ print(tok.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))
237
+ ```
238
+
239
+ `return_dict=True` matters on transformers 5.x, where `apply_chat_template` returns a
240
+ `BatchEncoding` rather than a bare tensor; the above works on both 4.x and 5.x.
241
+
242
+ For the example above, greedy decoding with the bf16 weights (transformers 5.14.1, CPU) produced,
243
+ in 57 tokens and ending on EOS:
244
+
245
+ > Entailment. The statement "All dogs are mammals" implies that any individual dog, such as Rex,
246
+ > must also be a mammal. Therefore, the conclusion "Rex is a mammal" is entailed by the premises.
247
+
248
+ The reported numbers use **greedy decoding** (`do_sample=False`) and a **2048-token** generation
249
+ budget for Track A. The shipped `generation_config.json` carries no sampling defaults, so greedy is
250
+ what you get unless you ask for otherwise. Track A generations average about 517 tokens and 0.9%
251
+ reach the 2048-token cap, so keep the budget at 2048 or more for formal-logic prompts.
252
+
253
+ ### GGUF / llama.cpp
254
+
255
+ Quantized GGUF builds ship in this repository alongside the safetensors weights. The `granite`
256
+ architecture is supported by llama.cpp, and the chat template is embedded in the GGUF metadata, so
257
+ chat mode needs no extra flags.
258
+
259
+ | file | quant | size | bits/weight | notes |
260
+ | ---------------------- | -------- | -------- | ----------- | ------------------------------------------------- |
261
+ | TwIL-LM2-Q4\_K\_M.gguf | Q4\_K\_M | 1.44 GiB | 4.88 | recommended default; runs on CPU |
262
+ | TwIL-LM2-Q5\_K\_M.gguf | Q5\_K\_M | 1.68 GiB | 5.70 | a little more headroom than Q4\_K\_M |
263
+ | TwIL-LM2-Q6\_K.gguf | Q6\_K | 1.94 GiB | 6.57 | close to Q8\_0 quality at about three quarters of the size |
264
+ | TwIL-LM2-Q8\_0.gguf | Q8\_0 | 2.51 GiB | 8.51 | near-lossless, for quality-sensitive use |
265
+ | TwIL-LM2-F16.gguf | F16 | 4.72 GiB | 16.01 | unquantized, for requantization or reference runs |
266
+
267
+ ```bash
268
+ llama-cli -m TwIL-LM2-Q4_K_M.gguf -cnv --temp 0 -n 2048
269
+ ```
270
+
271
+ Pass `--temp 0`, because the evaluation is greedy, and leave the generation budget at 2048 tokens
272
+ or more.
273
+
274
+ The published Track A and Track B numbers were measured on the **bf16** weights through vLLM, not
275
+ on any of these GGUF builds, so expect small deviations — most likely at Q4\_K\_M — that have not
276
+ been quantified here. Note also that F16 is not bit-identical to the bf16 weights: the two formats
277
+ carry the same 16 bits but trade exponent range against mantissa precision.
278
+
279
+ ## How it was built
280
+
281
+ Three stages on top of the base model:
282
+
283
+ 1. **LoRA supervised fine-tuning** on a synthetic formal-logic corpus covering the Track A
284
+ objectives (first-order-logic translation, entailment labelling, semantic parsing, Lean
285
+ formalisation and critique, procedural reasoning, rule induction), using the project's v5
286
+ SFT recipe.
287
+ 2. **WiSE-FT interpolation** toward the pretrained base,
288
+ `W = (1 − λ)·W_base + λ·W_finetuned` with **λ = 0.25** — only a quarter of the fine-tuned delta
289
+ is retained. This conservative λ is the main reason held-out capability survives; the λ = 0.5
290
+ sibling retains twice the delta and shows it in the strict formal-logic lanes and in the loss of
291
+ held-out math.
292
+ 3. **MGPO** — entropy-weighted GRPO reinforcement learning against a programmatic verifier, with
293
+ partial credit for loose matches and token-F1 so that all-fail prompt groups still produce
294
+ gradient. Published checkpoint is **step 1400**.
295
+
296
+ Unlike TwIL-LM3, there is no checkpoint-fusion stage between SFT and WiSE-FT in this arm.
297
+
298
+ ## Limitations and caveats
299
+
300
+ **Strict output form.** Strict MCQ accuracy is 0.0000 and `procedural` accuracy 0.0100;
301
+ `fol_translation` primary score is 0.0000; strict-7 is 0.1214. The model reasons near the required
302
+ form without reliably emitting it. If you need exactly-formatted formal objects, the λ = 0.5
303
+ checkpoint's profile (strict-7 0.2457, MCQ 0.2850) is better, at the cost of held-out math.
304
+
305
+ **Lean critique.** `lean_critic` accuracy is 0.3100, well below TwIL-LM3 (0.6600) and the
306
+ SmolLM2-1.7B-based model (0.5250). This is the weakest of the Track A lanes relative to its peers.
307
+
308
+ **No base comparison.** This card does not include a paired evaluation of
309
+ `granite-3.3-2b-instruct` on the same harness, so it makes no claim about how much the tuning
310
+ improved or regressed the base model on either track.
311
+
312
+ **Perplexity across tokenizers.** `lm_corpus` and `math_corpus` perplexity is a per-token quantity,
313
+ and the columns use different tokenizers (Granite and SmolLM2 have about 49K entries each but
314
+ distinct vocabularies; SmolLM3 has 128K). The perplexity rows are informative within a family and
315
+ only indicative across families.
316
+
317
+ **Result trees.** For Track B, TwIL-LM3 and the SmolLM2-1.7B-based model are scored from the
318
+ results tree that matches their published cards (rope-fixed), and the Granite checkpoints from the
319
+ default tree, with vLLM 0.19.1; the two TwIL macros reproduce their published values exactly
320
+ (0.7339 and 0.4333). Track A figures for all five columns come from GATE 2 reports at n = 200 per
321
+ lane and a 2048-token cap.
322
+
323
+ **Truncation of the comparison arms.** TwIL-LM3 and the SmolLM2-1.7B-based model are formally
324
+ `rankable: false` at 4.4% and 6.9% truncation. A truncated response scores zero regardless of
325
+ reasoning quality, so their Track A figures are understated: wherever this model is ahead of them
326
+ the true margin is smaller, and wherever it is behind, the true deficit is larger.
327
+
328
+ **Scope.** Tuned for formal logic. The Track B suite reported here does not cover code generation
329
+ or tool use, and no claim is made about either. Granite's base tool-calling and document-grounded
330
+ chat-template features are inherited but were not evaluated.
331
+
332
+ **Not a chat model.** It was optimised against automatic verifiers on logic tasks. It has had no
333
+ safety tuning beyond whatever the base model carries, and no instruction-following alignment work.
334
+
335
+ **GGUF builds.** Scores were measured on the bf16 weights only; the quantized builds have not been
336
+ evaluated.
337
+
338
+ ## Evaluation protocol
339
+
340
+ * Track A: `n = 200` per objective, greedy (`temperature = 0`), `max_new_tokens = 2048`, one retry
341
+ at 4096 for truncated rows.
342
+ * Track B: 300 examples per task, greedy, 4096 generation tokens, chat template applied, vLLM 0.19.1
343
+ backend. `musr` is the mean of the murder, object and team splits.
344
+ * Controlled decode bench: one idle H100 per model, 128 shared prompts, `ignore_eos`, hard 512-token
345
+ cap, engine initialisation excluded from the rate.
346
+ * All columns are scored on the same sampled rows within each track.
347
+
348
+ Track B is sampled at 300 examples per dataset for compute reasons. Absolute scores can shift on
349
+ the full sets, but the comparative ordering across models is expected to be stable.
350
+
351
+ ## Relationship to TwIL-LM
352
+
353
+ **TwIL-LM3** ([`webAI-Official/TwIL-LM3`](https://huggingface.co/webAI-Official/TwIL-LM3)) is the
354
+ SmolLM3-3B-based member of the family. It is stronger on held-out benchmarks and on `lean_critic`;
355
+ this model matches its in-domain gate at a smaller size and decodes 1.57x faster.
356
+
357
+ The name **TwIL-LM2** was previously used for a SmolLM2-1.7B-Instruct-based model. It is the
358
+ semantic-parsing specialist in the tables above, and the weakest of these models on held-out
359
+ benchmarks. This repository is a different model — Granite-based, 2.5B — that carries the TwIL-LM2
360
+ name; the SmolLM2-1.7B-based figures above are that model's published numbers, included as a
361
+ reference point and not as this model's results.
362
+
363
+ ## License and attribution
364
+
365
+ Released under the **webAI Non-Commercial License ver. 1.0** — see `LICENSE.md` in this repository.
366
+
367
+ The base model,
368
+ [`ibm-granite/granite-3.3-2b-instruct`](https://huggingface.co/ibm-granite/granite-3.3-2b-instruct),
369
+ is Apache 2.0; its licence text is retained as `apache-2.0-LICENSE.txt` and all credit for the base
370
+ model goes to the IBM Granite team. Apache 2.0 permits distributing derivative works under
371
+ different terms provided attribution is preserved, which is what the pair of licence files in this
372
+ repository does.
apache-2.0-LICENSE.txt ADDED
@@ -0,0 +1,202 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+
2
+ Apache License
3
+ Version 2.0, January 2004
4
+ http://www.apache.org/licenses/
5
+
6
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
7
+
8
+ 1. Definitions.
9
+
10
+ "License" shall mean the terms and conditions for use, reproduction,
11
+ and distribution as defined by Sections 1 through 9 of this document.
12
+
13
+ "Licensor" shall mean the copyright owner or entity authorized by
14
+ the copyright owner that is granting the License.
15
+
16
+ "Legal Entity" shall mean the union of the acting entity and all
17
+ other entities that control, are controlled by, or are under common
18
+ control with that entity. For the purposes of this definition,
19
+ "control" means (i) the power, direct or indirect, to cause the
20
+ direction or management of such entity, whether by contract or
21
+ otherwise, or (ii) ownership of fifty percent (50%) or more of the
22
+ outstanding shares, or (iii) beneficial ownership of such entity.
23
+
24
+ "You" (or "Your") shall mean an individual or Legal Entity
25
+ exercising permissions granted by this License.
26
+
27
+ "Source" form shall mean the preferred form for making modifications,
28
+ including but not limited to software source code, documentation
29
+ source, and configuration files.
30
+
31
+ "Object" form shall mean any form resulting from mechanical
32
+ transformation or translation of a Source form, including but
33
+ not limited to compiled object code, generated documentation,
34
+ and conversions to other media types.
35
+
36
+ "Work" shall mean the work of authorship, whether in Source or
37
+ Object form, made available under the License, as indicated by a
38
+ copyright notice that is included in or attached to the work
39
+ (an example is provided in the Appendix below).
40
+
41
+ "Derivative Works" shall mean any work, whether in Source or Object
42
+ form, that is based on (or derived from) the Work and for which the
43
+ editorial revisions, annotations, elaborations, or other modifications
44
+ represent, as a whole, an original work of authorship. For the purposes
45
+ of this License, Derivative Works shall not include works that remain
46
+ separable from, or merely link (or bind by name) to the interfaces of,
47
+ the Work and Derivative Works thereof.
48
+
49
+ "Contribution" shall mean any work of authorship, including
50
+ the original version of the Work and any modifications or additions
51
+ to that Work or Derivative Works thereof, that is intentionally
52
+ submitted to Licensor for inclusion in the Work by the copyright owner
53
+ or by an individual or Legal Entity authorized to submit on behalf of
54
+ the copyright owner. For the purposes of this definition, "submitted"
55
+ means any form of electronic, verbal, or written communication sent
56
+ to the Licensor or its representatives, including but not limited to
57
+ communication on electronic mailing lists, source code control systems,
58
+ and issue tracking systems that are managed by, or on behalf of, the
59
+ Licensor for the purpose of discussing and improving the Work, but
60
+ excluding communication that is conspicuously marked or otherwise
61
+ designated in writing by the copyright owner as "Not a Contribution."
62
+
63
+ "Contributor" shall mean Licensor and any individual or Legal Entity
64
+ on behalf of whom a Contribution has been received by Licensor and
65
+ subsequently incorporated within the Work.
66
+
67
+ 2. Grant of Copyright License. Subject to the terms and conditions of
68
+ this License, each Contributor hereby grants to You a perpetual,
69
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
70
+ copyright license to reproduce, prepare Derivative Works of,
71
+ publicly display, publicly perform, sublicense, and distribute the
72
+ Work and such Derivative Works in Source or Object form.
73
+
74
+ 3. Grant of Patent License. Subject to the terms and conditions of
75
+ this License, each Contributor hereby grants to You a perpetual,
76
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
77
+ (except as stated in this section) patent license to make, have made,
78
+ use, offer to sell, sell, import, and otherwise transfer the Work,
79
+ where such license applies only to those patent claims licensable
80
+ by such Contributor that are necessarily infringed by their
81
+ Contribution(s) alone or by combination of their Contribution(s)
82
+ with the Work to which such Contribution(s) was submitted. If You
83
+ institute patent litigation against any entity (including a
84
+ cross-claim or counterclaim in a lawsuit) alleging that the Work
85
+ or a Contribution incorporated within the Work constitutes direct
86
+ or contributory patent infringement, then any patent licenses
87
+ granted to You under this License for that Work shall terminate
88
+ as of the date such litigation is filed.
89
+
90
+ 4. Redistribution. You may reproduce and distribute copies of the
91
+ Work or Derivative Works thereof in any medium, with or without
92
+ modifications, and in Source or Object form, provided that You
93
+ meet the following conditions:
94
+
95
+ (a) You must give any other recipients of the Work or
96
+ Derivative Works a copy of this License; and
97
+
98
+ (b) You must cause any modified files to carry prominent notices
99
+ stating that You changed the files; and
100
+
101
+ (c) You must retain, in the Source form of any Derivative Works
102
+ that You distribute, all copyright, patent, trademark, and
103
+ attribution notices from the Source form of the Work,
104
+ excluding those notices that do not pertain to any part of
105
+ the Derivative Works; and
106
+
107
+ (d) If the Work includes a "NOTICE" text file as part of its
108
+ distribution, then any Derivative Works that You distribute must
109
+ include a readable copy of the attribution notices contained
110
+ within such NOTICE file, excluding those notices that do not
111
+ pertain to any part of the Derivative Works, in at least one
112
+ of the following places: within a NOTICE text file distributed
113
+ as part of the Derivative Works; within the Source form or
114
+ documentation, if provided along with the Derivative Works; or,
115
+ within a display generated by the Derivative Works, if and
116
+ wherever such third-party notices normally appear. The contents
117
+ of the NOTICE file are for informational purposes only and
118
+ do not modify the License. You may add Your own attribution
119
+ notices within Derivative Works that You distribute, alongside
120
+ or as an addendum to the NOTICE text from the Work, provided
121
+ that such additional attribution notices cannot be construed
122
+ as modifying the License.
123
+
124
+ You may add Your own copyright statement to Your modifications and
125
+ may provide additional or different license terms and conditions
126
+ for use, reproduction, or distribution of Your modifications, or
127
+ for any such Derivative Works as a whole, provided Your use,
128
+ reproduction, and distribution of the Work otherwise complies with
129
+ the conditions stated in this License.
130
+
131
+ 5. Submission of Contributions. Unless You explicitly state otherwise,
132
+ any Contribution intentionally submitted for inclusion in the Work
133
+ by You to the Licensor shall be under the terms and conditions of
134
+ this License, without any additional terms or conditions.
135
+ Notwithstanding the above, nothing herein shall supersede or modify
136
+ the terms of any separate license agreement you may have executed
137
+ with Licensor regarding such Contributions.
138
+
139
+ 6. Trademarks. This License does not grant permission to use the trade
140
+ names, trademarks, service marks, or product names of the Licensor,
141
+ except as required for reasonable and customary use in describing the
142
+ origin of the Work and reproducing the content of the NOTICE file.
143
+
144
+ 7. Disclaimer of Warranty. Unless required by applicable law or
145
+ agreed to in writing, Licensor provides the Work (and each
146
+ Contributor provides its Contributions) on an "AS IS" BASIS,
147
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
148
+ implied, including, without limitation, any warranties or conditions
149
+ of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
150
+ PARTICULAR PURPOSE. You are solely responsible for determining the
151
+ appropriateness of using or redistributing the Work and assume any
152
+ risks associated with Your exercise of permissions under this License.
153
+
154
+ 8. Limitation of Liability. In no event and under no legal theory,
155
+ whether in tort (including negligence), contract, or otherwise,
156
+ unless required by applicable law (such as deliberate and grossly
157
+ negligent acts) or agreed to in writing, shall any Contributor be
158
+ liable to You for damages, including any direct, indirect, special,
159
+ incidental, or consequential damages of any character arising as a
160
+ result of this License or out of the use or inability to use the
161
+ Work (including but not limited to damages for loss of goodwill,
162
+ work stoppage, computer failure or malfunction, or any and all
163
+ other commercial damages or losses), even if such Contributor
164
+ has been advised of the possibility of such damages.
165
+
166
+ 9. Accepting Warranty or Additional Liability. While redistributing
167
+ the Work or Derivative Works thereof, You may choose to offer,
168
+ and charge a fee for, acceptance of support, warranty, indemnity,
169
+ or other liability obligations and/or rights consistent with this
170
+ License. However, in accepting such obligations, You may act only
171
+ on Your own behalf and on Your sole responsibility, not on behalf
172
+ of any other Contributor, and only if You agree to indemnify,
173
+ defend, and hold each Contributor harmless for any liability
174
+ incurred by, or claims asserted against, such Contributor by reason
175
+ of your accepting any such warranty or additional liability.
176
+
177
+ END OF TERMS AND CONDITIONS
178
+
179
+ APPENDIX: How to apply the Apache License to your work.
180
+
181
+ To apply the Apache License to your work, attach the following
182
+ boilerplate notice, with the fields enclosed by brackets "[]"
183
+ replaced with your own identifying information. (Don't include
184
+ the brackets!) The text should be enclosed in the appropriate
185
+ comment syntax for the file format. We also recommend that a
186
+ file or class name and description of purpose be included on the
187
+ same "printed page" as the copyright notice for easier
188
+ identification within third-party archives.
189
+
190
+ Copyright [yyyy] [name of copyright owner]
191
+
192
+ Licensed under the Apache License, Version 2.0 (the "License");
193
+ you may not use this file except in compliance with the License.
194
+ You may obtain a copy of the License at
195
+
196
+ http://www.apache.org/licenses/LICENSE-2.0
197
+
198
+ Unless required by applicable law or agreed to in writing, software
199
+ distributed under the License is distributed on an "AS IS" BASIS,
200
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
201
+ See the License for the specific language governing permissions and
202
+ limitations under the License.
chat_template.jinja ADDED
@@ -0,0 +1,62 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {# Alias tools -> available_tools #}
2
+ {%- if tools and not available_tools -%}
3
+ {%- set available_tools = tools -%}
4
+ {%- endif -%}
5
+ {%- if messages[0]['role'] == 'system' %}
6
+ {%- set system_message = messages[0]['content'] %}
7
+ {%- set loop_messages = messages[1:] %}
8
+ {%- else %}
9
+ {%- set system_message = "Knowledge Cutoff Date: April 2024.
10
+ Today's Date: " + strftime_now('%B %d, %Y') + ".
11
+ You are Granite, developed by IBM." %}
12
+ {%- if available_tools and documents %}
13
+ {%- set system_message = system_message + " You are a helpful assistant with access to the following tools. When a tool is required to answer the user's query, respond only with <|tool_call|> followed by a JSON list of tools used. If a tool does not exist in the provided list of tools, notify the user that you do not have the ability to fulfill the request.
14
+ Write the response to the user's input by strictly aligning with the facts in the provided documents. If the information needed to answer the question is not available in the documents, inform the user that the question cannot be answered based on the available data." %}
15
+ {%- elif available_tools %}
16
+ {%- set system_message = system_message + " You are a helpful assistant with access to the following tools. When a tool is required to answer the user's query, respond only with <|tool_call|> followed by a JSON list of tools used. If a tool does not exist in the provided list of tools, notify the user that you do not have the ability to fulfill the request." %}
17
+ {%- elif documents %}
18
+ {%- set system_message = system_message + " Write the response to the user's input by strictly aligning with the facts in the provided documents. If the information needed to answer the question is not available in the documents, inform the user that the question cannot be answered based on the available data." %}
19
+ {%- elif thinking %}
20
+ {%- set system_message = system_message + " You are a helpful AI assistant.
21
+ Respond to every user query in a comprehensive and detailed way. You can write down your thoughts and reasoning process before responding. In the thought process, engage in a comprehensive cycle of analysis, summarization, exploration, reassessment, reflection, backtracing, and iteration to develop well-considered thinking process. In the response section, based on various attempts, explorations, and reflections from the thoughts section, systematically present the final solution that you deem correct. The response should summarize the thought process. Write your thoughts between <think></think> and write your response between <response></response> for each user query." %}
22
+ {%- else %}
23
+ {%- set system_message = system_message + " You are a helpful AI assistant." %}
24
+ {%- endif %}
25
+ {%- if 'citations' in controls and documents %}
26
+ {%- set system_message = system_message + '
27
+ Use the symbols <|start_of_cite|> and <|end_of_cite|> to indicate when a fact comes from a document in the search result, e.g <|start_of_cite|> {document_id: 1}my fact <|end_of_cite|> for a fact from document 1. Afterwards, list all the citations with their corresponding documents in an ordered list.' %}
28
+ {%- endif %}
29
+ {%- if 'hallucinations' in controls and documents %}
30
+ {%- set system_message = system_message + '
31
+ Finally, after the response is written, include a numbered list of sentences from the response with a corresponding risk value that are hallucinated and not based in the documents.' %}
32
+ {%- endif %}
33
+ {%- set loop_messages = messages %}
34
+ {%- endif %}
35
+ {{- '<|start_of_role|>system<|end_of_role|>' + system_message + '<|end_of_text|>
36
+ ' }}
37
+ {%- if available_tools %}
38
+ {{- '<|start_of_role|>available_tools<|end_of_role|>' }}
39
+ {{- available_tools | tojson(indent=4) }}
40
+ {{- '<|end_of_text|>
41
+ ' }}
42
+ {%- endif %}
43
+ {%- if documents %}
44
+ {%- for document in documents %}
45
+ {{- '<|start_of_role|>document {"document_id": "' + document['doc_id'] | string + '"}<|end_of_role|>
46
+ ' }}
47
+ {{- document['text'] }}
48
+ {{- '<|end_of_text|>
49
+ ' }}
50
+ {%- endfor %}
51
+ {%- endif %}
52
+ {%- for message in loop_messages %}
53
+ {{- '<|start_of_role|>' + message['role'] + '<|end_of_role|>' + message['content'] + '<|end_of_text|>
54
+ ' }}
55
+ {%- if loop.last and add_generation_prompt %}
56
+ {{- '<|start_of_role|>assistant' }}
57
+ {%- if controls %}
58
+ {{- ' ' + controls | tojson()}}
59
+ {%- endif %}
60
+ {{- '<|end_of_role|>' }}
61
+ {%- endif %}
62
+ {%- endfor %}
config.json ADDED
@@ -0,0 +1,34 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "GraniteForCausalLM"
4
+ ],
5
+ "attention_bias": false,
6
+ "attention_dropout": 0.0,
7
+ "attention_multiplier": 0.015625,
8
+ "bos_token_id": 0,
9
+ "dtype": "bfloat16",
10
+ "embedding_multiplier": 12.0,
11
+ "eos_token_id": 0,
12
+ "hidden_act": "silu",
13
+ "hidden_size": 2048,
14
+ "initializer_range": 0.02,
15
+ "intermediate_size": 8192,
16
+ "logits_scaling": 8.0,
17
+ "max_position_embeddings": 131072,
18
+ "mlp_bias": false,
19
+ "model_type": "granite",
20
+ "num_attention_heads": 32,
21
+ "num_hidden_layers": 40,
22
+ "num_key_value_heads": 8,
23
+ "pad_token_id": 0,
24
+ "residual_multiplier": 0.22,
25
+ "rms_norm_eps": 1e-05,
26
+ "rope_parameters": {
27
+ "rope_theta": 10000000.0,
28
+ "rope_type": "default"
29
+ },
30
+ "tie_word_embeddings": true,
31
+ "transformers_version": "5.5.0",
32
+ "use_cache": true,
33
+ "vocab_size": 49159
34
+ }
generation_config.json ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "bos_token_id": 0,
4
+ "eos_token_id": 0,
5
+ "pad_token_id": 0,
6
+ "transformers_version": "5.5.0"
7
+ }
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json ADDED
@@ -0,0 +1,15 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "backend": "tokenizers",
4
+ "bos_token": "<|end_of_text|>",
5
+ "clean_up_tokenization_spaces": true,
6
+ "eos_token": "<|end_of_text|>",
7
+ "errors": "replace",
8
+ "is_local": true,
9
+ "model_max_length": 9223372036854775807,
10
+ "pad_token": "<|end_of_text|>",
11
+ "padding_side": "left",
12
+ "tokenizer_class": "GPT2Tokenizer",
13
+ "unk_token": "<|end_of_text|>",
14
+ "vocab_size": 49152
15
+ }