Text Generation
Transformers
Safetensors
GGUF
English
granite
formal-logic
reasoning
lora
model-merging
wise-ft
reinforcement-learning
grpo
twil-lm
conversational
Instructions to use webAI-Official/TwIL-LM2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use webAI-Official/TwIL-LM2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="webAI-Official/TwIL-LM2") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("webAI-Official/TwIL-LM2") model = AutoModelForCausalLM.from_pretrained("webAI-Official/TwIL-LM2", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use webAI-Official/TwIL-LM2 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf webAI-Official/TwIL-LM2:Q4_K_M # Run inference directly in the terminal: llama cli -hf webAI-Official/TwIL-LM2:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf webAI-Official/TwIL-LM2:Q4_K_M # Run inference directly in the terminal: llama cli -hf webAI-Official/TwIL-LM2:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf webAI-Official/TwIL-LM2:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf webAI-Official/TwIL-LM2:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf webAI-Official/TwIL-LM2:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf webAI-Official/TwIL-LM2:Q4_K_M
Use Docker
docker model run hf.co/webAI-Official/TwIL-LM2:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use webAI-Official/TwIL-LM2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "webAI-Official/TwIL-LM2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "webAI-Official/TwIL-LM2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/webAI-Official/TwIL-LM2:Q4_K_M
- SGLang
How to use webAI-Official/TwIL-LM2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "webAI-Official/TwIL-LM2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "webAI-Official/TwIL-LM2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "webAI-Official/TwIL-LM2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "webAI-Official/TwIL-LM2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use webAI-Official/TwIL-LM2 with Ollama:
ollama run hf.co/webAI-Official/TwIL-LM2:Q4_K_M
- Unsloth Desktop
- Pi
How to use webAI-Official/TwIL-LM2 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webAI-Official/TwIL-LM2:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "webAI-Official/TwIL-LM2:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use webAI-Official/TwIL-LM2 with Docker Model Runner:
docker model run hf.co/webAI-Official/TwIL-LM2:Q4_K_M
- Lemonade
How to use webAI-Official/TwIL-LM2 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull webAI-Official/TwIL-LM2:Q4_K_M
Run and chat with the model
lemonade run user.TwIL-LM2-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use webAI-Official/TwIL-LM2 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webAI-Official/TwIL-LM2:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default webAI-Official/TwIL-LM2:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use webAI-Official/TwIL-LM2 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webAI-Official/TwIL-LM2:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "webAI-Official/TwIL-LM2:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
|
Download README.md from webAI-Official/TwIL-LM2: direct link, hf CLI and curl.
- Browser
- Download file 25.4 kB
-
https://huggingface.co/webAI-Official/TwIL-LM2/resolve/main/README.md
- Command line
-
hf download hf://webAI-Official/TwIL-LM2/README.md
-
curl -L -o README.md https://huggingface.co/webAI-Official/TwIL-LM2/resolve/main/README.md
25.4 kB
| language: | |
| - en | |
| library_name: transformers | |
| pipeline_tag: text-generation | |
| base_model: ibm-granite/granite-3.3-2b-instruct | |
| license: other | |
| license_name: webai-non-commercial-license-ver.-1.0 | |
| license_link: https://huggingface.co/webAI-Official/TwIL-LM2/blob/main/LICENSE.md | |
| tags: | |
| - granite | |
| - formal-logic | |
| - reasoning | |
| - lora | |
| - model-merging | |
| - wise-ft | |
| - reinforcement-learning | |
| - grpo | |
| - twil-lm | |
| - gguf | |
| # TwIL-LM2 | |
|  | |
| A 2.5B reasoning model for **formal logic** tasks, built from | |
| [`ibm-granite/granite-3.3-2b-instruct`](https://huggingface.co/ibm-granite/granite-3.3-2b-instruct) | |
| through LoRA supervised fine-tuning, WiSE-FT weight interpolation (λ = 0.25) and entropy-weighted | |
| GRPO reinforcement learning (MGPO, step 1400). | |
| On the in-domain macro gate it scores **0.4178** — fourth of the twelve models with a reported | |
| gate, behind only TwIL-LM3 (0.4218), Qwen3-8B (0.5336) and Gemma-4-26B-A4B-it (0.6344). It is ahead | |
| of the other eight, including LFM2.5-8B-A1B (0.3757), which has about three times its total | |
| parameters, and Granite-4.1-3B (0.3435). It also decodes **1.57x faster than TwIL-LM3** under | |
| identical forced work. | |
| It is not a strong strict-output or general-benchmark model. Its strict-7 score (0.1214) ranks | |
| eleventh of twelve, its `lean_critic` accuracy (0.3100) is the lowest of all thirteen models | |
| compared, and its 10-dataset held-out macro (0.6759) is below every model of comparable size in the | |
| comparison. No paired evaluation against its own base is included, so this card makes no claim | |
| about what the fine-tune did to held-out capability. See [Results](#results) and | |
| [Limitations](#limitations-and-caveats). | |
| ## Highlights | |
| * **Fourth on the in-domain gate.** 0.4178 against 0.4218 for TwIL-LM3, 0.3927 for the | |
| SmolLM2-1.7B-based TwIL-LM2 and 0.3757 for LFM2.5-8B-A1B. TwIL-LM3's figure is understated by | |
| truncation (4.4% of its Track A rows hit the token cap), so read the gap to it as approximate. | |
| * **Close to TwIL-LM3 at a smaller size.** 0.004 behind on the gate at 2.53B against 3.08B | |
| parameters — about 18% fewer. | |
| * **Formal-logic lanes where it is competitive.** `lean_formalize` token-F1 0.5159 is fourth of | |
| thirteen, ahead of Qwen3-8B (0.4022), Gemma-4-26B-A4B-it (0.4107) and LFM2.5-8B-A1B (0.4655). | |
| `rule_induction` 0.3292 is ahead of TwIL-LM3 (0.3192) and of Granite-4.1-3B (0.2476). | |
| * **Lowest `lm_corpus` perplexity in the comparison** (1.9808), and third on `math_corpus` | |
| (3.3073). Read these with the tokenizer caveat under [Limitations](#limitations-and-caveats). | |
| * **Fast.** 21,369 decode tokens/s on one H100 in a controlled bench — 1.57x TwIL-LM3 and 1.22x the | |
| SmolLM2-1.7B-based model — because Granite's architecture decodes quickly, not because it answers | |
| short. | |
| * **A cleaner measurement.** Only 0.9% of Track A generations hit the 2048-token cap, under the 2% | |
| threshold our protocol requires to mark a comparison `rankable`. | |
| * **Runs anywhere.** 2.53B parameters in bf16 (4.72 GiB), with a Q4\_K\_M GGUF at 1.44 GiB for CPU. | |
| Where it is weak: strict-7 (0.1214, eleventh of twelve), strict MCQ accuracy (0.0000), `lean_critic` | |
| (0.3100, last of thirteen), and held-out benchmarks (10-dataset macro 0.6759, eleventh of thirteen). | |
| It is not a general assistant. | |
| ## Model Details | |
| | Property | Value | | |
| | ------------------------- | ----------------------------------------------------------------------------------------------- | | |
| | Model ID | `webAI-Official/TwIL-LM2` | | |
| | Base model | [`ibm-granite/granite-3.3-2b-instruct`](https://huggingface.co/ibm-granite/granite-3.3-2b-instruct) | | |
| | Total parameters | 2.53B (2,533,539,840; tied input/output embeddings) | | |
| | Architecture | Granite decoder-only transformer; 40 layers, hidden size 2048, 32 attention heads, 8 KV heads | | |
| | Input / output | Text / text | | |
| | Language | English | | |
| | Vocabulary size | 49,159 embedding rows | | |
| | Context window | 131,072 tokens (inherited from the base; see note below) | | |
| | Checkpoint precision | bfloat16 (4.72 GiB), plus Q4\_K\_M / Q5\_K\_M / Q6\_K / Q8\_0 / F16 GGUF builds | | |
| | Post-training | LoRA SFT → WiSE-FT (λ = 0.25) → MGPO reinforcement learning (step 1400) | | |
| | Chat template | Granite chat template (`<\|start_of_role\|>…<\|end_of_role\|>`), EOS `<\|end_of_text\|>` | | |
| | Reasoning format | Answers directly under the default chat template; no `<think>` block was observed | | |
| | Evaluated decoding | Greedy; 2048 new tokens (Track A), 4096 new tokens (Track B) | | |
| | Specialisation | Formal logic: FOL translation, entailment, semantic parsing, Lean formalisation and critique | | |
| | License | webAI Non-Commercial License ver. 1.0 | | |
| The base model's 131,072-token context is carried through unchanged, but every score on this card | |
| was measured with generation budgets of 2048 (Track A) or 4096 (Track B) tokens. Longer contexts are | |
| inherited rather than validated here. Granite's optional `thinking=True` chat-template mode is also | |
| inherited from the base and was not evaluated for this card. | |
| ## Results | |
| Every model below was scored through the same harness, prompts and decoding settings described | |
| under [Evaluation protocol](#evaluation-protocol). The columns are this model; the two other TwIL | |
| models — [TwIL-LM3](https://huggingface.co/webAI-Official/TwIL-LM3) (SmolLM3-3B) and the earlier, | |
| SmolLM2-1.7B-Instruct-based TwIL-LM2 — with their published figures; the two SmolLM bases; and the | |
| external models reported alongside them on those cards. | |
| ### Track A — in-domain formal logic | |
| | lane / metric | TwIL-LM2 (this model) | TwIL-LM3 | TwIL-LM2 (SmolLM2-1.7B) | SmolLM3-3B base | SmolLM2-1.7B base | LFM2.5-1.2B-Thinking | LFM2-2.6B | Llama-3.2-3B | Granite-4.1-3B | LFM2.5-8B-A1B | Qwen3-8B | Gemma-4-26B-A4B-it | gpt-oss-120b ‡ | | |
| |---|---|---|---|---|---|---|---|---|---|---|---|---|---| | |
| | parameters | 2.53B | 3.08B | 1.7B | 3.08B | 1.7B | 1.2B | 2.6B | 3B | 3B | 8B (1B active) | 8B | 26B (4B active) | 120B | | |
| | `lean_formalize` token-F1 | 0.5159 | 0.5869 | 0.6199 | 0.4347 | 0.1087 | 0.1890 | 0.1321 | 0.3690 | 0.2652 | 0.4655 | 0.4022 | 0.4107 | **0.6306** | | |
| | `rule_induction` derivation | 0.3292 | 0.3192 | 0.5136 | 0.1029 | 0.1350 | 0.0837 | 0.0615 | 0.0825 | 0.2476 | 0.1936 | 0.3680 | **0.7319** | 0.6518 | | |
| | `entailment_label` accuracy | 0.5300 | 0.5750 | 0.5850 | 0.3750 | 0.2450 | 0.4700 | 0.4700 | 0.3300 | 0.4900 | 0.5400 | 0.5800 | 0.6200 | **0.7750** | | |
| | `mcq_answer` accuracy | 0.0000 | 0.1100 | **0.1600** | 0.0000 | 0.0000 | 0.0000 | 0.0150 | 0.0000 | 0.0100 | 0.0750 | 0.0000 | 0.0200 | 0.0700 | | |
| | `semantic_parse` token-F1 | 0.4013 | 0.4416 | **0.8428** | 0.4149 | 0.2155 | 0.4439 | 0.3665 | 0.3102 | 0.1953 | 0.3778 | 0.4257 | 0.4567 | 0.4331 | | |
| | `lean_critic` accuracy | 0.3100 | 0.6600 | 0.5250 | 0.6500 | 0.4950 | 0.5450 | 0.5900 | 0.5300 | 0.5150 | 0.5500 | **0.7950** | 0.7500 | 0.5550 | | |
| | `lm_corpus` perplexity ↓ | **1.9808** | 2.8972 | 2.2981 | 3.1818 | 2.5845 | 5.0065 | 4.3815 | 2.8478 | 2.4736 | 4.9472 | 2.5440 | 16.1145 | 912.23 § | | |
| | `math_corpus` perplexity ↓ | 3.3073 | 3.8229 | **3.0390** | 4.0685 | 3.2670 | 7.7402 | 6.7472 | 4.7531 | 4.1162 | 8.3323 | 4.0083 | 59.7838 | 1045.63 § | | |
| | **macro gate** | 0.4178 | 0.4218 | 0.3927 | 0.3466 † | 0.2590 † | 0.3067 | 0.3473 | 0.2925 | 0.3435 | 0.3757 | 0.5336 | **0.6344** | — | | |
| | **strict-7** | 0.1214 | 0.1971 | **0.2386** | 0.1493 | 0.1071 | 0.1450 | 0.1579 | 0.1229 | 0.1507 | 0.1714 | 0.2093 | 0.2050 | — | | |
| | macro\_primary | 0.4400 | 0.4475 | 0.3625 | 0.4075 | 0.2900 | 0.3625 | 0.4188 | 0.3450 | 0.3675 | 0.4213 | 0.5750 | **0.6100** | — | | |
| † The base columns come from the external-comparison run rather than the paired base-versus-TwIL | |
| run, hence SmolLM3-3B 0.3466 here against 0.3356 in its paired run and SmolLM2-1.7B 0.2590 against | |
| 0.2630. The paired run is the correct basis for an improvement claim. | |
| ‡ **gpt-oss-120b** runs MXFP4 weights at tensor-parallel 2 — quantized and multi-GPU, so it is not | |
| directly comparable to the single-GPU bf16 columns. Its `procedural` lane and the loose-match | |
| scorings were not collected, so its gate, strict-7 and macro\_primary cannot be computed; the — | |
| cells mean that, not zero. | |
| § The 120B's perplexities are three orders of magnitude off every other model because its harmony | |
| response format and tokenizer make the corpus lanes score a different quantity. They are reported | |
| for completeness and excluded from the perplexity ranking. | |
| Throughput and generation-length rows are left out of this table: the source cards report them | |
| from different runs, so they cannot be put in one column. The controlled decode bench under | |
| [Speed](#speed) is the like-for-like speed comparison. `average, 6 lanes`, `macro gate`, | |
| `macro_primary` and `strict-7` are the harness aggregates defined on the | |
| [TwIL-LM3 model card](https://huggingface.co/webAI-Official/TwIL-LM3); this card does not redefine | |
| them. They are not interchangeable, and the ordering changes between them. | |
| **Reading it.** Among the three TwIL models this is second on the gate — 0.4178, between 0.4218 for | |
| TwIL-LM3 and 0.3927 for the SmolLM2-1.7B-based model — but it is well behind both on strict-7 | |
| (0.1214 against 0.1971 and 0.2386). The gate credits loose matches on some lanes and | |
| strict-7 credits none, so the near-parity is on the gate, not on strict scoring. | |
| Against the wider set it is fourth on the gate and on `macro_primary`, behind TwIL-LM3, Qwen3-8B | |
| and Gemma-4-26B-A4B-it, and ahead of everything else with a reported gate. It is fourth of thirteen | |
| on `lean_formalize`, fifth on `rule_induction` and seventh on `entailment_label`. It is eighth on | |
| `semantic_parse` (0.4013, against 0.8428 for the SmolLM2-1.7B-based model) and last on | |
| `lean_critic` (0.3100, against 0.6600 for TwIL-LM3 and 0.7950 for Qwen3-8B). | |
| Strict MCQ accuracy is 0.0000. That is not unique to this model — Qwen3-8B, Llama-3.2-3B, | |
| LFM2.5-1.2B-Thinking and both SmolLM bases also score 0.0000, because they answer the lane without | |
| emitting the requested form — but it contributes to a strict-7 that is eleventh of twelve; only | |
| SmolLM2-1.7B base (0.1071) is lower. On its Track A `procedural` lane it scores 0.0100 and on | |
| `fol_translation` 0.0000. | |
| It does not beat the two largest models with reported gates: Qwen3-8B leads it 0.5336 to 0.4178 and | |
| Gemma-4-26B-A4B-it 0.6344 to 0.4178. Much of the Qwen gap is loose-match credit rather than | |
| capability (Qwen3-8B answers MCQ correctly but almost never in the requested format), though Gemma | |
| also genuinely leads on rule induction (0.7319), which no scoring convention explains away. | |
| ### Track B — held-out benchmarks | |
| Nothing in this suite was trained on. All models are scored by the same aggregation over 300 | |
| randomly sampled, model-identical examples per dataset. | |
| | dataset | TwIL-LM2 (this model) | TwIL-LM3 | TwIL-LM2 (SmolLM2-1.7B) | SmolLM3-3B base | SmolLM2-1.7B base | LFM2.5-1.2B-Thinking | LFM2-2.6B | Llama-3.2-3B | Granite-4.1-3B | LFM2.5-8B-A1B | Qwen3-8B | Gemma-4-26B-A4B-it | gpt-oss-120b ‡ | | |
| |---|---|---|---|---|---|---|---|---|---|---|---|---|---| | |
| | `gsm8k` | 0.7867 | 0.8733 | 0.4633 | 0.8833 | 0.4800 | 0.8400 | 0.8767 | 0.8300 | 0.9100 | 0.9133 | 0.9567 | 0.9733 | **0.9767** | | |
| | `svamp` | 0.8200 | 0.8500 | 0.3833 | 0.8567 | 0.4867 | 0.9167 | 0.9000 | 0.8200 | 0.9000 | 0.9133 | 0.9367 | **0.9500** | 0.9400 | | |
| | `gsm_symbolic` | 0.7233 | 0.7567 | 0.2600 | 0.7633 | 0.2200 | 0.6867 | 0.9767 | 0.8067 | 0.9533 | 0.9267 | 0.8133 | **0.9967** | 0.8467 | | |
| | `arc_cot` | 0.7367 | 0.8467 | 0.5200 | 0.8400 | 0.5100 | 0.8300 | 0.8667 | 0.7967 | 0.8633 | 0.9033 | 0.9633 | **0.9767** | 0.9667 | | |
| | `logicbench` | 0.6900 | 0.7167 | 0.5400 | 0.6467 | 0.5067 | 0.6700 | 0.6267 | 0.5733 | 0.7367 | 0.7200 | 0.8567 | **0.8667** | 0.8533 | | |
| | `strategyqa` | 0.6933 | 0.6500 | 0.5900 | 0.6333 | 0.6000 | 0.5933 | 0.6433 | 0.6533 | 0.6333 | 0.6667 | 0.7400 | 0.7700 | **0.7867** | | |
| | `drop` | 0.5833 | 0.7467 | 0.4367 | 0.7000 | 0.4233 | 0.6667 | 0.6900 | 0.6733 | 0.7600 | 0.6633 | **0.8833** | 0.7933 | 0.8500 | | |
| | `csqa` | 0.6900 | 0.7367 | 0.4333 | 0.7067 | 0.3967 | 0.6100 | 0.7433 | 0.7500 | 0.7633 | 0.7700 | **0.8633** | **0.8633** | 0.8367 | | |
| | `musr` | 0.4957 | 0.4957 | 0.3131 | 0.4997 | 0.4223 | 0.5227 | 0.4867 | 0.4932 | 0.5669 | 0.5703 | 0.6301 | 0.6369 | **0.6852** | | |
| | `mmlu_redux` | 0.5400 | 0.6667 | 0.3933 | 0.6633 | 0.4100 | 0.6400 | 0.7133 | 0.6000 | 0.6800 | 0.8367 | 0.8500 | **0.9633** | 0.9467 | | |
| | `ifeval` | — | 0.6433 | 0.4300 | 0.6767 | 0.4700 | 0.8233 | 0.7300 | 0.7167 | 0.7967 | **0.8900** | 0.8400 | 0.8733 | 0.7900 | | |
| | `rudas_ood` | — | 0.0365 | 0.0289 | 0.0209 | 0.0128 | 0.0089 | 0.0017 | 0.0733 | 0.0355 | 0.0061 | 0.0468 | **0.1547** | 0.0000 ¶ | | |
| | `bbh_logic` | — | 0.6633 | 0.2373 | 0.6667 | 0.2447 | 0.5327 | 0.5713 | 0.5333 | 0.7727 | 0.7700 | 0.6367 | 0.9940 | **0.9980** | | |
| | `math500` | — | 0.6900 | 0.2100 | 0.7000 | 0.1900 | 0.6867 | 0.7133 | 0.4233 | 0.6067 | 0.7800 | 0.6100 | **0.9000** | 0.8433 | | |
| | **macro (10 CoT datasets)** | 0.6759 | 0.7339 | 0.4333 | 0.7193 | 0.4456 | 0.6976 | 0.7523 | 0.6997 | 0.7767 | 0.7884 | 0.8493 | **0.8790** | 0.8689 | | |
| | **macro (all 14)** | — | 0.6694 | 0.3742 | 0.6612 | 0.3838 | 0.6448 | 0.6814 | 0.6245 | 0.7127 | 0.7378 | 0.7591 | **0.8366** | 0.8086 | | |
| ‡ **gpt-oss-120b**: MXFP4 weights, tensor-parallel 2 — quantized and multi-GPU, so not directly | |
| comparable to the single-GPU bf16 columns. ¶ 74% of its `rudas_ood` generations hit the length cap, | |
| so that cell is a truncation artefact rather than a measured score and is excluded from the bolding. | |
| The — cells for this model are lanes that were **not run** for this checkpoint (`ifeval`, | |
| `rudas_ood`, `bbh_logic`, `math500`, and therefore the 14-dataset macro); no instruction-following | |
| or 14-dataset result is claimed. Qwen3-8B `svamp` is 0.9367 as on the earlier TwIL-LM2 card; the | |
| TwIL-LM3 card lists 0.9400, which does not reproduce that card's own Qwen3-8B macros. | |
| **Reading it.** On the 10-dataset macro this model scores 0.6759, eleventh of thirteen. It is ahead | |
| of only the two SmolLM2-1.7B entries (0.4333 and 0.4456) and behind every model of comparable | |
| size: LFM2-2.6B (0.7523), SmolLM3-3B base (0.7193), Llama-3.2-3B (0.6997), LFM2.5-1.2B-Thinking | |
| (0.6976) and Granite-4.1-3B (0.7767). TwIL-LM3, the strongest TwIL model here, scores 0.7339. | |
| It is strongest on `strategyqa` (0.6933, fourth of thirteen, ahead of TwIL-LM3 at 0.6500, Llama-3.2-3B | |
| and Granite-4.1-3B) and holds a mid-table position on `logicbench` (0.6900, seventh of thirteen). It | |
| ranks eleventh of thirteen on `gsm8k` (0.7867), `arc_cot` (0.7367), `drop` (0.5833) and | |
| `mmlu_redux` (0.5400). | |
| The comparison that would say whether the fine-tune moved held-out performance is the one against | |
| `granite-3.3-2b-instruct` itself, and it is not in this card. Granite-4.1-3B, in the tables above, is | |
| a different and larger model, not this model's base. | |
| ### Speed | |
| Harness tokens-per-second and answers-per-second are confounded by how much each model writes. For | |
| an actual speed comparison, each model ran alone on one idle H100 (vLLM 0.19.1, torch 2.10.0+cu128, | |
| transformers 5.15.0) over the same 128 prompts with `ignore_eos` and a hard 512-token cap, so every | |
| model emitted exactly 65,536 output tokens. | |
| | controlled decode bench | **This model** | TwIL-LM3 | TwIL-LM2 (SmolLM2-1.7B) | | |
| | -------------------------------- | -------------- | -------- | ----------------------- | | |
| | decode tokens/s | **21,369** | 13,623 | 17,542 | | |
| | 512-token completions/s | **41.7** | 26.6 | 34.2 | | |
| | decode wall seconds (65,536 tok) | **3.07** | 4.81 | 3.74 | | |
| | relative to this model | 1.00x | 0.64x | 0.82x | | |
| Decode speed here is set by the architecture, not by anything the fine-tune changed. The shared | |
| prompt file is Granite-templated, so prefill differs slightly by tokenizer (30.7K tokens for | |
| Granite and SmolLM3, 38.2K for SmolLM2); it is a small share of the forced output and, if anything, | |
| slightly handicaps the SmolLM2-based model. The other models in the tables above were not run in | |
| this bench. | |
| ## Usage | |
| ```python | |
| import torch | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| model_id = "webAI-Official/TwIL-LM2" | |
| tok = AutoTokenizer.from_pretrained(model_id) | |
| model = AutoModelForCausalLM.from_pretrained( | |
| model_id, torch_dtype=torch.bfloat16, device_map="auto" | |
| ) | |
| messages = [{"role": "user", "content": | |
| "Does 'All dogs are mammals. Rex is a dog.' entail 'Rex is a mammal'? " | |
| "Answer entailment, contradiction, or neutral."}] | |
| inputs = tok.apply_chat_template( | |
| messages, add_generation_prompt=True, | |
| return_tensors="pt", return_dict=True, | |
| ).to(model.device) | |
| out = model.generate(**inputs, max_new_tokens=2048, do_sample=False) | |
| print(tok.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True)) | |
| ``` | |
| `return_dict=True` matters on transformers 5.x, where `apply_chat_template` returns a | |
| `BatchEncoding` rather than a bare tensor; the above works on both 4.x and 5.x. | |
| For the example above, greedy decoding with the bf16 weights (transformers 5.14.1, CPU) produced, | |
| in 57 tokens and ending on EOS: | |
| > Entailment. The statement "All dogs are mammals" implies that any individual dog, such as Rex, | |
| > must also be a mammal. Therefore, the conclusion "Rex is a mammal" is entailed by the premises. | |
| The reported numbers use **greedy decoding** (`do_sample=False`) and a **2048-token** generation | |
| budget for Track A. The shipped `generation_config.json` carries no sampling defaults, so greedy is | |
| what you get unless you ask for otherwise. Track A generations average about 517 tokens and 0.9% | |
| reach the 2048-token cap, so keep the budget at 2048 or more for formal-logic prompts. | |
| ### GGUF / llama.cpp | |
| Quantized GGUF builds ship in this repository alongside the safetensors weights. The `granite` | |
| architecture is supported by llama.cpp, and the chat template is embedded in the GGUF metadata, so | |
| chat mode needs no extra flags. | |
| | file | quant | size | bits/weight | notes | | |
| | ---------------------- | -------- | -------- | ----------- | ------------------------------------------------- | | |
| | TwIL-LM2-Q4\_K\_M.gguf | Q4\_K\_M | 1.44 GiB | 4.88 | recommended default; runs on CPU | | |
| | TwIL-LM2-Q5\_K\_M.gguf | Q5\_K\_M | 1.68 GiB | 5.70 | a little more headroom than Q4\_K\_M | | |
| | TwIL-LM2-Q6\_K.gguf | Q6\_K | 1.94 GiB | 6.57 | close to Q8\_0 quality at about three quarters of the size | | |
| | TwIL-LM2-Q8\_0.gguf | Q8\_0 | 2.51 GiB | 8.51 | near-lossless, for quality-sensitive use | | |
| | TwIL-LM2-F16.gguf | F16 | 4.72 GiB | 16.01 | unquantized, for requantization or reference runs | | |
| ```bash | |
| llama-cli -m TwIL-LM2-Q4_K_M.gguf -cnv --temp 0 -n 2048 | |
| ``` | |
| Pass `--temp 0`, because the evaluation is greedy, and leave the generation budget at 2048 tokens | |
| or more. | |
| The published Track A and Track B numbers were measured on the **bf16** weights through vLLM, not | |
| on any of these GGUF builds, so expect small deviations — most likely at Q4\_K\_M — that have not | |
| been quantified here. Note also that F16 is not bit-identical to the bf16 weights: the two formats | |
| carry the same 16 bits but trade exponent range against mantissa precision. | |
| ## How it was built | |
| Three stages on top of the base model: | |
| 1. **LoRA supervised fine-tuning** on a synthetic formal-logic corpus covering the Track A | |
| objectives (first-order-logic translation, entailment labelling, semantic parsing, Lean | |
| formalisation and critique, procedural reasoning, rule induction), using the project's v5 | |
| SFT recipe. | |
| 2. **WiSE-FT interpolation** toward the pretrained base and checkpoint fusion, | |
| `W = (1 − λ)·W_base + λ·W_finetuned`. λ was chosen to keep as much held-out capability as possible while still gaining | |
| in-domain. | |
| 3. **MGPO** — entropy-weighted GRPO reinforcement learning against a programmatic verifier, with | |
| partial credit for loose matches and token-F1 so that all-fail prompt groups still produce | |
| gradient. Published checkpoint is **step 1400**, chosen by probe Pass@1. | |
| Unlike TwIL-LM3, there is no checkpoint-fusion stage between SFT and WiSE-FT in this model. | |
| ## Limitations and caveats | |
| **Strict output form.** Strict MCQ accuracy is 0.0000, `procedural` accuracy 0.0100, | |
| `fol_translation` primary score 0.0000 and strict-7 0.1214 (eleventh of twelve). The model reasons | |
| near the required form without reliably emitting it. If you need exactly-formatted formal objects, | |
| the SmolLM2-1.7B-based TwIL-LM2 (strict-7 0.2386, semantic parsing 0.8428) is the stronger option | |
| in this comparison. | |
| **Lean critique.** `lean_critic` accuracy is 0.3100, the lowest of the thirteen models compared and | |
| well below TwIL-LM3 (0.6600). | |
| **Perplexity across tokenizers.** `lm_corpus` and `math_corpus` perplexity is a per-token quantity, | |
| and the columns use different tokenizers (Granite and SmolLM2 have about 49K entries each but | |
| distinct vocabularies; SmolLM3 has 128K). The perplexity rows are informative within a family and | |
| only indicative across families. | |
| **Result trees.** For Track B, TwIL-LM3 and the SmolLM2-1.7B-based model are scored from the | |
| results tree that matches their published cards (rope-fixed), and this model from the default tree, | |
| with vLLM 0.19.1; the two TwIL macros reproduce their published values exactly (0.7339 and 0.4333). | |
| Track A figures for this model come from GATE 2 reports at n = 200 per lane and a 2048-token cap. | |
| **Scope.** Tuned for formal logic. The Track B suite reported here does not cover code generation | |
| or tool use, and no claim is made about either. Granite's base tool-calling and document-grounded | |
| chat-template features are inherited but were not evaluated. | |
| **Not a chat model.** It was optimised against automatic verifiers on logic tasks. It has had no | |
| safety tuning beyond whatever the base model carries, and no instruction-following alignment work. | |
| **GGUF builds.** Scores were measured on the bf16 weights only; the quantized builds have not been | |
| evaluated. | |
| ## Evaluation protocol | |
| * Track A: `n = 200` per objective, greedy (`temperature = 0`), `max_new_tokens = 2048`, one retry | |
| at 4096 for truncated rows. | |
| * Track B: 300 examples per task, greedy, 4096 generation tokens, chat template applied, vLLM 0.19.1 | |
| backend. `musr` is the mean of the murder, object and team splits. | |
| * Controlled decode bench: one idle H100 per model, 128 shared prompts, `ignore_eos`, hard 512-token | |
| cap, engine initialisation excluded from the rate. | |
| * All models are scored on the same sampled rows within each track. The TwIL and external-model | |
| figures are those published on the TwIL-LM3 and earlier TwIL-LM2 cards. | |
| Track B is sampled at 300 examples per dataset for compute reasons. Absolute scores can shift on | |
| the full sets, but the comparative ordering across models is expected to be stable. | |
| ## Relationship to TwIL-LM | |
| **TwIL-LM3** ([`webAI-Official/TwIL-LM3`](https://huggingface.co/webAI-Official/TwIL-LM3)) is the | |
| SmolLM3-3B-based member of the family. It is stronger on held-out benchmarks, on strict-7 and on | |
| `lean_critic`; this model is 0.004 behind it on the gate at a smaller size and decodes 1.57x faster. | |
| The name **TwIL-LM2** was previously used for a SmolLM2-1.7B-Instruct-based model. It is the | |
| semantic-parsing specialist in the tables above, and the weakest of the TwIL models on held-out | |
| benchmarks. This repository is a different model — Granite-based, 2.5B — that carries the TwIL-LM2 | |
| name; the SmolLM2-1.7B-based figures above are that model's published numbers, included as a | |
| reference point and not as this model's results. | |
| ## License and attribution | |
| Released under the **webAI Non-Commercial License ver. 1.0** — see `LICENSE.md` in this repository. | |
| The base model, | |
| [`ibm-granite/granite-3.3-2b-instruct`](https://huggingface.co/ibm-granite/granite-3.3-2b-instruct), | |
| is Apache 2.0; its licence text is retained as `apache-2.0-LICENSE.txt` and all credit for the base | |
| model goes to the IBM Granite team. Apache 2.0 permits distributing derivative works under | |
| different terms provided attribution is preserved, which is what the pair of licence files in this | |
| repository does. | |