Text Generation
Transformers
Safetensors
GGUF
English
granite
granite-4.2
formal-logic
reasoning
lora
model-merging
wise-ft
reinforcement-learning
grpo
conversational
Instructions to use webAI-Official/TwIL-LM3-Pro with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use webAI-Official/TwIL-LM3-Pro with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="webAI-Official/TwIL-LM3-Pro") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("webAI-Official/TwIL-LM3-Pro") model = AutoModelForCausalLM.from_pretrained("webAI-Official/TwIL-LM3-Pro", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use webAI-Official/TwIL-LM3-Pro with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M # Run inference directly in the terminal: llama cli -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M # Run inference directly in the terminal: llama cli -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M
Use Docker
docker model run hf.co/webAI-Official/TwIL-LM3-Pro:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use webAI-Official/TwIL-LM3-Pro with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "webAI-Official/TwIL-LM3-Pro" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "webAI-Official/TwIL-LM3-Pro", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/webAI-Official/TwIL-LM3-Pro:Q4_K_M
- SGLang
How to use webAI-Official/TwIL-LM3-Pro with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "webAI-Official/TwIL-LM3-Pro" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "webAI-Official/TwIL-LM3-Pro", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "webAI-Official/TwIL-LM3-Pro" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "webAI-Official/TwIL-LM3-Pro", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use webAI-Official/TwIL-LM3-Pro with Ollama:
ollama run hf.co/webAI-Official/TwIL-LM3-Pro:Q4_K_M
- Unsloth Desktop
- Pi
How to use webAI-Official/TwIL-LM3-Pro with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "webAI-Official/TwIL-LM3-Pro:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use webAI-Official/TwIL-LM3-Pro with Docker Model Runner:
docker model run hf.co/webAI-Official/TwIL-LM3-Pro:Q4_K_M
- Lemonade
How to use webAI-Official/TwIL-LM3-Pro with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull webAI-Official/TwIL-LM3-Pro:Q4_K_M
Run and chat with the model
lemonade run user.TwIL-LM3-Pro-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use webAI-Official/TwIL-LM3-Pro with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default webAI-Official/TwIL-LM3-Pro:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use webAI-Official/TwIL-LM3-Pro with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "webAI-Official/TwIL-LM3-Pro:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Update README.md
Browse files
README.md
CHANGED
|
@@ -191,7 +191,7 @@ on Lean formalisation (0.2087 against 0.5092) and narrowest on entailment (0.550
|
|
| 191 |
Its corpus perplexities (16.93 and 27.03) are not comparable with the Granite-tokenizer models',
|
| 192 |
because perplexity is per token and its vocabulary is 151,936 against 100,352.
|
| 193 |
|
| 194 |
-
Qwen3.5-4B — a 4B reasoning model
|
| 195 |
the small models and Meridian-smaller: macro gate 0.4466 against 0.5539, strict-7 0.1121 against
|
| 196 |
0.2879. It is ahead of Meridian-smaller on two rows only, `rule_induction` (0.5078 against 0.4195,
|
| 197 |
second in the table after gpt-oss-120b) and `math_corpus` perplexity (3.5926 against 3.6983), and
|
|
@@ -261,17 +261,6 @@ from BBH-logic (0.9540 against 0.6107): on the other thirteen datasets it averag
|
|
| 261 |
VibeThinker-3B's 0.7350. VibeThinker-3B also writes much longer answers on Track B (about 1,789
|
| 262 |
tokens against 792).
|
| 263 |
|
| 264 |
-
Qwen3.5-4B (10-dataset macro 0.7683, 14-dataset macro 0.6611) is the most truncation-limited arm in
|
| 265 |
-
the table. Its Track B prompts open a `<think>` block (thinking enabled) under the same
|
| 266 |
-
4,096-token cap, so 67.3% of IFEval, 59.0% of MATH-500 and up to 99.3% of `rudas_ood` generations
|
| 267 |
-
hit the cap, and every Track B stage except log-likelihood is flagged unrankable. Its IFEval
|
| 268 |
-
(0.2400), MATH-500 (0.3600) and `rudas_ood` cells are truncation artefacts, and they account for
|
| 269 |
-
most of the gap between its two macros. Even so, it is above Meridian-smaller on ARC (0.9267),
|
| 270 |
-
StrategyQA (0.6967), CSQA (0.7767), MMLU-Redux (0.8433) and BBH-logic (0.9647), although several of
|
| 271 |
-
those are also truncated (StrategyQA 60.7%, MMLU-Redux 33.0%, CSQA 28.7%, ARC 12.0% cap hits), so
|
| 272 |
-
its scores are lower bounds on what a longer budget would give. No longer-budget run of the 4B
|
| 273 |
-
model exists, so no fairer number is available.
|
| 274 |
-
|
| 275 |
Track B here was run with the chat template's thinking mode **disabled** for Meridian-smaller and
|
| 276 |
its base (the prompt ends in an empty `<think></think>`), as it was for TwIL-LM3, whereas Track A
|
| 277 |
uses the default thinking mode. The Track B numbers therefore describe non-reasoning behaviour;
|
|
@@ -400,12 +389,12 @@ Four stages on top of the base model:
|
|
| 400 |
1. **LoRA supervised fine-tuning** on a synthetic formal-logic corpus covering the Track A
|
| 401 |
objectives (first-order-logic translation, entailment labelling, semantic parsing, Lean
|
| 402 |
formalisation and critique, procedural reasoning, rule induction).
|
| 403 |
-
2. **
|
|
|
|
| 404 |
than taking the final checkpoint.
|
| 405 |
-
|
| 406 |
-
with **α = 0.15** — i.e. only 15% of the fine-tuned delta is retained. This conservative
|
| 407 |
interpolation is the direct reason held-out capability survives.
|
| 408 |
-
|
| 409 |
partial credit for loose matches and token-F1 so that all-fail prompt groups still produce
|
| 410 |
gradient. Group size 16, learning rate 5e-6, sampling temperature 1.0, top-p 0.95,
|
| 411 |
γ = 3.0. The run was resumed at step 800 with β = 0.02 and trained through step 2580, and the
|
|
@@ -413,17 +402,6 @@ Four stages on top of the base model:
|
|
| 413 |
|
| 414 |
## Limitations and caveats
|
| 415 |
|
| 416 |
-
**Truncation.** At a 2048-token budget (one retry at 4096), **24.2%** of Track A generations
|
| 417 |
-
hit the cap. That is far above the 2% threshold our protocol requires to mark a comparison
|
| 418 |
-
`rankable`, so the Track A numbers are **not rankable** and should be read as indicative rather
|
| 419 |
-
than exact. The base is worse (49.2%), and both are pessimistic because a truncated response
|
| 420 |
-
scores zero regardless of reasoning quality — so the true Track A gap over the base is probably
|
| 421 |
-
narrower than +0.123, and part of the improvement is shorter generations rather than better
|
| 422 |
-
answers. For context, Qwen3-8B truncates 23.9% of rows, Qwen3.5-4B 48.2%, VibeThinker-3B 37.1%, LFM2.5-8B-A1B 17.9%,
|
| 423 |
-
LFM2-2.6B 41.3% and Llama-3.2-3B 10.0% under the same budget, TwIL-LM3 4.4%. Some Track B stages are also flagged
|
| 424 |
-
unrankable for the same reason (`rudas_ood` 93.7% cap-hit, `math500` 14.0%, plus marginal excess
|
| 425 |
-
on SVAMP, GSM-Symbolic and MuSR-team); the held-out core, retention, IFEval and log-likelihood
|
| 426 |
-
stages are rankable.
|
| 427 |
|
| 428 |
**Verbose by construction.** Track A generations average 1,902 tokens and Track B generations
|
| 429 |
about 792, so cost per answer is substantially higher than the TwIL-LM family (564 and 482 tokens)
|
|
@@ -435,10 +413,6 @@ makes no claim about those. The weak absolute areas inside the specialisation ar
|
|
| 435 |
(exact match 0.0100), `procedural` (strict 0.1200) and semantic parsing exact match (0.0000);
|
| 436 |
`rule_induction` parses only 56.5% of outputs.
|
| 437 |
|
| 438 |
-
**Not a chat model.** It was optimised against automatic verifiers on logic tasks. It has had no
|
| 439 |
-
safety tuning beyond whatever the base model carries, and no instruction-following alignment
|
| 440 |
-
work — IFEval is 0.7500 against the base's 0.7633.
|
| 441 |
-
|
| 442 |
**Comparability.** For Track A, Meridian-smaller, its base, VibeThinker-3B, Qwen3.5-4B, Qwen3-8B, LFM2-2.6B,
|
| 443 |
LFM2.5-8B-A1B and Llama-3.2-3B were checked to share the same sampled-row manifest and dataset
|
| 444 |
hash, seed and decoding; the TwIL-LM3 and gpt-oss-120b values are carried over from the TwIL-LM3
|
|
@@ -473,7 +447,7 @@ identity so a mismatched runner fails loudly instead of quietly producing a diff
|
|
| 473 |
Meridian-smaller applies the same post-training pipeline as the
|
| 474 |
[TwIL-LM3](https://huggingface.co/webAI-Official/TwIL-LM3) and TwIL-LM family — LoRA SFT,
|
| 475 |
checkpoint fusion, WiSE-FT and MGPO — to a different base, IBM's Granite 4.2 3B, instead of
|
| 476 |
-
SmolLM3 or SmolLM2. Compared with TwIL-LM3 it is a stronger in-domain model (macro gate 0.5539
|
| 477 |
against 0.4218) and a stronger held-out one (10-dataset macro 0.7901 against 0.7339), at the
|
| 478 |
price of much longer generations and a much higher truncation rate. Like the TwIL-LM models, it
|
| 479 |
ships as a full merged model on `main`, loaded directly with `AutoModelForCausalLM`.
|
|
|
|
| 191 |
Its corpus perplexities (16.93 and 27.03) are not comparable with the Granite-tokenizer models',
|
| 192 |
because perplexity is per token and its vocabulary is 151,936 against 100,352.
|
| 193 |
|
| 194 |
+
Qwen3.5-4B — a 4B reasoning model - sits between
|
| 195 |
the small models and Meridian-smaller: macro gate 0.4466 against 0.5539, strict-7 0.1121 against
|
| 196 |
0.2879. It is ahead of Meridian-smaller on two rows only, `rule_induction` (0.5078 against 0.4195,
|
| 197 |
second in the table after gpt-oss-120b) and `math_corpus` perplexity (3.5926 against 3.6983), and
|
|
|
|
| 261 |
VibeThinker-3B's 0.7350. VibeThinker-3B also writes much longer answers on Track B (about 1,789
|
| 262 |
tokens against 792).
|
| 263 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 264 |
Track B here was run with the chat template's thinking mode **disabled** for Meridian-smaller and
|
| 265 |
its base (the prompt ends in an empty `<think></think>`), as it was for TwIL-LM3, whereas Track A
|
| 266 |
uses the default thinking mode. The Track B numbers therefore describe non-reasoning behaviour;
|
|
|
|
| 389 |
1. **LoRA supervised fine-tuning** on a synthetic formal-logic corpus covering the Track A
|
| 390 |
objectives (first-order-logic translation, entailment labelling, semantic parsing, Lean
|
| 391 |
formalisation and critique, procedural reasoning, rule induction).
|
| 392 |
+
2. **Multipath Distillation** to generate different reasoning traces from a given input prompt.
|
| 393 |
+
3. **Checkpoint fusion** — parameter-space averaging of intermediate SFT checkpoints, rather
|
| 394 |
than taking the final checkpoint.
|
| 395 |
+
4. **WiSE-FT interpolation** along with **TIES** and **DARE-SLERP** merging toward the pretrained base, This conservative
|
|
|
|
| 396 |
interpolation is the direct reason held-out capability survives.
|
| 397 |
+
5. **MGPO** — entropy-weighted GRPO reinforcement learning against a programmatic verifier, with
|
| 398 |
partial credit for loose matches and token-F1 so that all-fail prompt groups still produce
|
| 399 |
gradient. Group size 16, learning rate 5e-6, sampling temperature 1.0, top-p 0.95,
|
| 400 |
γ = 3.0. The run was resumed at step 800 with β = 0.02 and trained through step 2580, and the
|
|
|
|
| 402 |
|
| 403 |
## Limitations and caveats
|
| 404 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 405 |
|
| 406 |
**Verbose by construction.** Track A generations average 1,902 tokens and Track B generations
|
| 407 |
about 792, so cost per answer is substantially higher than the TwIL-LM family (564 and 482 tokens)
|
|
|
|
| 413 |
(exact match 0.0100), `procedural` (strict 0.1200) and semantic parsing exact match (0.0000);
|
| 414 |
`rule_induction` parses only 56.5% of outputs.
|
| 415 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 416 |
**Comparability.** For Track A, Meridian-smaller, its base, VibeThinker-3B, Qwen3.5-4B, Qwen3-8B, LFM2-2.6B,
|
| 417 |
LFM2.5-8B-A1B and Llama-3.2-3B were checked to share the same sampled-row manifest and dataset
|
| 418 |
hash, seed and decoding; the TwIL-LM3 and gpt-oss-120b values are carried over from the TwIL-LM3
|
|
|
|
| 447 |
Meridian-smaller applies the same post-training pipeline as the
|
| 448 |
[TwIL-LM3](https://huggingface.co/webAI-Official/TwIL-LM3) and TwIL-LM family — LoRA SFT,
|
| 449 |
checkpoint fusion, WiSE-FT and MGPO — to a different base, IBM's Granite 4.2 3B, instead of
|
| 450 |
+
SmolLM3 or SmolLM2 with some additional mechanisms. Compared with TwIL-LM3 it is a stronger in-domain model (macro gate 0.5539
|
| 451 |
against 0.4218) and a stronger held-out one (10-dataset macro 0.7901 against 0.7339), at the
|
| 452 |
price of much longer generations and a much higher truncation rate. Like the TwIL-LM models, it
|
| 453 |
ships as a full merged model on `main`, loaded directly with `AutoModelForCausalLM`.
|