How to use from
Docker Model Runner
docker model run hf.co/logic65/whittle-dev:Q8_0
Quick Links

whittle-dev

β˜• Support this work

Whittle is built by one person on a grocery budget and rented GPU hours. If this research is useful to you, or you want to see it finished: ko-fi.com/davida81328. Every hour of GPU time goes straight into the next checkpoint, and every checkpoint, table and log lands in these repos.

Nightly training checkpoints (trainer state: HC/PLE/LoRA/gates) for Whittle-Next runs on Colab. colab1/latest.pt is overwritten as the run progresses; colab1/step<N>.pt are milestones (the earlier note on this card said every 1000 steps; the tree holds step200.pt to step4400.pt, every 200 steps, as the Whittle-Next-27B-A3B card states). Load with the trainer (RESUME_CKPT=), not as a standalone model.

Status

Development artefacts, not a release. These are the run directories referenced by the Whittle-Next-27B-A3B and Whittle-Qwen-3.8-35B-A3B cards: colab1/ is the v4.4 run of Whittle-Next-27B-A3B (its checkpoints and run-end table), and tbl1/, lw2/, lw5/ and agentfix2/ are the checkpoints of the same names described on the Whittle-Qwen-3.8-35B-A3B card. The other directories (colab2/, lw1/, lw3/, lw4/, agentfix/, ai2-archive/) are runs of the same line that no card describes.

What is here

  • colab1/ β€” step<N>.pt (22 milestones), latest.pt, ngram_table_final.npy.
  • colab2/ β€” step<N>.pt, latest.pt, ngram_table_final.npy, eval/ (CE curve, layer-wise eval logs and JSON, export log, TRAIN.log).
  • lw1/ β€” step<N>.pt, latest.pt, ngram_table_final.npy, eval/ (CE curve, layer-wise eval logs and JSON, export log, TRAIN.log).
  • lw2/, lw3/, lw4/, lw5/ β€” step<N>.pt, latest.pt, ngram_table.npy, ple_hash.json, TRAIN.log, eval/ (CE curve, layer-wise eval logs and JSON, math60_replies.jsonl + math60_score.txt, export/convert logs) and one Q8_0 GGUF per run: Whittle-Qwen-3.8-35B-A3B-lw2-Q8_0.gguf, …-lw3-…, …-lw4-…, …-lw5-….
  • tbl1/ β€” step<N>.pt, latest.pt, ngram_table.npy, ple_hash.json, eval/ (GGUF build and quantisation logs, gguf_sizes.json).
  • agentfix/ β€” latest.pt, step200.pt.
  • agentfix2/ β€” latest.pt, step300.pt, bf16/ (full weights, 14 safetensors shards + config, tokenizer and chat template) and Whittle-Qwen-3.8-35B-A3B-agentfix2-Q8_0.gguf.
  • ai2-archive/ β€” archived trainer states: reader7/, reader9/, reader11/, reader12/ (TRAIN.log, ngram_table.npy, ple_hash.json, and keep_final.pt for all but reader9), v3ref/ (resume_from.pt, ple_hash.json), v4/ (colab_onpolicy_next_step42.pt).

The checkpoints are trainer state, not standalone models: the trainer, exporter and the frozen body they apply to are described on the two model cards above.

Downloads last month
193
GGUF
Model size
36B params
Architecture
qwen4exp
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support