|
Download README.md from logic65/about: direct link, hf CLI and curl.
- Browser
- Download file 2.05 kB
-
https://huggingface.co/logic65/about/resolve/main/README.md
- Command line
-
hf download hf://logic65/about/README.md
-
curl -L -o README.md https://huggingface.co/logic65/about/resolve/main/README.md
2.05 kB
| tags: | |
| - about | |
| # David Aylward | |
| > ### ☕ Support this work | |
| > Whittle is built by one person on a grocery budget and rented GPU hours. If this research is useful to you, or you want to see it finished: | |
| > **[ko-fi.com/davida81328](https://ko-fi.com/davida81328)**. Every hour of GPU time goes straight into the next checkpoint, and every checkpoint, table and log lands in these repos. | |
| AI tinkerer in South Africa. I take language models apart on two 8GB gaming GPUs to | |
| see what's inside, then put smaller ones back together. Everything I break and | |
| everything I learn ships public: weights, measurements, mistakes and all. | |
| ## Current | |
| | | | | |
| |---|---| | |
| | **[Whittle-Qwen-3.8-35B-A3B](https://huggingface.co/logic65/Whittle-Qwen-3.8-35B-A3B)** | 35B-total, ~3B-active Qwen3.8-Flash-Next-format MoE with a 10B n-gram memory, distilled from Qwen3.8-27B; root = Phase-2 step 32010, published 30 Sep 2026. GGUFs on [Whittle-Qwen-3.8-35B-A3B-GGUF](https://huggingface.co/logic65/Whittle-Qwen-3.8-35B-A3B-GGUF). Parent: Whittle-Next-27B-A3B; base lineage: Qwen3.6-35B-A3B pruned 256→180 experts. | | |
| ## Earlier work (Aug 2026) | |
| | | | | |
| |---|---| | |
| | **[Qwen3.8-Whittle-16B](https://huggingface.co/logic65/Qwen3.8-Whittle-16B)** | A 27B whittled to 16.8B with a logit lens and a pricing table, healed with one A100 evening. 36/39 on the field battery at 20 tok/s on consumer GPUs. Research preview, v2. | | |
| | **[The un-repaired cut](https://huggingface.co/logic65/Qwen3.8-p44w75-16.8B-unrepaired)** | The raw surgery artifact: full measurement history, every pricing run, every script. The damage profile is the science. | | |
| ## How this works | |
| No lab, no cluster. Measurements run on a Ryzen 2700X with an RTX 4060 and an | |
| RTX 3050. Training runs on rented A100 hours. Every experiment publishes its wins, | |
| its bugs, and its dead ends on the model cards themselves. | |
| ## Support the tinkering | |
| **[☕ ko-fi.com/davida81328](https://ko-fi.com/davida81328)** | |
| Every donation becomes A100 hours, and every A100 hour ends up as a public model | |
| or a public measurement. | |