📱 POCKET — a 35-billion-parameter model that runs on your iPhone, and on your PC with no GPU
We're releasing POCKET, VIDRAFT's flagship Darwin-36B-Opus compressed for on-device use. No fork, no CUDA, no cloud — it runs on stock llama.cpp. It's a sparse Mixture-of-Experts model (256 experts, only 8 active per token), so the file can be large while the work per token stays small. That's what lets a 35B model run on a phone, and generate fast on a CPU with no graphics card.
Measured (POCKET-35B IQ1_M vs Bonsai-27B Q1_0): • CPU generate (Xeon, 16 threads): 27.0 vs 10.1 tok/s → 2.69× faster • GPU generate (H100): 197 vs 89 tok/s → 2.22× faster • GPU prompt processing (H100): 753 vs 1816 → 0.41× (Bonsai wins this one — MoE prefill wakes every expert, so sparsity stops helping there. We say so.) • Quality (HellaSwag, 400 q): 61.0% vs 60.0% → a tie (confidence intervals overlap)
On a real consumer laptop — MacBook M3 Pro (18 GB) — POCKET wins every axis, prompt processing included: • Metal generate: 25.4 vs 12.8 → 1.99× • CPU generate: 13.8 vs 4.4 → 3.13× • Metal prompt: 240.7 vs 73.4 → 3.28×
One more quiet fact: the same-size, quality-oriented rival Ternary-Bonsai-27B (7.2 GB) fails to load in upstream llama.cpp at all — it needs the PrismML fork. POCKET runs on the tools you already have: LM Studio, Ollama, PocketPal, MLX.
I’m doing a PhD in AI, which sounds impressive until you realize it mostly means I spend three years trying to make a computer say something slightly less stupid than it said yesterday.
People hear "AI researcher" and they think I’m building the future. No. I’m in a basement at 2 a.m. Googling, "CUDA error what the f**k does this mean."
And the worst part about AI research now is compute. You don’t even ask, "Is this idea good?" anymore. You ask, "Can I afford for this idea to be wrong?"
My advisor comes to me one day and says, "I think we should fine-tune our own language model."
I said, "Professor, with what money? I’m a PhD student. I have two bank accounts: checking and emotionally checking."
He goes, "Don’t worry. We have compute."
Now, in academia, "don’t worry" is never the beginning of a good sentence.
I said, "What do you mean we have compute?"
He said, "My friend knows the cluster admin. He can get us on the GPUs."
I said, "Okay… what do we have to do?"
He goes, "Nothing crazy. Just be very grateful in the acknowledgements."
I said, "How grateful?"
He said, "Maybe put him as co-author."
I said, "Co-author? Are we using the cluster, or is the cluster using us?"
Because at that point, that’s not a favor. That’s academic child support.
So I go to the server room, and the cluster admin walks up to me and goes, "So you’re the NLP student."
And in my head I’m like, "No, tonight you’re the principal investigator. You’re the provider. I’m just a little token waiting to be attended to."
Because whoever controls the GPUs controls the relationship. That’s lab romance.
He starts setting things up, and I’m trying to act casual, but I don’t understand any of the numbers he’s saying.
He’s like, "Yeah, I can probably give you four H100s for the weekend."
I’m nodding like, "Mmm. Four. Weekend. H. One hundred. Absolutely."
Inside I’m like, "Is that good? Is that prison time? Why did he say it like he was offering me organs?"