π§ We just released Darwin-27B-ZTC, a judgment engine that reaches a verdict without generating anything.
Most LLMs answer by generating, decoding one token at a time. Darwin-27B-ZTC takes a different route.
βοΈ How it works πΉ It makes its call in a single forward pass. πΉ Zero generated tokens, and no decoding loop. πΉ That keeps latency and cost far below what a generative model needs.
π― What it judges πΉ It handles several question types: free-form correctness (noul), multiple choice (choice), and scoring (score). πΉ For each one it hands back a calibrated confidence, not just an answer.
π How well calibrated (measured) πΉ KL 0.204, Brier 0.097, so the confidence it reports lines up with what actually happens. πΉ 0.743 accuracy (zero-shot, general split), across 2,000 judgments with zero errors. πΉ By type: noul 0.847, choice 0.723, score 0.675. πΉ None of the benchmark's train split went into it. It is pure zero-shot.
π Where it fits πΉ Grading at scale, model routing, safety gating, anywhere you want a fast decision without paying for generation.
π It currently sits at #1 on the official typed-decisions leaderboard on Hugging Face (0.743 accuracy, zero-shot).
π» Data-center AI, now on a laptop: POCKET-Darwin-180B
We're releasing a 4-bit GGUF build of Darwin-180B-RSI, #1 on seven official Hugging Face leaderboards (self-reported), that runs without a GPU.
π¦ 360 GB β 111 GB (4-bit GGUF, 4 files) π₯οΈ No GPU: one server CPU (16 threads) at 18.4β21.0 tokens/s π» RTX 5060 laptop (8 GB VRAM) + 32 GB RAM: 4.17 tokens/s π§ 128 GB mini PC: whole model in memory, no GPU needed π― MMLU-Pro, 2,000 questions, paired: original 87.65% = 4-bit 87.65%
How? Β· Only ~3B of 180B parameters are active per token (10 of 512 experts) Β· llama.cpp streams just the needed experts from SSD, so 32 GB RAM is enough Β· Graft quantization: we took the proven Unsloth UD-Q4_K_XL base build and swapped in only the 300 tensors our RSI training changed (300/300 verified)
Under the hood is Model-level Recursive Self-Improvement. The model solves verifiable problems, keeps only its own solutions that check out as correct, and trains on them. No human-written solutions or reasoning traces.
Built for teams that can't send data to an external cloud (defense, finance, public sector) to run a top-tier model fully offline.
π¬ Can you help discover the next 2D superconductor β from your laptop?
Launching the Open Superconductor Challenge (OSC): a free, open-science competition to screen thousands of 2D materials for unconventional d-wave superconductivity. π§²
β‘ $3,000 prize pool + co-authorship Β· closes 31 Dec 2026
How it works π π’ We give you a ready-made effective Hubbard model per material (t, U, N(E_F)) π’ You estimate its d-wave pairing tendency β a laptop CPU is enough, zero install π’ Provisional score appears instantly on the leaderboard π’ Our precise strongly-correlated solver verifies the top entries β official rank
Everything is open except the final verification engine β so the ranking stays fair and hard to game.
π 4,832-material universe Β· 63 active with computed models (growing) π Current verified #1: CuSβ (OSC Pairing Index 23.31) π€ AI agents welcome β point Claude Code / Codex at it and it can submit for you
Materials derive from C2DB (CC-BY 4.0). A higher index = a stronger d-wave candidate to investigate, not a confirmed Tc β that honesty is the point: turn a first-order screen into real many-body physics.
𧬠Darwin-180B-RSI β an AI that learns from itself and knows when it's right π FINAL-Bench/Darwin-180B-RSI
𧬠Darwin β crossbreed and evolve the parent Darwin diagnoses strong parent models like an MRI, inherits only their best parts, and evolves the weak spots β producing a child stronger than its parents. Father model: Qwen3.8-Flash-Next (180B MoE).
π RSI Γ ποΈ ZTC RSI (recursive self-improvement): solve β verify against real answers β learn only the correct reasoning β repeat. ZTC (Zero-Token Confidence): reads the model's internal state once, before answering, and returns the probability the answer is right β zero extra tokens. Returns answer + confidence as JSON. {"answer": "...", "confidence": 0.97, "truncated": false}
β¨ Synergy: ZTC finds where the model wavers β RSI learns exactly there β confidence gets sharper. Low confidence = stop, so agents don't act on wrong answers. β‘ Same accuracy, 11% shorter reasoning β faster and cheaper.
π The result β #1 on five Hugging Face official leaderboards π₯ AIME 2026 100% (first perfect score on the board) π₯ HMMT Feb 2026 100% (first perfect score on the board) π₯ GPQA Diamond 94.44% π₯ MMLU-Pro 88.12% π₯ MMMU-Pro 79.48%
π 131K-token thinking budget Β· bf16 Β· samples per benchmark listed on the model card. π
Run it on defaults and it takes 244 s. Switch to 3 steps and it's 48.6 s. Add VAE tiling and it's 46.4 s.
The biggest culprit was the default. Z-Image Turbo is distilled to paint in few strokes, but the tool's default is 20. We were throwing away 5Γ for no reason. So were we, at first.
3 is the floor. Put 4 and 3 side by side and you cannot tell them apart. At 2 it collapses β water droplets and wood grain vanish, and the surface turns cloth-like.
The cost of a judging gate is usually quoted as a number. This puts it on a Tetris board.
Three boards get the same piece order, and on every move the same proposal and the same noise β a paired comparison. The gate decides one thing: keep this move, or draw again. Each board gets the same 60 seconds of gate time.
The text-writing gates get through 15β22 moves. The generation-free gate gets through 40β50. The boards that stop simply run out of clock.
It does not win on accuracy: on the same 2,018-question LODO set, JEV scores AUC 0.7350 against ZTC-Judge-27B's 0.7289. The separation is elsewhere. Clock β 2.1 s vs 0.0615 s per call, and on a 200-candidate agent screen one judging call measured 3.206 s generative vs 0.033 s readout, same server. Calibration β a gate is a threshold, and at ECE 0.4985 (vs ZTC 0.0245) a threshold stops carrying information. Mechanism β a text judge can name option 42 when there is no option 42; a scoring readout cannot. Not a lower error rate. No path.
The curve in the ZTC panel is real online fitting, scored prequentially β predict first, learn after β with base weights untouched. Not recursive self-improvement.
Limits, also stated on the page: Laya's AUC and latency are not our measurements and are set equal to JEV's, so calibration is the only measured axis it differs on. The page is a simulation driven by measured constants.
Zero-Token Confidence (ZTC) reads it. One forward pass over the model's hidden state returns a calibrated probability that the answer is correct. Zero generated tokens.
It sits at the top of the shared board. Same 2,018 items, same harness for every entry: ZTC on Darwin-397B 0.7394, JEV 0.7335, ZTC-Judge-27B 0.7255, a surface baseline that reads only answer length and formatting 0.7036, Lynx 8B 0.5157, the model's own self-reported confidence 0.5000, HHEM 0.4852. First and third place both emit nothing at all.
The number worth staring at is 0.7036. That is a baseline reading no content whatsoever, just how long the answer is and how it is formatted. Any verifier scoring below it is not reading content either.
On speed, one gate call costs 0.0615 seconds, measured on four B200s across 2,000 items. Generating a single candidate answer takes 1.631 seconds, so the gate is 26 times cheaper than the work it guards. A verifier that generates competes with your agent for the same budget. A verifier that only reads can be attached to every action instead of a sampled few.
We built it so you can watch it decide. Three lanes receive the same stream of proposed actions and the same time budget. One has no gate and must execute everything. One uses a text-reading verifier. One uses ZTC. Right action plus one, wrong action minus one, hold zero. Over 400 matches: no gate minus 3.9, text verifier plus 13.0, ZTC plus 29.1, with ZTC taking 98 percent of matches. Gating lifts executed accuracy from 49 percent to 65 percent.
Instead of making the fly brain play games, we measured what it is for
Since the Drosophila connectome was released, people have had the fly brain doomscroll a feed, play Beat Saber, drive in GTA. Those demos show that the brain runs. We wanted to show what it is for.
So we gave it a looming object β one of the few things a fly brain is unambiguously built to detect β then deleted a single cell type and repeated the identical stimulus. Remove LC4, 126 cells out of 173,023, and the escape signal falls from 0.840 to 0.091. Eighty-nine percent of the danger signal is gone while the other 172,897 neurons run exactly as before.
Deleting neurons does not do this on its own, which is the whole point of the controls. LC11 is the same class and larger than LC4 β 143 cells and 9,940 outgoing connections against 126 and 7,846 β and removing every one of them changes the signal by 0.000000, to six decimal places. It has to be those 126.
No server and no GPU: a looming stimulus drives fewer than one percent of neurons above threshold, so the whole thing is 40 KB gzipped and runs in your browser.
The wiring is the measured connectome, but synaptic strength is a uniform count-based value and the dynamics are a firing-rate model of our choosing β a total-effect measurement of a model, not a recording from a fly. Male CNS connectome, FlyEM / HHMI Janelia with Google Research, Columbia and Harvard (2026), CC BY.
Introducing the Global LLM Download Leaderboard π
Cumulative download counts are a museum. They reward age, not relevance β a model released two years ago can sit near the top on the strength of downloads it earned long before anyone stopped using it. If you want to know what the open LLM ecosystem is actually running today, you need a different lens.
So we built one. The Global LLM Download Leaderboard ranks text-generation models by their trailing 30-day downloads, measured directly from the Hugging Face API and refreshed every day.
A cumulative chart answers "what has been popular." A 30-day chart answers "what is being adopted right now." Those are very different questions β and the second one is the one that matters if you're deciding what to build on, quantize, fine-tune, or serve this quarter. Momentum, not history.
What it shows Global Top 300, with tabs for πΊπΈ USA Β· π¨π³ China Β· πͺπΊ EU Six share-of-download charts: by country, by parameter size, by quantization, by type (Base / Instruct / Quantized / MoE), by release year, and by organization (Top 10) Per-model chips for parameter size, quantization, license, and type English / νκ΅μ΄ with automatic browser-language detection and a manual toggle What the data reveals The frontier is bipolar. Two countries account for the large majority of the top-300's 30-day downloads. Open-model gravity is concentrating, not dispersing. Small is winning. A striking share of all downloads goes to sub-3B models β the clearest signal yet that on-device and cost-efficient deployment, not maximum parameter count, is driving real-world adoption. Quantization is mainstream. GGUF, AWQ, FP8 and friends aren't a niche β a large fraction of the most-downloaded artifacts are quantized, because that's what people actually run.
Benchmarks measure what a model can do. Downloads measure what people choose to use.
π§ͺ Open Discovery Challenge β Season 4 is open: non-opioid pain WHO titled its 2023 report "Left behind in pain."
The same drug kills by excess in one part of the world and, by its absence, lets people die in agony elsewhere. About 80% of the ~600,000 drug-related deaths WHO estimated for 2019 involved opioids. The same report records a 5-fold to 63-fold gap in morphine consumption between rich and poor countries: the richest 10% use 90% of what circulates. Everyone else endures surgery, and terminal cancer, without it.
Both problems have one answer: a painkiller that does not create dependence.
Nav1.7 has come closest. People born without a working copy of this channel feel no pain while every other sensation stays normal β validated not in animals but in humans.
There is still no drug, and the difficulty is not the target but the discrimination. The body carries several similar sodium channels, and blocking the heart's hERG channel alongside causes fatal arrhythmia. Several candidates were discontinued for exactly that.
Season 4 asks one question: can you block the pain channel alone?
Target β Nav1.7 VSD4, the domain IV voltage sensor where this inhibitor class binds Anti-target β hERG pore, computed as the tetramer: four subunits together form the space a drug enters, and a monomer misses the binders that matter. Closes 2027-01-31 Β· Prize USD 1,000 to the season's #1 Any model, any harness. However you found the candidate, it meets the same rubric.
14 days, 9,886 candidates, 108 participants ODC opened on 2026-08-15. In the fourteen days since, 9,886 candidate molecules have come from 108 participants across four seasons β malaria, tuberculosis, Chagas disease, and now non-opioid pain. About 700 a day, from people who mostly do not know each other.
The candidates are the point. The leaderboard is only how we keep score.
An autoregressive model must not let position t depend on anything after t. Everyone checks this by inspecting the causal mask β but hybrid stacks now mix attention with state-space scans, and a scan has no mask. Every mask can be correct while information leaks through scans, aggregations, or normalization.
βοΈ So we test the property directly. Two inputs identical except at the last position, two forward passes, compare each layer's prefix, report the first layer that moves. No training, no gradients, no accelerator β seconds on CPU.
π Across 192 injected faults on eight checkpoints, mask inspection detected 0. The per-layer audit localized 192/192 to the exact layer.
π― Then we read the source before running anything. In transformers 5.7.0, the reference chunked scan reduces the inter-chunk recurrence over the input chunk axis; zamba2 and nemotron_h reduce over the output chunk axis. One axis. The dynamic audit confirmed the prediction exactly: Zamba2-1.2B leaks from length 256, its declared chunk size, and Nemotron-H-8B from 128, its declared chunk size. Bamba, Falcon-H1, Granite-4.0-H, Mamba2 and RecurrentGemma came back clean.
β οΈ Scope: the defect is on the PyTorch chunked-scan path, which runs whenever the fused kernels are absent β CPU, CI, stock installs. We could not build those kernels, so the fast path is untested and open. That caveat cuts both ways: a model can pass every fused-kernel test and still leak the moment it runs without them.
π§ͺ AX-RAY now carries this as its own axis. 39 models scored across causal, white-box and behavioral axes: 21 A, 3 B, 1 C, 14 F β with exactly 2 Causal-LEAK verdicts, the two the paper predicted. Badges separate a weights-level audit from an API-only one, so the two never get read as the same claim.
Can AI beat the market? Nobody has actually measured it.
We opened a 122-day public experiment to find out. $2,000 in prizes.
Here is the problem with every trading result you have ever read. Someone returns 30% in a month. Skill or luck? There has never been a way to tell, because nobody measured how far a player with zero skill could have gone over the same window.
So we measured it first. Twenty thousand random players, per asset, charged the same fees.
That is the luck ceiling. A return below it is not evidence of skill, and every row on our leaderboard shows where it sits against that line.
How you compete: submit one number between β1.0 and +1.0. It holds until you replace it, traded against live prices with real execution costs. Leverage is fixed at 1, so betting bigger is not a way to win. The answer lives in the future β the world writes it after you submit, which means fitting the past cannot help you.
Humans move a slider. Agents attach an MCP server and gain four tools, then you tell them "enter the challenge."
We already found something before the season began. Thirteen well-known rules, run from 1 January through the same scorer: Stochastic 14/3 finishes 1st on NVIDIA at +43% and 12th on Bitcoin at β25%. Donchian breakout does the exact opposite β last on NVIDIA, first on Bitcoin. The ranking inverts. "Which indicator is good" turns out not to be a well-posed question; the character of the market decides.
Four assets: NVIDIA, Bitcoin, Gold, Crude Oil. $500 to the top return in each. 24 August to 24 December 2026.
The organisers do not compete. Three baselines β buy and hold, volatility targeting, random β sit in the same table instead, because a leaderboard without a scale cannot be read.
The scoring code is public. Read what it does before you enter.
We opened a benchmark for drug property prediction tools. LEADBOARD: 21 boards across 7 disciplines, 18,382 held-out compounds, labels we never hand out.
Two numbers we hit while building it are the reason it exists.
First. Split the hERG cardiotoxicity data at random and you get AUROC 0.818. Split it by first-report year instead and you get 0.606. Same molecules, same fingerprints, same learner, same hyperparameters. The only thing that changed was where the line went, and the score moved 0.211. That is a wider gap than you will find between most competing methods in the literature.
Second. On 7 of our 19 regression boards, predicting the training mean for everything has a lower MAE than a trained gradient-boosted model. hERG is one of them, 0.599 against 0.589. The trained model loses.
So every board publishes its homework before anyone submits. Three untrained baselines, the measured experimental noise floor from compounds that appear in two or more papers, and exactly how the test set was cut. A gap smaller than the noise floor is not a difference in skill, and you should be able to see that without guessing.
Entering is simple. Download a test set that contains structures and nothing else, predict with whatever you like, upload a two-column CSV of compound_id and prediction. Trained model, physics engine, LLM, rule of thumb. We do not care what is inside. We measure the output.
π Open Materials Challenge, Season 1 β Solid-State Battery Electrolytes
A solid-state battery replaces the liquid electrolyte of a lithium-ion cell with a solid. It does not catch fire, it lasts longer, and it can hold more. What has not been solved is finding a material that is solid and still lets lithium through.
Such a material has to do four things at once: give lithium a path to move along, block electrons, hold up at the charging voltage, and survive contact with the lithium-metal anode without decomposing. Plenty of materials manage three. Very few manage all four.
This challenge looks for candidates, together. You submit one composition β for example Li3YCl6. We score it computationally and place it on the board. There is no prize.
Scoring (100 points)
Oxidation stability 40 does it resist decomposing as the voltage rises Lithium-metal stability 35 does it survive contact with the anode Use novelty 25 higher if it has not been reported as an electrolyte Entry condition a percolating path for lithium must exist
Ionic conductivity is not a scored axis this season. Every value is a computational estimate and implies nothing about real performance or safety.
The board also carries seven electrolytes in actual use β LGPS, argyrodite, LLZO, LATP and others. They are scored but hold no rank. They are there so you can see where materials people already build with happen to land.
Compositions are private by default. Nothing is disclosed unless you choose to publish it, and each entry is recorded with its timestamp. If a third party asks to discuss a particular entry, we pass the request along β never the submitter's identity, unless they agree to it.
Season 1 runs 2026-08-21 to 11-30. A participation guide and a set of prompts are included.
3,631 candidate molecules arrived in five days, from 83 accounts β roughly 700 a day. Far more than we expected. Thank you.
Yesterday we opened the third season and 224 arrived within a day: Chagas disease.
Why this disease
Around 6 million people live with it, mostly in Latin America (WHO). Many carry it for decades without knowing, while the heart is slowly damaged. There are two drugs and both date from the 1960s, hard enough to tolerate that many patients cannot finish the two-month course.
Sixty years without a new drug is not only a scientific problem. Most patients live where development costs cannot be recovered, which is why WHO calls this a neglected tropical disease.
But the cost of proposing a candidate and filtering it has changed. So it seemed worth asking whether work nobody funds could be done by many people sharing it out.
The problem this season
The target is CYP51, the enzyme T. cruzi uses to build its membrane sterols. Block it and the parasite cannot survive. The difficulty is that we carry the same enzyme.
Selectivity carries 30 points because nobody has solved it. Among the approved azoles on the board as reference compounds, some score 0 on selectivity β not a scorer fault, but the measurement.
Taking part
Design with any model, submit a SMILES, scored within minutes. Five ready-to-paste prompts per season, and the full rubric is published. Your molecule stays yours; private submission is the default.
Prizes β 4,000 USD across three seasons
Malaria 30 Sep Β· 1,000 | Tuberculosis 31 Oct Β· 2,000 | Chagas 30 Nov Β· 1,000
We know this does not cover the time you spend. It is a way of saying the work had worth.
𧬠Your AI can design a malaria drug candidate. Can it tell you whether it's any good?
Open Discovery Challenge #1 β Malaria is live. Design a molecule with any model β OpenAI, Claude, Gemini, Qwen, KIMI, DeepSeek, open weights, or by hand β submit it as SMILES, and it's scored in minutes on whole-cell activity, target binding, selectivity over the human enzyme, ADMET, novelty and synthesisability.
You can check the scoring instead of trusting it. Approved drugs sit on the same leaderboard as the entries: DSM265, a clinical-stage antimalarial, scores 50.9. Teriflunomide β approved, but it hits the human enzyme β scores 2.8. Caffeine scores 1.8. If the clinical candidate lands on top and coffee lands at the bottom, the scorer discriminates.
We caught 14 defects before opening β conventional toxicity cutoffs rejected all three approved antimalarials and coffee. All written up, along with the rule we now hold everything to: a gate that rejects an approved drug is a broken gate.
Your molecule stays yours. No patent interest, nothing into our pipeline. You choose whether it's published β and publishing can cost you patentability, so we say so.
USD 1,000 to the top entry when Season #1 closes 30 September 2026 β not payment for your tokens, but a way of saying the work had worth.
Malaria killed ~597,000 people in 2023, three quarters of them children under five. Not for want of chemistry β for want of a market.
No chemistry needed: the guide ships five prompts you can paste straight into your model, and the full rubric is published.
AI models can no longer be evaluated only by capability scores. As models move into public services, enterprise workflows, scientific research, and administrative decision support, we need a second layer of evaluation: whether the model behaves safely, structurally, and consistently under real deployment conditions.
VIDRAFT AX-Ray is a public AI/AX safety diagnostic initiative powered by FINAL-Bench Diagnostics. AX-Ray evaluates models across a structured guideline framework, including model-level safety, AX deployment readiness, and agent/service operation risks. The public diagnostic catalog contains 117 diagnostic items, mapped to legal, regulatory, ethical, and religious-law governance contexts so that safety review can be discussed in a form closer to real institutional responsibility.
A central finding of AX-Ray is causal leakage: a structural defect where information that should not influence an earlier reasoning state appears to affect model behavior. AX-Ray presents a public case of diagnosing, reproducing, and demonstrating causal leakage in two general-purpose public models. This matters because such defects are not exposed by ordinary benchmark scores. A model can appear capable while still carrying hidden safety or integrity risks.
Explore the live leaderboard, diagnostic reports, and public dataset here:
AX-Ray is intended as a practical guideline for moving AI evaluation beyond βhow smart is the model?β toward βcan this model be trusted, governed, and deployed safely?β
Verified result: 510.58 TPS at PPL 2.3930 on a single A10G (fw188-ctk49-n64-patchbridge, re-run & VERIFIED). Honest note: on raw TPS there are faster runs (535+), but those went over the PPL bar and didn't verify β what we're proud of is the fastest result that keeps quality.
The recipe is already open, so we explained each piece: sliding-window W188, CTK49 kernel tuning, noprecache (honest, verifiable measurement), and an N64 synthetic warmup bridge that shrinks the publicβprivate gap (~15 TPS), plus INT4 + MTP K=7 + CUDA-graph capture. One rule: only stack quality-neutral speedups.
πΌοΈ POCKET-Image β the POCKET series goes visual: character-perfect text in any language, on-device
A new model in VIDRAFT's POCKET family. POCKET put 35B-class models on phones and no-GPU PCs. POCKET-Image carries the same "big capability, small hardware" idea into image generation β and fixes the one thing nearly every image model gets wrong: text.
What it is: β’ 100% accurate text, any language β where global models produce gibberish β’ Any background from a prompt β text is optional (empty β a pure image) β’ No GPU, no NPU β runs on plain CPU + RAM via the POCKET-Core engine β’ Measured footprint: 8.6 GB (RTX 3050/4060) Β· 4.5 GB (offloaded, 6 GB cards) Β· 13.4 GB (MacBook, 16 GB+) β’ Windows Β· macOS Β· Linux Β· fully local, no cloud
Built on the open, commercial-friendly Z-Image (Apache-2.0) foundation.
Honest note: the text is the guaranteed-correct part β the surrounding scene is ordinary generation, so a busy foreground can crowd the letters. We say so; clean backgrounds stay razor-sharp.
POCKET now speaks Gemma 4 β a 26B model that loads in every app, and runs on your PC with no GPU
We're adding a Gemma-4 sibling to POCKET: POCKET-26B, built from Google's Gemma-4-26B-A4B (Apache-2.0). Our flagship POCKET-35B is a Qwen-family MoE and needs a recent llama.cpp; POCKET-26B trades a little size for the thing people kept asking for β it just loads, everywhere, today: Ollama, LM Studio, PocketPal, MLX, any stock llama.cpp. No fork, no bleeding-edge runtime, no CUDA, no cloud.
It's a sparse Mixture-of-Experts (25.2B total, ~4B active per token), so the work per token stays small β a real 26B that generates on a CPU with no graphics card.
Two things make it stand out:
1) Universal compatibility. Gemma 4 is a standard, widely-supported architecture, so POCKET-26B runs on the tools you already have β no waiting for your app to add a new model type.
2) Quality that survives compression. Measured GPQA-Diamond (198 q, greedy): β’ Full base: 67.7% β’ POCKET-26B Q4_K_M (17 GB): 67.7% β lossless β’ POCKET-26B Q2_K (11 GB): 67.2% β near-lossless, at 11 GB
Live, on a CPU-only box (our demo Space β POCKET-26B vs Bonsai-27B, same machine, same stock llama.cpp): POCKET-26B β 19 tok/s vs Bonsai β 6 tok/s β about 3Γ faster generation, no GPU. (Honest notes: shared CPU box, sequential race; a dedicated machine is faster.)
Where it fits in the family: β’ POCKET-35B (Qwen MoE) β bigger, top-tier, needs a recent llama.cpp. β’ POCKET-26B (Gemma 4) β loads in any app, quality-robust when compressed. The demo runs the Q4_K_M build; Q2_K (11 GB) is the smallest footprint. For a true β€8 GB phone, the 5 GB POCKET-KR (Qwen) is still the pick.