π± POCKET β a 35-billion-parameter model that runs on your iPhone, and on your PC with no GPU
We're releasing POCKET, VIDRAFT's flagship Darwin-36B-Opus compressed for on-device use. No fork, no CUDA, no cloud β it runs on stock llama.cpp. It's a sparse Mixture-of-Experts model (256 experts, only 8 active per token), so the file can be large while the work per token stays small. That's what lets a 35B model run on a phone, and generate fast on a CPU with no graphics card.
Measured (POCKET-35B IQ1_M vs Bonsai-27B Q1_0): β’ CPU generate (Xeon, 16 threads): 27.0 vs 10.1 tok/s β 2.69Γ faster β’ GPU generate (H100): 197 vs 89 tok/s β 2.22Γ faster β’ GPU prompt processing (H100): 753 vs 1816 β 0.41Γ (Bonsai wins this one β MoE prefill wakes every expert, so sparsity stops helping there. We say so.) β’ Quality (HellaSwag, 400 q): 61.0% vs 60.0% β a tie (confidence intervals overlap)
On a real consumer laptop β MacBook M3 Pro (18 GB) β POCKET wins every axis, prompt processing included: β’ Metal generate: 25.4 vs 12.8 β 1.99Γ β’ CPU generate: 13.8 vs 4.4 β 3.13Γ β’ Metal prompt: 240.7 vs 73.4 β 3.28Γ
One more quiet fact: the same-size, quality-oriented rival Ternary-Bonsai-27B (7.2 GB) fails to load in upstream llama.cpp at all β it needs the PrismML fork. POCKET runs on the tools you already have: LM Studio, Ollama, PocketPal, MLX.
We are thrilled to announce the release of **NRS_QWEN_MYTHOS_1M**, a high-performance reasoning model built on the powerful **Qwen 3.5 9B** base. At **SKT AI LABS**, weβve applied our proprietary **Neural Reasoning System (NRS)** to push the boundaries of what a 9B model can do.
π₯ **Why this model is a Game-Changer:**
β **100x High Reasoning Capacity:** Deep logical thinking and complex problem-solving via NRS Boosting. β **1 Million Token Context:** Handle massive codebases, long documents, and multi-turn agentic tasks with ease (YaRN Scaling). β **Advanced Thinking Mode:** Native <think> tags for step-by-step Chain-of-Thought reasoning. β **Tool-Use Ready:** Optimized for Python execution and Web Search with self-correction. β **Blazing Fast:** Efficient 9B architecture that runs smoothly on consumer hardware (RTX 3090/4090).
π οΈ **Technical Highlights:** * **Base:** Qwen 3.5 9B * **Tuning:** NRS Specific Tuning high-quality samples. * **License:** NRS DOCS Whether you are a developer building coding agents, a researcher dealing with long-context data, or just someone who loves deep reasoning, this model is built for you.
We are excited to share that SKT-NRS is now live on Hugging Face. Weβve developed a Neural Reasoning System (NRS) designed to enhance the capabilities of foundation models β giving them stronger reasoning, improved performance, and more reliable outputs across a wide range of tasks.
Our goal is to bring meaningful quality improvements to both new and existing models. Youβll start seeing boosted versions of various models released here soon, each refined with our NRS approach.
Regular releases of Neural Reasoning-enhanced models Clear focus on better reasoning and overall model quality Ongoing improvements based on community feedback
If youβd like to stay updated, feel free to follow this space β weβll be posting the first boosted models very soon.
**Community Requests**
Have a specific model youβd like us to work on? Looking for improvements on an existing model, or have any other requests? Weβre happy to hear from you. Please share your suggestions here: