Starting…
Private AI on this device
Models run in this browser. What you type is not sent to any server; web search runs only when you turn it on.
📱 On a phone, big models won’t fit. Point LocalMind at your computer’s Ollama / LM Studio (Settings → Models → Local endpoint) and use any model over your network.
About each model
- Ternary Bonsai 4B (~1.1 GB) — text + agent (tool calling). 1.58-bit ternary weights, Qwen3 backbone, Apache-2.0 — the small ternary tier.
- Ternary Bonsai 2 27B (~5.9 GB) — text + agent. PrismML's Ternary Bonsai 2 27B (Sept 2026; ternary 1.75-bit weights, Qwen3.8-27B backbone, 98% of the full-precision model's benchmark score) on a from-scratch WebGPU engine: hand-written WGSL for ternary dequant + Gated-DeltaNet linear attention, reading the PTQ1_0 GGUF directly. Ported from webml-community/ternary-bonsai-2-webgpu-kernels; Apache-2.0 weights. WebGPU-only — the biggest in-browser model here. A reasoning model: flip Show reasoning in Settings for its chain-of-thought. Text-only here (its optional vision tower is not loaded).
- Underdog Saluki 27B (~7.9 GB) — text + agent. ConwayResearch's Underdog Saluki 27B 1.0 (Oct 2026): Qwen3.8-27B as one 2-bit-class IQ2-mix GGUF, tuned to keep tool calling intact (the vendor reports 88 of 120 on a frozen BFCL v4 tool-calling set, against 84 for the full-size model and 70 for Ternary Bonsai 2). Runs on a from-scratch WebGPU engine that reads the GGUF’s eleven llama.cpp block types directly; greedy output matches llama.cpp on the same file. The first load copies the 7.9 GB file into this browser’s on-device storage once (or pick the file you already have with Load .gguf from disk…); it needs about 8 GB of GPU memory. Apache-2.0 weights. A reasoning model: flip Show reasoning in Settings for its chain-of-thought. Text-only here (its optional vision add-on is not loaded).
- Qwen3.5 4B (~3 GB) — text + image + agent (tool calling). Qwen3.5 at q4f16, Apache-2.0; its native 262K context is capped to 32K to keep browser memory bounded.
- LFM2.5 1.2B (~760 MB) — text + agent (tool calling). Liquid AI's LFM2.5-1.2B-Instruct (Jan 2026) — same LFM2 hybrid arch as the 230M default, RL-tuned for instruction following + tool use. q4f16, 32K context.
- Gemini Nano (0 MB) — Google's on-device model built into Chrome, via the Prompt API. Zero download, no WebGPU. Chrome/Edge only; runs where your browser exposes
self.LanguageModel. When Nano is already downloaded on desktop Chrome/Edge, LocalMind starts on it by default so you can chat with no model download. - LFM2.5 230M · WebGPU kernels (~140 MB) — the same LFM2.5 230M run on a from-scratch WebGPU engine whose every kernel (RoPE, RMSNorm, Q4_0 dequant, the LFM2 short-conv, GQA attention) is hand-written WGSL, reading a Q4_0 GGUF directly (no ONNX, no llama.cpp). Ported from webml-community/lfm2-webgpu-kernels. WebGPU-only, tuned for maximal decode throughput.
- Gemma 4 E2B · WebGPU kernels (~2 GB) — text + agent. Google's Gemma 4 E2B (QAT mobile, int4) on a from-scratch WebGPU engine: hand-written WGSL for QAT int4 matmul, embed-gather-norm, RoPE/RMSNorm and GQA + sliding-window attention, reading the QAT-mobile weights directly. Ported from webml-community/gemma-4-webgpu-kernels. WebGPU-only. About 4× the ONNX Gemma 4 E2B’s decode speed on the same Mac (171 vs 42 tok/s, M4 Pro); text-only.
- Gemma 4 E4B · WebGPU kernels (~3.6 GB) — text + agent. The larger Gemma 4 (QAT mobile) on the same hand-written WebGPU engine. Its per-layer embedding table can also live on disk (Settings → Models, below).
- Gemma 4 E2B (~1.5 GB) — multimodal (image + audio) + agent.
- Gemma 4 E4B (~4.9 GB) — multimodal + agent, best quality.
- LFM2.5 230M · GGUF (~170 MB) — the same LFM2.5 230M through wllama (llama.cpp → WebAssembly), reading a Q5_K_M GGUF. WebGPU when your browser has it, CPU otherwise, so it also runs where WebGPU is missing.
- MiniCPM5 2B · GGUF (~1.6 GB) — text + agent (tool calling). OpenBMB’s MiniCPM5-2B (Sept 2026, Apache-2.0) loaded from its first-party GGUF through wllama (llama.cpp → WebAssembly) — there is no ONNX export, so this is the only way to run it in the tab. WebGPU when your browser has it, CPU otherwise. Gets 16K context, twice the other GGUF models, because its 2 KV heads make the cache unusually cheap; its native 131K is capped since wllama reserves the whole cache up front. Long-document retrieval is what it is best at. It emits tool calls in its own XML format rather than the
<tool_call>JSON the other models use; LocalMind parses both. Short-horizon tool use is its strength — don’t expect long multi-step agent runs to hold up. - Qwen3.6 35B-A3B · SSD-streamed (experimental) (~37 GB) — text, with reasoning. Alibaba's Qwen3.6-35B-A3B mixture-of-experts at Q8_0 (36.9 GB, more than most laptops' RAM) on a from-scratch WebGPU engine: the dense layers, router, KV cache and Gated-DeltaNet state stay on the GPU while the 10,240 routed experts stream from on-device storage (OPFS) as the router picks them. The first load copies the model into this browser's storage once (~37 GB of disk; a Hugging Face token in Settings speeds the download). About 9 tokens/s on an M4 Pro with 24 GB; prompts are read one token at a time for now. Greedy output matches llama.cpp. Flip Show reasoning for its chain-of-thought.
- Gemma 4 26B-A4B · SSD-streamed (experimental) (~14 GB) — text, with reasoning. Google's Gemma 4 26B-A4B mixture-of-experts (128 experts, 8 active per token) from its official 4-bit QAT GGUF (14.4 GB) on the same from-scratch WebGPU engine family: attention, the dense MLP, routers and KV cache stay on the GPU (~1.6 GB plus a 4 GB expert pool) while the 3,840 routed experts stream from on-device storage (OPFS) as the router picks them. The first load copies the model into this browser's storage once (~14 GB of disk; about 12 minutes on a fast connection). About 23 tokens/s on an M4 Pro with 24 GB, and prompts are read in chunks at about 55 tokens/s. Greedy output matches llama.cpp. Flip Show reasoning for its thinking.
onnx/. Multimodal
custom models are not yet supported. Added models appear in the model selector.
Authorization: Bearer, when LocalMind downloads
weights for the WebGPU-kernel models (LFM2.5, Gemma 4, Ternary Bonsai 2) or looks up a custom
model. Unlocks gated repos and lifts you from the anonymous per-IP rate limit to your
account's. Stored on this device only — never synced to the data folder, never sent anywhere else.
Create a read token.
POST <endpoint> with the model-generated
args as JSON body; the response JSON is fed back to the model. The endpoint must send
CORS headers for this origin. Name must be [a-zA-Z_][a-zA-Z0-9_]* and not
collide with a built-in.
Local servers (models that run on this machine, outside the browser)
OLLAMA_ORIGINS; LM Studio: enable CORS; Atomic Chat: add this site under Trusted Hosts).
Pick a preset to fill the URL: Ollama :11434/v1 · LM Studio :1234/v1 · llama.cpp :8080/v1 · Atomic :1337/v1 · dflash serve :8790/v1.
mcp_
to avoid collisions with built-ins. The server must allow CORS for this origin and
accept JSON-RPC 2.0 requests at the given URL. Connections re-established on page load.
Your image appears here. Describe it below and press Send.
chrome://flags/#enable-unsafe-webgpu). A lighter recognizer for low-end devices is planned for a later release. The rest of LocalMind keeps working.