Drop image, audio, or video here

Starting…

Private AI on this device

Models run in this browser. What you type is not sent to any server; web search runs only when you turn it on.

📱 On a phone, big models won’t fit. Point LocalMind at your computer’s Ollama / LM Studio (Settings → Models → Local endpoint) and use any model over your network.

Preparing to download model...
Loading SAM model...
Mirror all data (skills, memory, documents, conversations) to a folder on disk — point it at a git / iCloud / Dropbox folder to keep it permanent and synced across devices. API keys stay in this browser only.
Not connected — data is stored in this browser only.
About each model
  • Ternary Bonsai 4B (~1.1 GB) — text + agent (tool calling). 1.58-bit ternary weights, Qwen3 backbone, Apache-2.0 — the small ternary tier.
  • Ternary Bonsai 2 27B (~5.9 GB) — text + agent. PrismML's Ternary Bonsai 2 27B (Sept 2026; ternary 1.75-bit weights, Qwen3.8-27B backbone, 98% of the full-precision model's benchmark score) on a from-scratch WebGPU engine: hand-written WGSL for ternary dequant + Gated-DeltaNet linear attention, reading the PTQ1_0 GGUF directly. Ported from webml-community/ternary-bonsai-2-webgpu-kernels; Apache-2.0 weights. WebGPU-only — the biggest in-browser model here. A reasoning model: flip Show reasoning in Settings for its chain-of-thought. Text-only here (its optional vision tower is not loaded).
  • Underdog Saluki 27B (~7.9 GB) — text + agent. ConwayResearch's Underdog Saluki 27B 1.0 (Oct 2026): Qwen3.8-27B as one 2-bit-class IQ2-mix GGUF, tuned to keep tool calling intact (the vendor reports 88 of 120 on a frozen BFCL v4 tool-calling set, against 84 for the full-size model and 70 for Ternary Bonsai 2). Runs on a from-scratch WebGPU engine that reads the GGUF’s eleven llama.cpp block types directly; greedy output matches llama.cpp on the same file. The first load copies the 7.9 GB file into this browser’s on-device storage once (or pick the file you already have with Load .gguf from disk…); it needs about 8 GB of GPU memory. Apache-2.0 weights. A reasoning model: flip Show reasoning in Settings for its chain-of-thought. Text-only here (its optional vision add-on is not loaded).
  • Qwen3.5 4B (~3 GB) — text + image + agent (tool calling). Qwen3.5 at q4f16, Apache-2.0; its native 262K context is capped to 32K to keep browser memory bounded.
  • LFM2.5 1.2B (~760 MB) — text + agent (tool calling). Liquid AI's LFM2.5-1.2B-Instruct (Jan 2026) — same LFM2 hybrid arch as the 230M default, RL-tuned for instruction following + tool use. q4f16, 32K context.
  • Gemini Nano (0 MB) — Google's on-device model built into Chrome, via the Prompt API. Zero download, no WebGPU. Chrome/Edge only; runs where your browser exposes self.LanguageModel. When Nano is already downloaded on desktop Chrome/Edge, LocalMind starts on it by default so you can chat with no model download.
  • LFM2.5 230M · WebGPU kernels (~140 MB) — the same LFM2.5 230M run on a from-scratch WebGPU engine whose every kernel (RoPE, RMSNorm, Q4_0 dequant, the LFM2 short-conv, GQA attention) is hand-written WGSL, reading a Q4_0 GGUF directly (no ONNX, no llama.cpp). Ported from webml-community/lfm2-webgpu-kernels. WebGPU-only, tuned for maximal decode throughput.
  • Gemma 4 E2B · WebGPU kernels (~2 GB) — text + agent. Google's Gemma 4 E2B (QAT mobile, int4) on a from-scratch WebGPU engine: hand-written WGSL for QAT int4 matmul, embed-gather-norm, RoPE/RMSNorm and GQA + sliding-window attention, reading the QAT-mobile weights directly. Ported from webml-community/gemma-4-webgpu-kernels. WebGPU-only. About 4× the ONNX Gemma 4 E2B’s decode speed on the same Mac (171 vs 42 tok/s, M4 Pro); text-only.
  • Gemma 4 E4B · WebGPU kernels (~3.6 GB) — text + agent. The larger Gemma 4 (QAT mobile) on the same hand-written WebGPU engine. Its per-layer embedding table can also live on disk (Settings → Models, below).
  • Gemma 4 E2B (~1.5 GB) — multimodal (image + audio) + agent.
  • Gemma 4 E4B (~4.9 GB) — multimodal + agent, best quality.
  • LFM2.5 230M · GGUF (~170 MB) — the same LFM2.5 230M through wllama (llama.cpp → WebAssembly), reading a Q5_K_M GGUF. WebGPU when your browser has it, CPU otherwise, so it also runs where WebGPU is missing.
  • MiniCPM5 2B · GGUF (~1.6 GB) — text + agent (tool calling). OpenBMB’s MiniCPM5-2B (Sept 2026, Apache-2.0) loaded from its first-party GGUF through wllama (llama.cpp → WebAssembly) — there is no ONNX export, so this is the only way to run it in the tab. WebGPU when your browser has it, CPU otherwise. Gets 16K context, twice the other GGUF models, because its 2 KV heads make the cache unusually cheap; its native 131K is capped since wllama reserves the whole cache up front. Long-document retrieval is what it is best at. It emits tool calls in its own XML format rather than the <tool_call> JSON the other models use; LocalMind parses both. Short-horizon tool use is its strength — don’t expect long multi-step agent runs to hold up.
  • Qwen3.6 35B-A3B · SSD-streamed (experimental) (~37 GB) — text, with reasoning. Alibaba's Qwen3.6-35B-A3B mixture-of-experts at Q8_0 (36.9 GB, more than most laptops' RAM) on a from-scratch WebGPU engine: the dense layers, router, KV cache and Gated-DeltaNet state stay on the GPU while the 10,240 routed experts stream from on-device storage (OPFS) as the router picks them. The first load copies the model into this browser's storage once (~37 GB of disk; a Hugging Face token in Settings speeds the download). About 9 tokens/s on an M4 Pro with 24 GB; prompts are read one token at a time for now. Greedy output matches llama.cpp. Flip Show reasoning for its chain-of-thought.
  • Gemma 4 26B-A4B · SSD-streamed (experimental) (~14 GB) — text, with reasoning. Google's Gemma 4 26B-A4B mixture-of-experts (128 experts, 8 active per token) from its official 4-bit QAT GGUF (14.4 GB) on the same from-scratch WebGPU engine family: attention, the dense MLP, routers and KV cache stay on the GPU (~1.6 GB plus a 4 GB expert pool) while the 3,840 routed experts stream from on-device storage (OPFS) as the router picks them. The first load copies the model into this browser's storage once (~14 GB of disk; about 12 minutes on a fast connection). About 23 tokens/s on an M4 Pro with 24 GB, and prompts are read in chunks at about 55 tokens/s. Greedy output matches llama.cpp. Flip Show reasoning for its thinking.
Loading...
no download · this session · Gemma 4 26B-A4B or Qwen3.6 35B-A3B's GGUF from your Hugging Face or LM Studio folder skips their download (macOS picker: ⌘⇧. shows ~/.cache)
Must be a Hugging Face repo with ONNX files under onnx/. Multimodal custom models are not yet supported. Added models appear in the model selector.
After the model loads, fetches a 1.1 GB DFlash 2 drafter (once, cached) and lets the 27B model check several drafted tokens per step: ~1.1–1.2× faster on code and structured text, unchanged output. Uses ~1 GB more GPU memory. Applies to the next model load.
Keeps the model's 1.2 GB per-layer embedding table in this browser's private storage and only its ~33K most-used rows on the GPU: ~1 GB less GPU memory, unchanged output. Decoding is up to ~8% slower on new text and ~3–5% slower on repeated text or code. The first load writes the table to disk once (~1 s). Also applies to Gemma 4 E4B (WebGPU kernels), whose 2-bit table is 749 MB. Applies to the next model load.
Sent only to huggingface.co, as Authorization: Bearer, when LocalMind downloads weights for the WebGPU-kernel models (LFM2.5, Gemma 4, Ternary Bonsai 2) or looks up a custom model. Unlocks gated repos and lifts you from the anonymous per-IP rate limit to your account's. Stored on this device only — never synced to the data folder, never sent anywhere else. Create a read token.
On a tool call, LocalMind does POST <endpoint> with the model-generated args as JSON body; the response JSON is fed back to the model. The endpoint must send CORS headers for this origin. Name must be [a-zA-Z_][a-zA-Z0-9_]* and not collide with a built-in.

Local servers (models that run on this machine, outside the browser)

0 = greedy. Required for lossless speculative decoding on dflash serve.
Run bigger / faster models without the WebGPU limit — inference runs on your machine, nothing leaves your device. Discovered models appear in the model menu above. The server must allow CORS for this origin (Ollama: set OLLAMA_ORIGINS; LM Studio: enable CORS; Atomic Chat: add this site under Trusted Hosts). Pick a preset to fill the URL: Ollama :11434/v1 · LM Studio :1234/v1 · llama.cpp :8080/v1 · Atomic :1337/v1 · dflash serve :8790/v1.
Used when the Image screen's model is "Local server". Same field as the Image panel's Server box.
Tools discovered from an MCP server are registered with the prefix mcp_ to avoid collisions with built-ins. The server must allow CORS for this origin and accept JSON-RPC 2.0 requests at the given URL. Connections re-established on page load.
Lets other JavaScript on this page call the loaded model via an OpenAI-shaped method. Same-tab only — cross-origin scripts cannot reach it. No tool calling, memory access, or web search is exposed.
API
Pick 2–3 models 0 selected
On-device image generation
Type a prompt below and Send. First run downloads the model (~3 GB, cached after). Best on Chrome/Edge.

Your image appears here. Describe it below and press Send.

Masked-diffusion text · Qwen3 0.6B (~1.4 GB)
Type a prompt below and Send. First run downloads ~1.4 GB (cached after). Watch the answer denoise out of the fog. Best on Chrome/Edge.
Document OCR · GLM-OCR 0.9B (~1.4 GB) · on-device
Drop an image or PDF here, or choose a file. The document never leaves your device.
Choose an image or PDF, pick a format, then Extract. First use downloads the model once (cached after, then offline). WebGPU — best on Chrome/Edge.
Voice · Moonshine STT + Kokoro TTS · on-device
Hold the mic (or press and hold V) to talk. Release to send — it transcribes, answers, and reads the reply back. Everything stays on your device.
Vision · SmolVLM-500M · on-device (WebGPU)
Start the camera or pick an image, then Describe. First use downloads SmolVLM (~0.9 GB) once, then runs offline. WebGPU — best on Chrome/Edge.
Voice Clone · Chatterbox 0.5B · on-device (WebGPU)
No reference voice yet.
Record or upload a 5–10s voice sample, then type text and Speak. First use downloads Chatterbox (~1.5 GB) once, then runs offline. WebGPU — best on Chrome/Edge.