loopl — Agents that run on your phone.

license: Apache 2.0 bf16: transformers · 9.08 GB vision: ✓ 297 tensors kept tool calling: native + loopl JSON params: 4B version: v1.1 Get it on: TestFlight site: loopl.dev collection: 6 models

loopl 4B · v1.1

An agent that lives on your phone. loopl runs the model, the loop and the tools on the device. It talks, uses the phone, asks before it acts — and works with the network off. This repo is the merged bf16 weights — the base for Train your own — of the 4B loopl model, a Qwen3.5 post-tuned to run loopl's agent loop well, know who it is, and read images.

It uses the phone — every tool call is a row.It asks before anything leaves the device.Show it a photo. It reads it on-device.
It uses the phone — every tool call is a row.It asks before anything leaves the device.Show it a photo. It reads it on-device.

Real captures from the free iPhone app.

What loopl is

  • A Swift SDK — Agent(model:tools:) and a small, honest loop: model → tools → model until it answers. Tool errors go back to the model (it fixes the call instead of repeating it); anything that leaves the device waits for a tap. Runtimes for MLX, llama.cpp and Apple's on-device model; hooks, checkpoints, structured output, tracing.
  • A free iPhone app built on it — pick a model, talk to it in airplane mode, let it use the phone (torch, haptics, translation, speech, image generation, other apps, memory, HTTP), show it photos, record a take, and train it on your own conversations from the phone. No account, nothing collected.
  • These models — the same Qwen3.5 the app offers, post-tuned so the loop, the identity and the tool manners are in the weights, not in a system prompt you cannot see.

loopl.dev · Get it on TestFlight · github.com/cagataycali/loopl

This model

probe this model previous (v1) base Qwen3.5 4-bit
agent-SDK knowledge quiz (first 472 questions, exact-answer) 158 111 83
tool calls, native template dialect 9/9 9/9 9/9
tool calls, loopl JSON dialect 9/9 9/9 9/9
identity (5 classic probes, with system prompt) 5/5 4/5 4/5
identity · bio facts (8, with system prompt) 8/8 — —
identity · no system prompt (13) 11/13 — —

Scored by train/eval.py on the 4-bit MLX export (greedy). The quiz asks for exact file-level facts about an agent SDK's source — hard for every small model; the delta over base is the point. Tool probes count a well-formed call with the right arguments. identity_nosys is the one that matters in the app: the system prompt there is the user's to edit.

Vision: the base's tower is frozen and kept — 297 vision_tower.* tensors in the MLX export, so the app shows the photo button.

Use it today

In the loopl app (iPhone). Install from TestFlight → Models → the loopl models section lists this size (no account, no token) → download. The app sees the vision tower and shows the photo button; the system prompt is yours to edit — identity and tool manners are in the weights.

On a Mac with MLX (vision included):

pip install -U mlx-vlm
python -m mlx_vlm.generate --model cagataydev/loopl-4b-4bit \
  --image photo.jpg --prompt "What is in this picture? Then tell me who built you." --max-tokens 200

With llama.cpp: the GGUF Q4_K_M export (cagataydev/loopl-4b-GGUF) is not public yet; export your own from the bf16 weights with llama.cpp's convert_hf_to_gguf.py --no-mtp (text only — the converter has no vision path).

From Swift with the loopl SDK — the snippet below is docs/start.md § "Run the loop" in the loopl repo, compile-checked in CI against the package (products Loopl + LooplRuntimes); spec is this repo's ModelSpec and dir the downloaded folder:

import Loopl
import LooplRuntimes
import Foundation

func chat(_ spec: ModelSpec, at dir: URL) async throws {
    let model = MLXModel()                                       // Metal — a real device
    try await model.load(spec, from: dir) { _ in }

    let agent = Agent(model: model,
                      tools: [CurrentTimeTool(), CalculatorTool()],
                      systemPrompt: "You are a concise assistant on the user's phone.")

    for try await event in agent.stream("What time is it in Istanbul, and what is 23 × 47?") {
        switch event {
        case .text(let t):            print(t, terminator: "")
        case .toolStarted(let use):   print("\n→ \(use.name)")
        case .toolResult(let r):      print("← \(r.content)")
        default: break
        }
    }
}

Keep post-tuning it. The bf16 weights (cagataydev/loopl-4b) are the base for continual post-tuning: in the app, Models › Train your own takes a loopl model as the base and trains on your own shared conversations (train/loopl_sft.py --base cagataydev/loopl-4b); on a Mac they load with transformers ≥ 5 as Qwen3_5ForConditionalGeneration. Export your own quantisation with mlx_vlm.convert -q (keeps vision) — mlx_lm.convert silently drops the tower.

Recipe

base Qwen/Qwen3.5-4B (vision tower frozen, visual.* never trained)
data cagataydev/loopl-train (private) @ db6fa452 · 10596 rendered rows / 16,710,241 tokens · both tool dialects · oversample c=2
method LoRA r=64 α=128 on the language model · lr 0.0001 · batch 2×8 · max_len 4096 · 2.0 epochs, best-epoch checkpoint kept
eval loss ep1 1.643, ep2 1.712 → best 1.643 (eval split of the same dataset revision)
compute a100-large (Hugging Face Jobs), 14064 s train · job 6ac4fa06404719ba3765fc0d
stack transformers 5.18.0 · trl 1.14.1 · peft 0.21.2
exports MLX 4-bit via mlx_vlm.convert -q (297 vision tensors of 1221, 3.03 GB) · GGUF Q4_K_M via llama.cpp --no-mtp (2.71 GB, text only)

Script: train/loopl_sft.py — the same single file the loopl app launches when you tap Models › Train your own on your own conversations. Assistant turns that seed a bad tool call carry weight: 0 so the model learns the recovery, not the mistake. Tokenizer files are normalised to the base's (Qwen2Tokenizer, chat_template.jinja with tools) so Swift loaders accept them; train/check_mlx_repo.py gates every export on that plus the vision triad (vision_config ⇔ vision_tower.* ⇔ preprocessor_config.json).

Limitations

  • A 4B model: fluent and well-behaved in the loop, not an encyclopedia. It will get arithmetic and obscure facts wrong; give it tools (calculator, http, recall) and it does better.
  • Identity holds in most no-prompt probes (see the scores), not all; a one-line system prompt ("You are loopl…") makes it consistent.
  • The GGUF export is text-only. Vision needs the MLX export (or the bf16 weights).
  • Trained on English plus a little Turkish; other languages are the base's.
  • Knowledge about agent SDKs is file-level and dated to the dataset revision; it does not know your repo.

Data and privacy

Every training row is agent-synthesised from public sources (the owner's public GitHub repos and the agent-SDK source they build on), critic-reviewed (score ≥ 4/5 kept) and filtered by a denylist for secrets, private names, phone numbers and addresses — see the dataset card. No user conversations from the app are in this model.

License

Apache-2.0, inherited from the Qwen3.5 base. The fine-tune, exports and dataset are © Cagatay Cali, same license.

The family

Six public repos, one recipe. The app lists the three MLX rows under Models › loopl models (2B is the sweet spot for speed on an iPhone 15/16; 4B is the most knowledgeable; 0.8B fits anywhere). Every bf16 repo is a valid --base for the next round.

MLX 4-bit · vision · what the app downloads bf16 · the base to keep training
0.8B cagataydev/loopl-0.8b-4bit · 625 MB cagataydev/loopl-0.8b · 1.71 GB
2B cagataydev/loopl-2b-4bit · 1.72 GB cagataydev/loopl-2b · 4.43 GB
4B cagataydev/loopl-4b-4bit · 3.03 GB cagataydev/loopl-4b · 9.08 GB ← this repo

All of them: the loopl collection.

Links

Downloads last month
-
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for cagataydev/loopl-4b

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(883)
this model
Quantizations
1 model

Collection including cagataydev/loopl-4b