Prism Caption 3 Micro

Prism Caption 3 Micro is a chat-titling model: given the first user message of a conversation, it writes a short, specific, correctly formatted title (4-6 words, title case, naming the actual subject). It is the same ~350M model as Prism Caption 2.5 Micro, retrained on a 30,000-example dataset, fine-tuned with LoRA on LFM2-350M. Part of the Prism family of small, single-purpose models.

Evaluation

Prism Caption 3 comes in two sizes, Micro (354M) and Pico (135M). Both are compared below with their untuned base models and the previous generation, on the same inputs: 275 held-out topics that never appear in any training bank, and 30 hand-written, realistic multi-sentence chat openers (the training prompts are short templated phrasings, so the second set checks that the model generalizes beyond the template).

275 held-out topics

System Format issues Relevant 3-6 words Avg words Speed (per title)
Base LFM2-350M 231/275 247/275 90/275 10.2 69 ms
Prism Caption 2.5 Micro (published) 27/275 274/275 222/275 5.2 52 ms
Prism Caption 3 Micro 0/275 275/275 264/275 4.9 63 ms
Base SmolLM2-135M 273/275 255/275 4/275 20.8 121 ms
Prism Caption 3 Pico 2/275 275/275 219/275 5.3 48 ms

30 realistic multi-sentence openers

System Format issues Relevant 3-6 words Avg words Speed (per title)
Base LFM2-350M 25/30 27/30 9/30 14.2 85 ms
Prism Caption 2.5 Micro (published) 9/30 29/30 19/30 6.8 60 ms
Prism Caption 3 Micro 0/30 30/30 27/30 4.8 64 ms
Base SmolLM2-135M 29/30 29/30 1/30 20.3 126 ms
Prism Caption 3 Pico 2/30 29/30 25/30 4.7 47 ms

"Format issues" = formatting problems (too long, too terse, leaked preamble, trailing punctuation, multiline). "Relevant" = the title shares a content word with the message. "3-6 words" = title length within the target range. "Speed" = average wall-clock time to generate one title (up to 28 new tokens, greedy decoding, single request) with mlx-lm on an Apple M4 (16 GB); speeds from separate runs differ by roughly +/-15 ms, so treat gaps smaller than that as noise. The Caption 3 rows above are the published default build (MLX 6-bit); every build is in the table further down.

The relevance check is a word-overlap rule and the format check is rule-based: this is a regression-style check, not a human or model judge, and both Caption 3 sizes are near its ceiling. The Caption 2.5 Micro row was measured on the published (fused, 4-bit) weights; the 0/275 issues figure on its own model card was measured earlier with the unfused adapter loaded.

Formats

This repo holds both an MLX and a GGUF build, plus extra MLX versions:

Format Location Notes
MLX 6-bit (default) repo root (model.safetensors + config) mlx-lm on Apple Silicon. Best balance: no measurable quality loss versus fp16 at lower latency
MLX fp16 mlx-fp16/ Full-precision reference
MLX 4-bit, group size 32 mlx-4bit-g32/ Fastest MLX option; small quality trade-off versus 6-bit (see the table below)
GGUF Q4_K_M prism_caption_3_micro_Q4_K_M.gguf llama.cpp and compatible runtimes (LM Studio, Ollama, ...)
GGUF Q8_0 prism_caption_3_micro_Q8_0.gguf Higher-fidelity GGUF

Why not plain 4-bit

The model was fine-tuned on a 4-bit base and then fused. Fusing a LoRA into 4-bit weights and re-quantizing with the default settings loses part of the fine-tune: the plain 4-bit MLX build measured 16/275 formatting issues versus 0/275 for fp16 (same model, same prompts), so it is not published. The builds here fuse the LoRA at full precision first and quantize afterwards. The 6-bit build holds quality; the group-32 4-bit build is the faster option and trades a little quality for it (see the table: lengths loosen a bit, and Pico picks up a few format issues). GGUF K-quants held quality too.

All builds measured

275 held-out topics

Build Format issues Relevant 3-6 words Avg words Speed (per title)
MLX 6-bit (repo root, default) 0/275 275/275 264/275 4.9 63 ms
MLX fp16 (mlx-fp16/) 0/275 275/275 267/275 4.9 105 ms
MLX 4-bit group-32 (mlx-4bit-g32/) 0/275 275/275 240/275 5.3 56 ms
GGUF Q8_0 0/275 275/275 267/275 4.9 53 ms
GGUF Q4_K_M 0/275 275/275 236/275 5.3 77 ms
MLX 8-bit (measured, not published) 0/275 275/275 267/275 4.8 69 ms
MLX plain 4-bit (measured, not published) 16/275 275/275 205/275 5.6 54 ms

30 realistic multi-sentence openers

Build Format issues Relevant 3-6 words Avg words Speed (per title)
MLX 6-bit (repo root, default) 0/30 30/30 27/30 4.8 64 ms
MLX fp16 (mlx-fp16/) 1/30 30/30 27/30 5.0 106 ms
MLX 4-bit group-32 (mlx-4bit-g32/) 0/30 30/30 24/30 5.3 57 ms
GGUF Q8_0 1/30 30/30 27/30 5.0 77 ms
GGUF Q4_K_M 0/30 30/30 27/30 5.0 59 ms
MLX 8-bit (measured, not published) 1/30 30/30 27/30 4.9 74 ms
MLX plain 4-bit (measured, not published) 8/30 30/30 20/30 6.9 60 ms

GGUF rows were measured through llama-server (Metal), which adds a few ms of local HTTP overhead; MLX rows through mlx-lm.

Usage -- MLX

from mlx_lm import load, generate

model, tokenizer = load("VertexAGI/prism-caption-3-micro")   # 6-bit default. For mlx-fp16/ or mlx-4bit-g32/, download the repo and pass that subfolder's local path

messages = [{"role": "system", "content": (
    "You name chat conversations. Given the user's first message, reply with ONLY a short, specific chat title (4-6 words, title case, no quotes, no punctuation at the end, no preamble). The title MUST name the main subject of the message -- do not over-abbreviate into something vague. Nothing else -- just the title."
)}, {"role": "user", "content": "Any advice on how to fix a leaking kitchen faucet?"}]
text = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
print(generate(model, tokenizer, prompt=text, max_tokens=24))

Usage -- GGUF (llama.cpp)

hf download VertexAGI/prism-caption-3-micro prism_caption_3_micro_Q4_K_M.gguf --local-dir .
llama-cli -m prism_caption_3_micro_Q4_K_M.gguf -st -n 24 --temp 0 \
  -sys "You name chat conversations. Given the user's first message, reply with ONLY a short, specific chat title (4-6 words, title case, no quotes, no punctuation at the end, no preamble). The title MUST name the main subject of the message -- do not over-abbreviate into something vague. Nothing else -- just the title." \
  -p "Any advice on how to fix a leaking kitchen faucet?"

Model Details

Base model LiquidAI/LFM2-350M
Fine-tuning base checkpoint the 4-bit MLX checkpoint mlx-community/LFM2-350M-4bit
Architecture LFM2: hybrid short-convolution / attention
Fine-tuning method LoRA (rank 8, scale 20.0, 16 (full depth) layers; 2.998M (0.846%) trainable parameters)
Framework MLX / mlx-lm, on Apple Silicon
License LFM Open License v1.0, inherited from the LFM2-350M base model. Free for research/non-commercial use and for commercial use under $10M annual revenue.

Training Data

A 30,000-example chat-titling dataset (27,000 train / 3,000 validation), the 13,000-example set behind Prism Caption 2.5 Micro extended with 17,000 new examples over a much larger topic bank: 9,207 unique topics (up from 1,207), 28,367 unique prompts, 23,241 unique titles. Each example is a first user message paired with a teacher-written title. Teachers (fast models, cycled/switched adaptively by recent success rate):

Teacher Examples Share
openai/gpt-oss-20b (NIM) 23,690 79.0%
nvidia/nemotron-3.5-lightning-30b-a3b (NIM) 4,250 14.2%
poolside/laguna-s-2.1:free (OpenRouter) 1,027 3.4%
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning (NIM) 600 2.0%
openai/gpt-oss-120b (NIM, retired 2026-09-03) 433 1.4%

New examples were filtered for format (3-8 words, no preamble) and relevance (the title must share a content word with the message). The 275 evaluation topics are excluded from every training bank.

Training Procedure

  • Method: LoRA, rank 8, scale 20.0, dropout 0.0, 16 (full depth) layers; Adam, learning rate 1e-5, batch size 4, sequence length 256
  • Steps: 6,750 iterations (exactly one epoch of the 27,000 training examples), validation every 250 steps
  • Validation loss: 7.122 at initialization, best 0.333 at iteration 6,500 (final 0.345) -- the best checkpoint was used
  • Throughput: ~2.35 it/s, ~1,050 tokens/s, peak memory ~1.2 GB (Apple M4)
  • Hyperparameters are identical to Prism Caption 2.5 Micro, so the comparison with it isolates the data (13k to 30k examples, larger topic bank) rather than the recipe.
  • Release builds: the LoRA adapter was fused into the de-quantized base at full precision, then quantized (MLX 6-bit / group-32 4-bit) or converted to GGUF (Q8_0 / Q4_K_M) from that fp16 model.

Eval scripts and raw results for every build are in eval/.

Limitations

Trained on synthetic titles distilled from a mix of teacher models, so some stylistic inconsistency between teachers may remain. Validated only on English, conversational, everyday-topic inputs; highly technical or non-English inputs are untested. Titles are short and extractive-leaning; a message with several unrelated requests will get a title for only one of them. Evaluation uses rule-based checks on 275 + 30 prompts, which are near ceiling for both sizes, so differences between Micro and Pico on this eval are small and should not be over-read.

License

LFM Open License v1.0, inherited from the LFM2-350M base model. Free for research/non-commercial use and for commercial use under $10M annual revenue.

Downloads last month
202
Safetensors
Model size
0.4B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for VertexAGI/prism-caption-3-micro

Adapter
(23)
this model

Collection including VertexAGI/prism-caption-3-micro