Maple mmproj โ SigLIP2 NaFlex vision projector
Vision projector (mmproj) for Maple,
for use with llama.cpp. It pairs with a Maple GGUF and gives the model image
understanding.
llama-serverใงไฝฟใใจใใฏchat_template_nothink.jinjaใ--chat-template-fileใงๆธกใใฆใใ ใใใ ๆธกใใชใใจๅฟ็ญใ็ฉบๆๅญๅใซใชใใ ใใใธใงใณใๅฃใใฆใใใใใใซ่ฆใใพใใ่ฉณใใใฏ Serving withllama-serverใllama-mtmd-cliใงใฏไธ่ฆใงใใ
| Vision encoder | google/siglip2-base-patch16-naflex (frozen) |
| Projector | 2-layer MLP with GELU, 11.5M trainable params |
| File size | 196.6 MB (F16) |
| Base model | deepgrove/maple-preview (frozen) |
The base Maple weights are frozen; only the projector is trained.
Usage
This requires a patched llama.cpp (see below). With it:
llama-mtmd-cli \
-m maple-preview.gguf \
--mmproj mmproj-maple-step8000.gguf \
--image photo.jpg \
-p "ใฉใใชๅ็ฉใใใพใใ๏ผ\nๆฅๆฌ่ชใงใ็ญใ็ญใใฆใใ ใใใ" \
-ngl 99 --repeat-penalty 1.3
-m is any GGUF conversion of
deepgrove/maple-preview.
โ Two prompt-side settings matter a lot
--repeat-penalty 1.3 โ without it the model falls into repetition loops
on short VQA answers. Measured over 24 Japanese questions:
| penalty | Japanese answers | correct |
|---|---|---|
| 1.0 | 13/24 | 1/24 |
| 1.3 | 13/24 | 2/24 |
| 1.5 | 9/24 | 0/24 |
1.5 is too strong โ it damages Japanese vocabulary and drops the answer rate.
ๆฅๆฌ่ชใงใ็ญใ็ญใใฆใใ ใใใ ("Answer in Japanese, briefly.") โ adding
this line roughly triples both the Japanese answer rate and the accuracy:
| prompt | Chinese output | Japanese answers | <think> loops |
correct |
|---|---|---|---|---|
| question only | 4/24 | 6/24 | 6/24 | 2/24 |
+ ๆฅๆฌ่ชใง็ญใใฆใใ ใใ |
4/24 | 18/24 | 1/24 | 3/24 |
+ ๆฅๆฌ่ชใงใ็ญใ็ญใใฆใใ ใใ |
2/24 | 18/24 | 1/24 | 6/24 |
+ ๅฟ
ใๆฅๆฌ่ชใฎใฟใงโฆ (too forceful) |
3/24 | 11/24 | 5/24 | 2/24 |
This is inference-time only โ no retraining needed. The model's training
data is pure Japanese, but because Japanese kanji share many glyphs with
simplified Chinese, the projector slightly prefers Chinese continuations. A
explicit language instruction removes almost all of it, and "briefly" fixes
the tendency to over-explain. Making the instruction forceful (ๅฟ
ใโฆใฎใฟ) is
counter-productive and brings the <think> loops back.
โ
Serving with llama-server
llama-mtmd-cli builds the prompt itself and works as-is, but
llama-server needs two extra settings. Without them the model appears to
return empty responses, and it is easy to wrongly conclude that the vision
adapter is broken.
Maple's built-in chat template ends with:
{%- if add_generation_prompt %}
{{- '<|im_start|>assistant\n<think>\n' }}
{%- endif %}
Maple is a reasoning model, so this is correct for standalone use โ but this
projector was trained on <|im_start|>assistant\n{answer}, with no
<think> block. llama-server always routes through the template, so the
model starts "thinking" instead of answering, and points its whole generation
budget at reasoning. The server then strips that reasoning out of
content, leaving an empty string:
finish_reason: "length" # ran out of tokens while still thinking
usage.completion_tokens: 20
content: "" # <- looks like the model is broken
The image is fine โ usage.prompt_tokens is 271 (256 image tokens + 15
text), confirming the mmproj is loaded and working. Only the output is affected.
Two files/settings fix it.
1. Use the supplied template โ chat_template_nothink.jinja,
which is identical to the built-in one except for that single line:
{%- if add_generation_prompt %}
- {{- '<|im_start|>assistant\n<think>\n' }}
+ {{- '<|im_start|>assistant\n' }}
{%- endif %}
2. Ask for the raw output with "reasoning_format": "none" (optional โ
setting 1 alone is enough, but this guarantees the reasoning splitter cannot
swallow the answer).
llama-server \
-m maple-preview.gguf \
--mmproj mmproj-maple-step8000.gguf \
--chat-template-file chat_template_nothink.jinja \
-ngl 99 --host 127.0.0.1 --port 8080
curl http://127.0.0.1:8080/v1/chat/completions -H 'Content-Type: application/json' -d '{
"messages": [{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": "data:image/png;base64,..."}},
{"type": "text", "text": "ไฝใๅใฃใฆใใพใใ๏ผ\nๆฅๆฌ่ชใงใ็ญใ็ญใใฆใใ ใใใ"}
]}],
"temperature": 0, "n_predict": 64, "reasoning_format": "none"
}'
Images must go through /v1/chat/completions with image_url. The
/completion endpoint does not accept images (files is a // dummy
there), and image_url must be the first element of content so the image
tokens land where training put them.
With these settings a 24-sample Japanese VQA run completes in ~20 s (โ0.85 s/sample), versus ~30 s/sample through the PyTorch reference stack.
Required llama.cpp patch
llama.cpp picks the NaFlex resize grid with image_max_pixels (a pixel count)
and floor rounding, but the reference implementation in transformers uses
max_num_patches (a patch count) with a binary search and ceil rounding.
The two disagree on most inputs, which changes the interpolated position
embeddings and stops the trained projector from working correctly.
Measured over 200 distinct image sizes, only 38.5% of images got the same
grid; tuning image_max_pixels alone raises that to at most 77.5%, because the
underlying algorithm differs. For example:
800x600 input
transformers 288x224 (252 patches) <- training
llama.cpp 288x208 (234 patches) <- unpatched inference
The patch is in the maple-vlm branch:
https://github.com/shibadogcap/llama.cpp/tree/maple-vlm
It adds mtmd_image_preprocessor_naflex, which reproduces
transformers.image_transforms.get_image_size_for_max_num_patches
(binary search for the largest scale whose patch count stays within
max_num_patches, then round each side up). max_num_patches is derived from
the existing image_max_pixels key as image_max_pixels / patch_sizeยฒ, so no
metadata change is needed.
Only PROJECTOR_TYPE_PHI4 uses the new preprocessor; calc_size_preserved_ratio,
which 12 other projector types share, is untouched. Files without the key keep
the previous behaviour.
Training
Trained as a LLaVA-style projector in two stages, with English and Japanese data.
| Stage | Data | Purpose |
|---|---|---|
| 1 | captions (coco, textcaps) |
learn the vision-to-language mapping |
| 2 | VQA + captions, Japanese-heavy | learn to answer |
The share of Japanese is set by character count, not sample count. The projector's gradient is proportional to the number of supervised answer tokens, so a few long English configs can dominate even when most samples are Japanese:
config answer length share of training signal
aokvqa 571 chars 30.5%
gqa 490 chars 26.1%
allava_instruct_vflan4v 403 chars 21.5%
cauldron_ja (x6) 24 chars 2.1% <- Japanese
ja_vg_vqa 5 chars 0.3% <- Japanese
--------------------------------------------------------------
Japanese 1.8% / English 98.2%
An earlier run with exactly this mixture, weighted 31% Japanese by sample count, made the model answer in English 3 times out of 10. Dropping the four long English configs and raising the Japanese weight brought the Japanese share to 38% by character count, and Japanese VQA loss fell from 2.7479 to 2.0904 in 500 steps.
| Component | License |
|---|---|
mvp-lab/LLaVA-OneVision-1.5-Instruct-Data |
Apache-2.0 |
SakanaAI/JA-VG-VQA-500 |
CC-BY-4.0 (derived from Visual Genome, CC-BY-4.0) |
turing-motors/Cauldron-JA |
see dataset card |
Known limitations
- Serving needs the extra settings above. Through
llama-serverthe model returns nothing without--chat-template-file, because the built-in template opens a<think>block that this projector was not trained for. This is not a loading or vision problem โ the image tokens are present and correct. - Needs the "answer in Japanese, briefly" line. Without it only ~25% of answers come out in Japanese, and some contain simplified Chinese characters. This is a base-model behaviour the projector cannot override.
- About 1 in 8 answers never leaves
<think>. Maple is a reasoning model; the projector was trained without reasoning blocks, so it cannot steer the model out once it starts.--repeat-penalty 1.3and the brief-answer instruction reduce this to ~1/24, but do not eliminate it. - Confuses similar subjects. e.g. a train may be called a boat.
- Weak on charts and documents. Numeric and table QA did not improve much.
- Accuracy is low in absolute terms. On 24 Japanese VQA questions with the recommended settings, 6 answers contained the reference answer (25%).
Evaluation
Japanese VQA (SakanaAI/JA-VG-VQA-500), 24 questions, served through
llama-server. repeat_penalty 1.3, prompt suffix
ๆฅๆฌ่ชใงใ็ญใ็ญใใฆใใ ใใใ.
| step | Japanese answers | correct | <think> loops |
|---|---|---|---|
| 7000 | 13/24 | 2/24 | 7/24 |
| 8000 (recommended) | 17/24 | 6/24 | 1/24 |
| 8750 | 17/24 | 4/24 | 1/24 |
step 8750 is worse than 8000 despite 750 extra steps: answers grew from 136
to 175 characters and accuracy fell from 6/24 to 4/24. Past ~8000 steps the
long Cauldron-JA answers appear to erode the short-answer behaviour. Training
was stopped at 8780.
Stage 1 (captions only, before Cauldron-JA)
| step 4000 | step 4500 | |
|---|---|---|
| Loss (training format) | 1.7367 | 1.4818 |
| Loss with shuffled images | 2.9373 | 3.3235 |
| Shuffle delta | +1.2006 | +1.8417 |
| Japanese VQA loss | 2.7520 | 2.3587 |
The growing shuffle delta shows the model depends on the image being correct: feeding a mismatched image destroys the output entirely.
- Downloads last month
- 36
Model tree for shibadogcap/maple-mmproj
Base model
deepgrove/maple-preview