Maple mmproj โ€” SigLIP2 NaFlex vision projector

Vision projector (mmproj) for Maple, for use with llama.cpp. It pairs with a Maple GGUF and gives the model image understanding.

llama-server ใงไฝฟใ†ใจใใฏ chat_template_nothink.jinja ใ‚’ --chat-template-file ใงๆธกใ—ใฆใใ ใ•ใ„ใ€‚ ๆธกใ•ใชใ„ใจๅฟœ็ญ”ใŒ็ฉบๆ–‡ๅญ—ๅˆ—ใซใชใ‚Šใ€ ใ€Œใƒ“ใ‚ธใƒงใƒณใŒๅฃŠใ‚Œใฆใ„ใ‚‹ใ€ใ‚ˆใ†ใซ่ฆ‹ใˆใพใ™ใ€‚่ฉณใ—ใใฏ Serving with llama-serverใ€‚ llama-mtmd-cli ใงใฏไธ่ฆใงใ™ใ€‚

Vision encoder google/siglip2-base-patch16-naflex (frozen)
Projector 2-layer MLP with GELU, 11.5M trainable params
File size 196.6 MB (F16)
Base model deepgrove/maple-preview (frozen)

The base Maple weights are frozen; only the projector is trained.

Usage

This requires a patched llama.cpp (see below). With it:

llama-mtmd-cli \
    -m maple-preview.gguf \
    --mmproj mmproj-maple-step8000.gguf \
    --image photo.jpg \
    -p "ใฉใ‚“ใชๅ‹•็‰ฉใŒใ„ใพใ™ใ‹๏ผŸ\nๆ—ฅๆœฌ่ชžใงใ€็Ÿญใ็ญ”ใˆใฆใใ ใ•ใ„ใ€‚" \
    -ngl 99 --repeat-penalty 1.3

-m is any GGUF conversion of deepgrove/maple-preview.

โ˜… Two prompt-side settings matter a lot

--repeat-penalty 1.3 โ€” without it the model falls into repetition loops on short VQA answers. Measured over 24 Japanese questions:

penalty Japanese answers correct
1.0 13/24 1/24
1.3 13/24 2/24
1.5 9/24 0/24

1.5 is too strong โ€” it damages Japanese vocabulary and drops the answer rate.

ๆ—ฅๆœฌ่ชžใงใ€็Ÿญใ็ญ”ใˆใฆใใ ใ•ใ„ใ€‚ ("Answer in Japanese, briefly.") โ€” adding this line roughly triples both the Japanese answer rate and the accuracy:

prompt Chinese output Japanese answers <think> loops correct
question only 4/24 6/24 6/24 2/24
+ ๆ—ฅๆœฌ่ชžใง็ญ”ใˆใฆใใ ใ•ใ„ 4/24 18/24 1/24 3/24
+ ๆ—ฅๆœฌ่ชžใงใ€็Ÿญใ็ญ”ใˆใฆใใ ใ•ใ„ 2/24 18/24 1/24 6/24
+ ๅฟ…ใšๆ—ฅๆœฌ่ชžใฎใฟใงโ€ฆ (too forceful) 3/24 11/24 5/24 2/24

This is inference-time only โ€” no retraining needed. The model's training data is pure Japanese, but because Japanese kanji share many glyphs with simplified Chinese, the projector slightly prefers Chinese continuations. A explicit language instruction removes almost all of it, and "briefly" fixes the tendency to over-explain. Making the instruction forceful (ๅฟ…ใšโ€ฆใฎใฟ) is counter-productive and brings the <think> loops back.

โ˜… Serving with llama-server

llama-mtmd-cli builds the prompt itself and works as-is, but llama-server needs two extra settings. Without them the model appears to return empty responses, and it is easy to wrongly conclude that the vision adapter is broken.

Maple's built-in chat template ends with:

{%- if add_generation_prompt %}
    {{- '<|im_start|>assistant\n<think>\n' }}
{%- endif %}

Maple is a reasoning model, so this is correct for standalone use โ€” but this projector was trained on <|im_start|>assistant\n{answer}, with no <think> block. llama-server always routes through the template, so the model starts "thinking" instead of answering, and points its whole generation budget at reasoning. The server then strips that reasoning out of content, leaving an empty string:

finish_reason: "length"      # ran out of tokens while still thinking
usage.completion_tokens: 20
content: ""                  # <- looks like the model is broken

The image is fine โ€” usage.prompt_tokens is 271 (256 image tokens + 15 text), confirming the mmproj is loaded and working. Only the output is affected.

Two files/settings fix it.

1. Use the supplied template โ€” chat_template_nothink.jinja, which is identical to the built-in one except for that single line:

 {%- if add_generation_prompt %}
-    {{- '<|im_start|>assistant\n<think>\n' }}
+    {{- '<|im_start|>assistant\n' }}
 {%- endif %}

2. Ask for the raw output with "reasoning_format": "none" (optional โ€” setting 1 alone is enough, but this guarantees the reasoning splitter cannot swallow the answer).

llama-server \
    -m maple-preview.gguf \
    --mmproj mmproj-maple-step8000.gguf \
    --chat-template-file chat_template_nothink.jinja \
    -ngl 99 --host 127.0.0.1 --port 8080
curl http://127.0.0.1:8080/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "messages": [{"role": "user", "content": [
    {"type": "image_url", "image_url": {"url": "data:image/png;base64,..."}},
    {"type": "text", "text": "ไฝ•ใŒๅ†™ใฃใฆใ„ใพใ™ใ‹๏ผŸ\nๆ—ฅๆœฌ่ชžใงใ€็Ÿญใ็ญ”ใˆใฆใใ ใ•ใ„ใ€‚"}
  ]}],
  "temperature": 0, "n_predict": 64, "reasoning_format": "none"
}'

Images must go through /v1/chat/completions with image_url. The /completion endpoint does not accept images (files is a // dummy there), and image_url must be the first element of content so the image tokens land where training put them.

With these settings a 24-sample Japanese VQA run completes in ~20 s (โ‰ˆ0.85 s/sample), versus ~30 s/sample through the PyTorch reference stack.

Required llama.cpp patch

llama.cpp picks the NaFlex resize grid with image_max_pixels (a pixel count) and floor rounding, but the reference implementation in transformers uses max_num_patches (a patch count) with a binary search and ceil rounding. The two disagree on most inputs, which changes the interpolated position embeddings and stops the trained projector from working correctly.

Measured over 200 distinct image sizes, only 38.5% of images got the same grid; tuning image_max_pixels alone raises that to at most 77.5%, because the underlying algorithm differs. For example:

800x600 input
  transformers  288x224  (252 patches)   <- training
  llama.cpp     288x208  (234 patches)   <- unpatched inference

The patch is in the maple-vlm branch:

https://github.com/shibadogcap/llama.cpp/tree/maple-vlm

It adds mtmd_image_preprocessor_naflex, which reproduces transformers.image_transforms.get_image_size_for_max_num_patches (binary search for the largest scale whose patch count stays within max_num_patches, then round each side up). max_num_patches is derived from the existing image_max_pixels key as image_max_pixels / patch_sizeยฒ, so no metadata change is needed.

Only PROJECTOR_TYPE_PHI4 uses the new preprocessor; calc_size_preserved_ratio, which 12 other projector types share, is untouched. Files without the key keep the previous behaviour.

Training

Trained as a LLaVA-style projector in two stages, with English and Japanese data.

Stage Data Purpose
1 captions (coco, textcaps) learn the vision-to-language mapping
2 VQA + captions, Japanese-heavy learn to answer

The share of Japanese is set by character count, not sample count. The projector's gradient is proportional to the number of supervised answer tokens, so a few long English configs can dominate even when most samples are Japanese:

config                    answer length   share of training signal
  aokvqa                       571 chars              30.5%
  gqa                          490 chars              26.1%
  allava_instruct_vflan4v      403 chars              21.5%
  cauldron_ja (x6)              24 chars               2.1%   <- Japanese
  ja_vg_vqa                     5 chars               0.3%   <- Japanese
  --------------------------------------------------------------
  Japanese 1.8%  /  English 98.2%

An earlier run with exactly this mixture, weighted 31% Japanese by sample count, made the model answer in English 3 times out of 10. Dropping the four long English configs and raising the Japanese weight brought the Japanese share to 38% by character count, and Japanese VQA loss fell from 2.7479 to 2.0904 in 500 steps.

Component License
mvp-lab/LLaVA-OneVision-1.5-Instruct-Data Apache-2.0
SakanaAI/JA-VG-VQA-500 CC-BY-4.0 (derived from Visual Genome, CC-BY-4.0)
turing-motors/Cauldron-JA see dataset card

Known limitations

  • Serving needs the extra settings above. Through llama-server the model returns nothing without --chat-template-file, because the built-in template opens a <think> block that this projector was not trained for. This is not a loading or vision problem โ€” the image tokens are present and correct.
  • Needs the "answer in Japanese, briefly" line. Without it only ~25% of answers come out in Japanese, and some contain simplified Chinese characters. This is a base-model behaviour the projector cannot override.
  • About 1 in 8 answers never leaves <think>. Maple is a reasoning model; the projector was trained without reasoning blocks, so it cannot steer the model out once it starts. --repeat-penalty 1.3 and the brief-answer instruction reduce this to ~1/24, but do not eliminate it.
  • Confuses similar subjects. e.g. a train may be called a boat.
  • Weak on charts and documents. Numeric and table QA did not improve much.
  • Accuracy is low in absolute terms. On 24 Japanese VQA questions with the recommended settings, 6 answers contained the reference answer (25%).

Evaluation

Japanese VQA (SakanaAI/JA-VG-VQA-500), 24 questions, served through llama-server. repeat_penalty 1.3, prompt suffix ๆ—ฅๆœฌ่ชžใงใ€็Ÿญใ็ญ”ใˆใฆใใ ใ•ใ„ใ€‚.

step Japanese answers correct <think> loops
7000 13/24 2/24 7/24
8000 (recommended) 17/24 6/24 1/24
8750 17/24 4/24 1/24

step 8750 is worse than 8000 despite 750 extra steps: answers grew from 136 to 175 characters and accuracy fell from 6/24 to 4/24. Past ~8000 steps the long Cauldron-JA answers appear to erode the short-answer behaviour. Training was stopped at 8780.

Stage 1 (captions only, before Cauldron-JA)

step 4000 step 4500
Loss (training format) 1.7367 1.4818
Loss with shuffled images 2.9373 3.3235
Shuffle delta +1.2006 +1.8417
Japanese VQA loss 2.7520 2.3587

The growing shuffle delta shows the model depends on the image being correct: feeding a mismatched image destroys the output entirely.

Downloads last month
36
GGUF
Model size
97.4M params
Architecture
clip
Hardware compatibility
Log In to add your hardware
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for shibadogcap/maple-mmproj

Quantized
(15)
this model