You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

MIAI-VLM 0.2

Korean/English vision-language model: google/gemma-4-E4B-it fine-tuned with LoRA on a Korean-heavy mix of image-text and text data.

This repository holds two snapshots of the same adapter lineage. Stage 1 sits at the repository root as merged bf16 weights, ready to load directly. Stage 2 continues that adapter as a new run on a cleaned data mix and the official non-thinking chat format, and is published under stage2/ as a LoRA adapter while the run is still going.

stage 1 — repository root stage 2 — stage2/
Snapshot step 1,000,000 of 2,531,640 · epoch 1.19 of 3 step 325,000 of 562,997 · epoch 0.58 of 1 · training in progress
Published as merged bf16 weights + adapter/ stage2/adapter/ (LoRA adapter only)
Data 191 datasets · 54.0M samples · image-text 29% · Korean 46% 167 datasets · 36.0M samples · image-text 34% · Korean 44%
Sequence window 1,024 tokens, over-length truncated 1,280 tokens, over-length samples dropped
Chat format thinking format — enable_thinking=True non-thinking format — enable_thinking=False
Learning rate 2e-4 peak, cosine 1e-4 peak, cosine, restarted on the stage-1 adapter
Training loss 0.921 0.907

training loss across both stages

Stage 2 keeps the base model and the LoRA shape of stage 1 and changes what goes into them: a de-duplicated dataset selection (near-duplicate image sets and reasoning-trace sets removed, remaining reasoning blocks stripped to their final answers), samples longer than the window dropped rather than cut mid-answer, a 1024×1024 vision budget, and half the learning rate. Loss values are not comparable across the boundary — the data mix, the sequence window and the chat format all changed with it.

Shared setup

Base google/gemma-4-E4B-it (8.0B params incl. vision/audio towers)
Fine-tuning LoRA r=32, α=64 on all linear layers of the language model (69.8M trainable params); vision/audio towers frozen
Compute 16 × RTX 3090 (2 nodes), effective batch 64, bf16
Framework LLaMA-Factory 0.9.6 · transformers 5.6 · PEFT 0.18

Per-100-step loss logs are in training/trainer_state.json (stage 1) and stage2/training/trainer_state.json (stage 2); the full training configs sit beside them.

dataset composition

Top-20 dataset families of the stage-1 mix. Stage 2 draws on the same pool minus the removed sets; its percentages above are of the 42.5M selected rows, 15% of which the length and format filters then dropped.

Usage

The two stages expect different prompt formats. Stage 1 was trained with the system prompt You are a helpful assistant. and a thinking channel; stage 2 was trained on the plain non-thinking format with no default system prompt. Passing the wrong enable_thinking flag is the most common way to get degraded output.

Stage 2 — latest adapter

import torch
from PIL import Image
from transformers import AutoProcessor, AutoModelForImageTextToText
from peft import PeftModel

base_id = "google/gemma-4-E4B-it"
processor = AutoProcessor.from_pretrained(base_id)
model = AutoModelForImageTextToText.from_pretrained(base_id, dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(model, "Yong-Hoon/MIAI_VLM_0.2", subfolder="stage2/adapter").eval()

def chat(question, image=None, max_new_tokens=256):
    content = ([{"type": "image", "image": image}] if image is not None else []) + [{"type": "text", "text": question}]
    messages = [{"role": "user", "content": content}]
    inputs = processor.apply_chat_template(messages, add_generation_prompt=True, tokenize=True,
                                           return_dict=True, return_tensors="pt",
                                           enable_thinking=False).to(model.device)
    out = model.generate(**inputs, max_new_tokens=max_new_tokens, do_sample=False)
    return processor.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True).strip()

print(chat("이 사진에 무엇이 보이나요? 두 문장으로 설명해주세요.", image=Image.open("photo.jpg")))

The model answers directly — there is no thought block to strip. A system message is optional; when you use one, keep passing enable_thinking=False. Call model.merge_and_unload() to fold the adapter into the base weights for serving.

Stage 1 — merged weights at the repository root

import re, torch
from PIL import Image
from transformers import AutoProcessor, AutoModelForImageTextToText

model_id = "Yong-Hoon/MIAI_VLM_0.2"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(model_id, dtype=torch.bfloat16, device_map="auto").eval()

def chat(question, image=None, max_new_tokens=256):
    content = ([{"type": "image", "image": image}] if image is not None else []) + [{"type": "text", "text": question}]
    messages = [
        {"role": "system", "content": [{"type": "text", "text": "You are a helpful assistant."}]},
        {"role": "user", "content": content},
    ]
    inputs = processor.apply_chat_template(messages, add_generation_prompt=True, tokenize=True,
                                           return_dict=True, return_tensors="pt", enable_thinking=True).to(model.device)
    out = model.generate(**inputs, max_new_tokens=max_new_tokens, do_sample=False)
    text = processor.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=False)
    return re.sub(r"<\|channel\>thought\n.*?<channel\|>", "", text, flags=re.S).replace("<turn|>", "").strip()

print(chat("이 사진에 무엇이 보이나요? 두 문장으로 설명해주세요.", image=Image.open("photo.jpg")))

Stage 1 answers after an empty thought channel, which the snippet strips. Its adapter alone is at adapter/: PeftModel.from_pretrained(base_model, "Yong-Hoon/MIAI_VLM_0.2", subfolder="adapter").

Files

model-*.safetensors            stage 1, merged bf16 (~16 GB)
tokenizer / processor configs  shared by both stages
adapter/                       stage 1 LoRA adapter (~280 MB)
training/                      stage 1 config + per-100-step loss log
stage2/adapter/                stage 2 LoRA adapter (~280 MB)
stage2/training/               stage 2 config + per-100-step loss log
assets/                        charts

Both stages are intermediate snapshots of runs in progress; later snapshots follow as new commits. Governed by the Gemma license. Built with LLaMA-Factory. Developed by Yong-Hoon (KETI).

Downloads last month
-
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Yong-Hoon/MIAI_VLM_0.2

Adapter
(389)
this model