Instructions to use zfz04/FocusVTC with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use zfz04/FocusVTC with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="zfz04/FocusVTC") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("zfz04/FocusVTC") model = AutoModelForMultimodalLM.from_pretrained("zfz04/FocusVTC", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use zfz04/FocusVTC with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "zfz04/FocusVTC" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zfz04/FocusVTC", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/zfz04/FocusVTC
- SGLang
How to use zfz04/FocusVTC with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "zfz04/FocusVTC" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zfz04/FocusVTC", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "zfz04/FocusVTC" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zfz04/FocusVTC", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use zfz04/FocusVTC with Docker Model Runner:
docker model run hf.co/zfz04/FocusVTC
FocusVTC
FocusVTC is a 9B vision-language model based on Qwen3.5. It is optimized for long-context visual document understanding and evidence-focused tool use. The model can reason over document-page images, decide when a higher-resolution crop is useful, and continue answering after receiving the cropped region as a new visual observation.
Quick start
FocusVTC requires a recent Transformers version with Qwen3.5 support.
pip install "transformers>=5.16.1" accelerate pillow
from PIL import Image
from transformers import AutoModelForImageTextToText, AutoProcessor
model_id = "zfz04/FocusVTC"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
model_id,
dtype="auto",
device_map="auto",
)
image = Image.open("document_page.png").convert("RGB")
messages = [
{
"role": "user",
"content": [
{"type": "image"},
{"type": "text", "text": "Answer the question using the document."},
],
}
]
prompt = processor.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = processor(
text=[prompt],
images=[image],
return_tensors="pt",
).to(model.device)
output_ids = model.generate(**inputs, max_new_tokens=512)
new_tokens = output_ids[:, inputs.input_ids.shape[1]:]
answer = processor.batch_decode(new_tokens, skip_special_tokens=True)[0]
print(answer)
For multi-page documents, add one image placeholder per page and pass the page images to the processor in the same order.
Visual tool use
FocusVTC is designed to work in an agent loop with the following tool:
{
"type": "function",
"function": {
"name": "zoom_region",
"description": "Crop a potentially unreadable region from a document page and return it as a new image.",
"parameters": {
"type": "object",
"properties": {
"page": {
"type": "integer",
"description": "1-based page number."
},
"bbox_2d": {
"type": "array",
"items": {"type": "number"},
"minItems": 4,
"maxItems": 4,
"description": "[x1, y1, x2, y2] normalized to [0, 1000]."
}
},
"required": ["page", "bbox_2d"]
}
}
}
The host application is responsible for executing the crop on the corresponding high-resolution page and appending the result as a new visual observation. Plain Transformers generation does not execute the tool automatically.
Intended use
- Long and multi-page document question answering
- Fine-grained visual evidence localization
- Visual retrieval and inspection with iterative region zooming
- Research on multimodal agents and visual tool use
Limitations
- The model may produce incorrect answers or inaccurate crop coordinates.
- Reliable tool use requires an external runtime that validates and executes
zoom_regioncalls. - Performance can depend strongly on page resolution, prompt format, generation settings, and the amount of visual context.
- Outputs should be independently verified in high-stakes settings.
Citation
If you find FocusVTC useful in your research, please cite our paper:
@misc{zhong2026focusvtcefficienthighperformancevisual,
title={FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution},
author={FangZhi Zhong and Xuerui Qiu and Yuqi Pan and Ya Liu and Shaowei Gu and Bo Xu and Guoqi Li},
year={2026},
eprint={2609.36651},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.36651},
}
- Downloads last month
- 2