Instructions to use pianzhikuang/VisAlloc with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use pianzhikuang/VisAlloc with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="pianzhikuang/VisAlloc") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("pianzhikuang/VisAlloc") model = AutoModelForMultimodalLM.from_pretrained("pianzhikuang/VisAlloc", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use pianzhikuang/VisAlloc with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "pianzhikuang/VisAlloc" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pianzhikuang/VisAlloc", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/pianzhikuang/VisAlloc
- SGLang
How to use pianzhikuang/VisAlloc with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "pianzhikuang/VisAlloc" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pianzhikuang/VisAlloc", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "pianzhikuang/VisAlloc" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pianzhikuang/VisAlloc", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use pianzhikuang/VisAlloc with Docker Model Runner:
docker model run hf.co/pianzhikuang/VisAlloc
VisAlloc — LookAway 9B
This repository contains the completed step 390 checkpoint from the
LookAway-Qwen3.5-9B-b16-765d1f5 training run, based on
Qwen/Qwen3.5-9B and the
LookAway codebase.
The FSDP actor checkpoint was merged into a complete Hugging Face model.
The release contains 9,409,813,744 parameters in BF16, tokenizer, image
processor, generation configuration, and the checkpoint's chat template.
The base-model revision is c202236235762e1c871ad0ccb60c8ee5ba337b9a.
The training code was based on commit 765d1f5 with local runtime and memory
adjustments. The checkpoint completed 390 training steps and one epoch.
Usage
The evaluation environment used Transformers 5.5.0, PyTorch 2.10.0, and vLLM 0.18.0. Keep the supplied processor and chat template when loading this checkpoint.
import torch
from PIL import Image
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration
model_id = "pianzhikuang/VisAlloc"
processor = AutoProcessor.from_pretrained(model_id)
model = Qwen3_5ForConditionalGeneration.from_pretrained(
model_id, dtype=torch.bfloat16, device_map="auto"
).eval()
image = Image.open("image.jpg").convert("RGB")
messages = [{"role": "user", "content": [
{"type": "image", "image": image},
{"type": "text", "text": "Describe this image."},
]}]
inputs = processor.apply_chat_template(
messages, tokenize=True, add_generation_prompt=True,
return_dict=True, return_tensors="pt",
).to(model.device)
with torch.inference_mode():
output = model.generate(**inputs, max_new_tokens=256, do_sample=False)
print(processor.decode(
output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True
))
Evaluation
The following results were measured on this checkpoint on October 9, 2026.
Inference used full images, BF16, greedy decoding (temperature=0), up to
4,096 generated tokens, and a 32,768-token engine context. The checkpoint's
chat template was used without an enable_thinking override.
Judging used strict matching for unambiguous answers and GPT-6 Astra for the remaining semantic decisions. All 16,074 inference requests completed; 16,073 samples have valid ground-truth labels and were scored.
| Benchmark | Metric | Score (%) | Scored samples |
|---|---|---|---|
| V* | Accuracy | 83.25 | 191 |
| ZoomBench | Valid-subset accuracy | 58.29 | 844 / 845 |
| HR-4K | Accuracy | 74.625 | 800 |
| HR-8K | Accuracy | 72.25 | 800 |
| MMVP | Per-question accuracy | 78.00 | 300 |
| CV-Bench | Mean of 2D and 3D accuracy | 74.77 | 2,638 |
| MMStar | Accuracy | 62.20 | 1,500 |
| POPE | Accuracy | 89.03 | 9,000 |
- ZoomBench contains one question with only A/B choices but a ground-truth label of D. That sample remains unscored; 58.29% is 492/844, not a full-dataset score. The exception is recorded in the evaluation manifest.
- MMVP paired accuracy is 57.33% (86/150 pairs).
- POPE F1 is 88.78%. Evaluation includes all three 3,000-question subsets.
- HR-4K is exactly 597/800 = 74.625%; the CSV uses Python's two-decimal formatting.
See results.json, results.csv, and the evaluation manifest for exact scores, pinned dataset revisions, generation settings, and integrity hashes. Results depend on this prompting and judging protocol.
Release integrity
model.safetensors is the exact weight file used in the evaluation:
d7977253c2b469499e0b9f9e509d9cf60c0ae16d112b119cb2ebce072e59b0ef
The published top-level config.json dtype is corrected to bfloat16 to
match all saved tensors and the evaluated inference dtype. The tensor
weights and chat template are unchanged.
The base model is released under Apache-2.0; its license is included in LICENSE. This repository contains the model artifacts and aggregate evaluation metadata; training images and benchmark images are separate resources.
- Downloads last month
- -