Image-Text-to-Text
Transformers
Safetensors
English
qwen3_vl
pathology
histopathology
vision-language
medical
conversational
Instructions to use Mikezcy/ASPECT-8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Mikezcy/ASPECT-8B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Mikezcy/ASPECT-8B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Mikezcy/ASPECT-8B") model = AutoModelForMultimodalLM.from_pretrained("Mikezcy/ASPECT-8B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Mikezcy/ASPECT-8B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Mikezcy/ASPECT-8B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Mikezcy/ASPECT-8B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Mikezcy/ASPECT-8B
- SGLang
How to use Mikezcy/ASPECT-8B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Mikezcy/ASPECT-8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Mikezcy/ASPECT-8B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Mikezcy/ASPECT-8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Mikezcy/ASPECT-8B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Mikezcy/ASPECT-8B with Docker Model Runner:
docker model run hf.co/Mikezcy/ASPECT-8B
|
Download README.md from Mikezcy/ASPECT-8B: direct link, hf CLI and curl.
- Browser
- Download file 3.18 kB
-
https://huggingface.co/Mikezcy/ASPECT-8B/resolve/main/README.md
- Command line
-
hf download hf://Mikezcy/ASPECT-8B/README.md
-
curl -L -H "Authorization: Bearer $HF_TOKEN" -o README.md https://huggingface.co/Mikezcy/ASPECT-8B/resolve/main/README.md
3.18 kB
| license: cc-by-nc-sa-4.0 | |
| base_model: Qwen/Qwen3-VL-8B-Instruct | |
| pipeline_tag: image-text-to-text | |
| library_name: transformers | |
| language: | |
| - en | |
| tags: | |
| - pathology | |
| - histopathology | |
| - vision-language | |
| - medical | |
| extra_gated_prompt: >- | |
| ASPECT-8B is released for non-commercial research use under CC BY-NC-SA 4.0. It is not a medical | |
| device and must not be used for diagnosis or clinical decision-making. Access requests are reviewed | |
| manually. | |
| extra_gated_fields: | |
| Full name: text | |
| Affiliation: text | |
| Country: country | |
| Intended use: text | |
| I will use ASPECT-8B for non-commercial research only: checkbox | |
| I will not use ASPECT-8B for clinical decision-making: checkbox | |
| # ASPECT-8B | |
| ASPECT is a pathology vision-language model that reports the nucleus counts behind its answers. | |
| It is built on Qwen3-VL-8B-Instruct with 8 pathology-feature tokens and 6 cell tokens, trained with | |
| three-stage supervised fine-tuning (Perceive, Generate, Reason) and GRPO with an answer–observation | |
| consistency reward. | |
| Responses follow the format | |
| ``` | |
| <think> the patch feature of the image is <|anchor_start|>...<|anchor_end|>, and the cell composition of the image is <|anchor_start|>...<|anchor_end|>. </think> | |
| <observe> description {"<count name>": <count>, ...}</observe> | |
| <answer> reasoning | |
| FINAL: <option> </answer> | |
| ``` | |
| Paper: [https://arxiv.org/abs/2609.34277](https://arxiv.org/abs/2609.34277) | |
| Code: https://github.com/ChyaZhang/ASPECT | |
| ## Files | |
| The repository root holds the supervised model (Qwen3-VL-8B with the visual-token embeddings and the SFT | |
| LoRA merged in); `rl_adapter/` holds the LoRA adapter from reinforcement learning. Load both, as below; | |
| this is the configuration evaluated on PathoVernier. | |
| ## Usage | |
| ```python | |
| import torch | |
| from peft import PeftModel | |
| from PIL import Image | |
| from transformers import AutoModelForImageTextToText, AutoProcessor | |
| model_id = "Mikezcy/ASPECT-8B" | |
| processor = AutoProcessor.from_pretrained(model_id, max_pixels=1360 * 28 * 28) | |
| model = AutoModelForImageTextToText.from_pretrained(model_id, dtype=torch.bfloat16, device_map="cuda") | |
| model = PeftModel.from_pretrained(model, model_id, subfolder="rl_adapter") | |
| image = Image.open("example.png").convert("RGB").resize((512, 512)) | |
| question = "..." | |
| messages = [{"role": "user", "content": [{"type": "image"}, {"type": "text", "text": question}]}] | |
| text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) | |
| inputs = processor(text=[text], images=[image], return_tensors="pt").to(model.device) | |
| out = model.generate(**inputs, max_new_tokens=512, do_sample=False) | |
| print(processor.tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=False)) | |
| ``` | |
| Images are resized to 512x512. No system prompt is used; the question should list the answer options. | |
| ## Citation | |
| ```bibtex | |
| @article{zhang2026see, | |
| title = {See, Measure, and Reason: Learning Visually Grounded Reasoning in Pathology}, | |
| author = {Zhang, Chengyang and Zhang, Wenchuan and Li, Bo and Li, Mengran and Liu, Xinyu and Yang, Jiaming and Chen, Jie and Zhang, Zhang and Yi, Yuhao and Bu, Hong and Lv, Jiancheng}, | |
| journal = {arXiv preprint arXiv:2609.34277}, | |
| year = {2026} | |
| } | |
| ``` | |