ASPECT-8B / README.md
Mikezcy's picture
Add arXiv citation
b92bcb9 verified
|
Raw History Blame Contribute Delete
3.18 kB
---
license: cc-by-nc-sa-4.0
base_model: Qwen/Qwen3-VL-8B-Instruct
pipeline_tag: image-text-to-text
library_name: transformers
language:
- en
tags:
- pathology
- histopathology
- vision-language
- medical
extra_gated_prompt: >-
ASPECT-8B is released for non-commercial research use under CC BY-NC-SA 4.0. It is not a medical
device and must not be used for diagnosis or clinical decision-making. Access requests are reviewed
manually.
extra_gated_fields:
Full name: text
Affiliation: text
Country: country
Intended use: text
I will use ASPECT-8B for non-commercial research only: checkbox
I will not use ASPECT-8B for clinical decision-making: checkbox
---
# ASPECT-8B
ASPECT is a pathology vision-language model that reports the nucleus counts behind its answers.
It is built on Qwen3-VL-8B-Instruct with 8 pathology-feature tokens and 6 cell tokens, trained with
three-stage supervised fine-tuning (Perceive, Generate, Reason) and GRPO with an answer–observation
consistency reward.
Responses follow the format
```
<think> the patch feature of the image is <|anchor_start|>...<|anchor_end|>, and the cell composition of the image is <|anchor_start|>...<|anchor_end|>. </think>
<observe> description {"<count name>": <count>, ...}</observe>
<answer> reasoning
FINAL: <option> </answer>
```
Paper: [https://arxiv.org/abs/2609.34277](https://arxiv.org/abs/2609.34277)
Code: https://github.com/ChyaZhang/ASPECT
## Files
The repository root holds the supervised model (Qwen3-VL-8B with the visual-token embeddings and the SFT
LoRA merged in); `rl_adapter/` holds the LoRA adapter from reinforcement learning. Load both, as below;
this is the configuration evaluated on PathoVernier.
## Usage
```python
import torch
from peft import PeftModel
from PIL import Image
from transformers import AutoModelForImageTextToText, AutoProcessor
model_id = "Mikezcy/ASPECT-8B"
processor = AutoProcessor.from_pretrained(model_id, max_pixels=1360 * 28 * 28)
model = AutoModelForImageTextToText.from_pretrained(model_id, dtype=torch.bfloat16, device_map="cuda")
model = PeftModel.from_pretrained(model, model_id, subfolder="rl_adapter")
image = Image.open("example.png").convert("RGB").resize((512, 512))
question = "..."
messages = [{"role": "user", "content": [{"type": "image"}, {"type": "text", "text": question}]}]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = processor(text=[text], images=[image], return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(processor.tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=False))
```
Images are resized to 512x512. No system prompt is used; the question should list the answer options.
## Citation
```bibtex
@article{zhang2026see,
title = {See, Measure, and Reason: Learning Visually Grounded Reasoning in Pathology},
author = {Zhang, Chengyang and Zhang, Wenchuan and Li, Bo and Li, Mengran and Liu, Xinyu and Yang, Jiaming and Chen, Jie and Zhang, Zhang and Yi, Yuhao and Bu, Hong and Lv, Jiancheng},
journal = {arXiv preprint arXiv:2609.34277},
year = {2026}
}
```