Instructions to use lucent517/v2v2r_cbct_vqa with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use lucent517/v2v2r_cbct_vqa with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="lucent517/v2v2r_cbct_vqa") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("lucent517/v2v2r_cbct_vqa") model = AutoModelForMultimodalLM.from_pretrained("lucent517/v2v2r_cbct_vqa", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use lucent517/v2v2r_cbct_vqa with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "lucent517/v2v2r_cbct_vqa" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lucent517/v2v2r_cbct_vqa", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/lucent517/v2v2r_cbct_vqa
- SGLang
How to use lucent517/v2v2r_cbct_vqa with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "lucent517/v2v2r_cbct_vqa" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lucent517/v2v2r_cbct_vqa", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "lucent517/v2v2r_cbct_vqa" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lucent517/v2v2r_cbct_vqa", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use lucent517/v2v2r_cbct_vqa with Docker Model Runner:
docker model run hf.co/lucent517/v2v2r_cbct_vqa
Qwen3.5-9B-AWQ-dental-cbct-sft
Qwen3.5-9B fine-tuned to read rendered dental CBCT images and answer a fixed clinical schema as JSON. Built for the ODIN2026 / ToothFairy4 challenge, where a CBCT volume and its segmentation mask become a narrative radiology report.
⚠️ Research model. Not a medical device and not validated for clinical use. It has been measured only on the ToothFairy4 data, against reference reports, by an automatic judge. Do not use it to inform patient care.
What it does
The model is one stage of a pipeline, not a standalone assistant. Upstream, a
CBCT volume and its mask are rendered into a fixed set of images — a panoramic
reconstruction, 3D surface views, maxillary sinus crops, and one close-up
composite per present tooth. The model is then asked ~38 questions per case, each
one sent with the images that can actually settle it and a JSON Schema as
OpenAI-standard response_format. Downstream, deterministic rules reconcile the
answers across image sources and a string template writes the report.
So it expects those renders and those prompts. It is not a general-purpose chat model for dental radiographs, and it will not give useful free-text answers about an arbitrary CBCT slice or photograph.
Results
Official ToothFairy4 ranking on the 40-case validation split, scored end-to-end
(renders → VQA → rule-based postprocess → templated report) against the reference
reports, with a local Qwen3-14B judge for RadFact:
| metric | score |
|---|---|
| Final score | 0.4658 |
| Clinical score (RadFact logical F1) | 0.5139 |
| ├ logical precision | 0.5306 |
| └ logical recall | 0.4982 |
| Captioning (BLEU-4 + METEOR) | 0.2737 |
These are pipeline numbers. The checkpoint's own contribution is the structured extraction; the rendering and the report writer are fixed and deterministic, and changing either moves the score without retraining.
Training
LoRA on a frozen base, trained against bf16 Qwen3.5-9B and merged tensor-wise
into the AWQ checkpoint that is served. The 214 targeted tensors are bit-identical
between the two bases, so the merge is exact rather than an approximation;
provenance for this checkpoint is in merge_meta.json.
| Method | LoRA, rank 16, α 32, dropout 0.05, bias none |
| Placement | 214 modules — 108 in the vision tower (27 blocks × 4), 2 in the vision–language merger, 32 in the full-attention layers, 72 in the GatedDeltaNet layers |
| Trainable | 22,928,896 (0.24 % of the 9.41 B base) |
| Data | 9,520 VQA calls over 528 CBCT cases; 30 cases held out for eval loss |
| Schedule | 2 epochs, LR 1e-4 cosine with 3 % warm-up, effective batch 16, bf16, gradient checkpointing |
| Loss | Token cross-entropy on the values of supervised fields only — not field names, braces or separators |
| Hardware | 1 × A100 40 GB, ~12.5 h |
Targets are built from the mask and the reference reports, so supervision is on the clinical decision each field asks for rather than on report prose.
Serving
Runs under vLLM on a single 24 GiB card; weights are 11.2 GiB.
vllm serve lucent517/Qwen3.5-9B-AWQ-dental-cbct-sft \
--max-model-len 32768 \
--limit-mm-per-prompt '{"image": 3}' \
--gpu-memory-utilization 0.90
Do not shorten --max-model-len: the context has to hold a ~12.5 k-character
prompt and up to 8,192 tokens of JSON reply. At 8192 the larger calls fail with
a 400 and the pipeline quietly falls back to defaults.
Questions are sent one per request with the schema as response_format; greedy
decoding (temperature=0) is recommended for reproducibility.
Limitations
- Tied to one pipeline's renders, captions and prompt wording — different images or a different schema are out of distribution.
- Trained on 528 cases from a single challenge dataset; no external validation.
- Answers are schema-constrained JSON, English only.
- Recall is the weaker half of the clinical score: it misses findings more often than it invents them.
Base model
Qwen/Qwen3.5-9B, AWQ 4-bit quantized.
The 4-bit kernels need compute capability 8.0+ (Ampere or newer); a T4 will not
run it. Apache-2.0, inherited from the base.
- Downloads last month
- -