You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Qwen3.5-9B-AWQ-dental-cbct-sft

Qwen3.5-9B fine-tuned to read rendered dental CBCT images and answer a fixed clinical schema as JSON. Built for the ODIN2026 / ToothFairy4 challenge, where a CBCT volume and its segmentation mask become a narrative radiology report.

⚠️ Research model. Not a medical device and not validated for clinical use. It has been measured only on the ToothFairy4 data, against reference reports, by an automatic judge. Do not use it to inform patient care.

What it does

The model is one stage of a pipeline, not a standalone assistant. Upstream, a CBCT volume and its mask are rendered into a fixed set of images — a panoramic reconstruction, 3D surface views, maxillary sinus crops, and one close-up composite per present tooth. The model is then asked ~38 questions per case, each one sent with the images that can actually settle it and a JSON Schema as OpenAI-standard response_format. Downstream, deterministic rules reconcile the answers across image sources and a string template writes the report.

So it expects those renders and those prompts. It is not a general-purpose chat model for dental radiographs, and it will not give useful free-text answers about an arbitrary CBCT slice or photograph.

Results

Official ToothFairy4 ranking on the 40-case validation split, scored end-to-end (renders → VQA → rule-based postprocess → templated report) against the reference reports, with a local Qwen3-14B judge for RadFact:

metric score
Final score 0.4658
Clinical score (RadFact logical F1) 0.5139
├ logical precision 0.5306
└ logical recall 0.4982
Captioning (BLEU-4 + METEOR) 0.2737

These are pipeline numbers. The checkpoint's own contribution is the structured extraction; the rendering and the report writer are fixed and deterministic, and changing either moves the score without retraining.

Training

LoRA on a frozen base, trained against bf16 Qwen3.5-9B and merged tensor-wise into the AWQ checkpoint that is served. The 214 targeted tensors are bit-identical between the two bases, so the merge is exact rather than an approximation; provenance for this checkpoint is in merge_meta.json.

Method LoRA, rank 16, α 32, dropout 0.05, bias none
Placement 214 modules — 108 in the vision tower (27 blocks × 4), 2 in the vision–language merger, 32 in the full-attention layers, 72 in the GatedDeltaNet layers
Trainable 22,928,896 (0.24 % of the 9.41 B base)
Data 9,520 VQA calls over 528 CBCT cases; 30 cases held out for eval loss
Schedule 2 epochs, LR 1e-4 cosine with 3 % warm-up, effective batch 16, bf16, gradient checkpointing
Loss Token cross-entropy on the values of supervised fields only — not field names, braces or separators
Hardware 1 × A100 40 GB, ~12.5 h

Targets are built from the mask and the reference reports, so supervision is on the clinical decision each field asks for rather than on report prose.

Serving

Runs under vLLM on a single 24 GiB card; weights are 11.2 GiB.

vllm serve lucent517/Qwen3.5-9B-AWQ-dental-cbct-sft \
    --max-model-len 32768 \
    --limit-mm-per-prompt '{"image": 3}' \
    --gpu-memory-utilization 0.90

Do not shorten --max-model-len: the context has to hold a ~12.5 k-character prompt and up to 8,192 tokens of JSON reply. At 8192 the larger calls fail with a 400 and the pipeline quietly falls back to defaults.

Questions are sent one per request with the schema as response_format; greedy decoding (temperature=0) is recommended for reproducibility.

Limitations

  • Tied to one pipeline's renders, captions and prompt wording — different images or a different schema are out of distribution.
  • Trained on 528 cases from a single challenge dataset; no external validation.
  • Answers are schema-constrained JSON, English only.
  • Recall is the weaker half of the clinical score: it misses findings more often than it invents them.

Base model

Qwen/Qwen3.5-9B, AWQ 4-bit quantized. The 4-bit kernels need compute capability 8.0+ (Ampere or newer); a T4 will not run it. Apache-2.0, inherited from the base.

Downloads last month
-
Safetensors
Model size
10B params
Tensor type
F32
·
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for lucent517/v2v2r_cbct_vqa

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(982)
this model