ResNet-50 (dual head) β PitVQA phase and step recognition
Supervised dual-head ResNet-50 baseline for joint surgical-phase and surgical-step recognition on endoscopic pituitary surgery frames from PitVQA.
Trained as a baseline for the SDSC Γ Chicago Booth surgical video understanding leaderboard (Clinical context tab).
Prompt example
This closed-set example mirrors the leaderboard format, not a text-input API for this checkpoint.
[surgical frame]
What is the current surgical phase and surgical step in this endoscopic pituitary frame?
Choose one phase and one step.
Phase (choose one)
- closure
- nasal sphenoid
- sellar
Step (choose one)
- anterior sphenoidotomy
- debris clearance
- dural sealant
- durotomy
- fat graft placement
- gasket seal construct
- haemostasis
- nasal corridor creation
- nasal packing
- sellotomy
- septum displacement
- sphenoid sinus clearance
- synthetic graft placement
- tumour excision
Model
torchvision.models.resnet50backbone initialized fromIMAGENET1K_V2weights- Two classification heads on the pooled features: 3-way phase head and 14-way step head (linear, dropout 0.5)
- Cross-entropy per head, argmax decoding; batch size 64, lr 1e-4, weight decay 1e-4, 4 epochs, seed 42
- Full training code (including the model class needed to load the checkpoint) in
s68_pitvqa_supervised.py
Evaluation
Full 24,767-frame video-level validation split; exact match requires both phase and step to be correct (95% bootstrap CI):
| Metric | Value |
|---|---|
| Exact match (phase AND step correct) | 66.4% (65.9β67.0) |
| Micro-averaged F1 over both slots | 77.1% (76.6β77.5) |
This supervised baseline leads the Clinical context leaderboard as of Aug 2026; see the leaderboard for VLM comparisons.
Usage
Instantiate the dual-head module defined in s68_pitvqa_supervised.py, then:
import torch
state = torch.load("best_model.pt", map_location="cpu")
model.load_state_dict(state)
model.eval()
# phase = phase_logits.argmax(); step = step_logits.argmax()
Class vocabularies and per-class scores are in class_metrics.csv.
References
- Skobelev, K., Fithian, E., Baranovski, Y., et al. A Comparative Study in Surgical AI: Potential and Limitations of Data, Compute, and Scaling. arXiv:2603.27341, 2026.
- Dataset: He, R., Xu, M., Das, A., et al. PitVQA: Image-grounded Text Embedding LLM for Visual Question Answering in Pituitary Surgery. MICCAI 2024.
Limitations
Research baseline only. Not a medical device. Trained on a single center's endoscopic pituitary videos; expect degraded performance elsewhere.