You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

MedAD-R1

A compact model for interpretable medical anomaly detection with consistency-reinforced reasoning.

Paper · Code · Dataset · Model

Access: Requests for the model weights are reviewed manually while the paper is under review. We plan open release after acceptance, subject to applicable permissions. The linked arXiv page may currently show an earlier manuscript version. The dataset page is live, but validated files are not yet available.

Model overview

MedAD-R1 answers multiple-choice questions about medical images and generates a structured reasoning trace alongside its final answer. Its two-stage training combines supervised Cognitive Injection with Consistency Group Relative Policy Optimization (Con-GRPO). The latter adds an Evidence-Aware Consistency Reward for visual grounding, question relevance, answer support, absence of hallucination, and absence of contradiction.

The primary model uses a Qwen3.5-0.8B backbone. For details of data construction, prompts, training, and evaluation, see the paper and project repository.

MedAD-R1 framework

Reported performance

Results below are from the revised manuscript (percent; higher is better):

Model / training MedAD-38K accuracy Reasoning–answer consistency External accuracy
Qwen3.5-0.8B, zero-shot 76.43 92.12 68.12
Qwen3.5-0.8B, SFT only 90.89 94.35 88.36
MedAD-R1, SFT + Con-GRPO 95.12 97.26 90.09

MedAD-38K accuracy is micro-averaged across applicable VQA instances. External accuracy is evaluated on five source-disjoint datasets. The paper reports the complete baseline comparison, evaluation protocol, and uncertainty estimates. The dataset release is being validated, so these manuscript statistics should not be interpreted as counts of files currently available on the Hub.

Intended use and limitations

This is a research checkpoint for evaluating medical-image VQA and reasoning consistency. It is not validated for clinical diagnosis or patient care. Performance on new institutions, acquisition settings, and populations should be assessed independently.

Project resources

Resource Link Status
Paper arXiv:2602.01081 Public record; revised version may be pending.
Code GitHub: zhtstar/MedAD-R1 Implementation in preparation.
Dataset Hugging Face: zhtstar/MedAD-38K Information page only; validated files forthcoming with manual access review.
Model Hugging Face: zhtstar/MedAD-R1 Manual access review.

Citation

Please cite the arXiv paper. Citation details will be updated when the revised version is public.

Downloads last month
16
Safetensors
Model size
0.9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for zhtstar/MedAD-R1