MedAD-R1
A compact model for interpretable medical anomaly detection with consistency-reinforced reasoning.
Paper · Code · Dataset · Model
Access: Requests for the model weights are reviewed manually while the paper is under review. We plan open release after acceptance, subject to applicable permissions. The linked arXiv page may currently show an earlier manuscript version. The dataset page is live, but validated files are not yet available.
Model overview
MedAD-R1 answers multiple-choice questions about medical images and generates a structured reasoning trace alongside its final answer. Its two-stage training combines supervised Cognitive Injection with Consistency Group Relative Policy Optimization (Con-GRPO). The latter adds an Evidence-Aware Consistency Reward for visual grounding, question relevance, answer support, absence of hallucination, and absence of contradiction.
The primary model uses a Qwen3.5-0.8B backbone. For details of data construction, prompts, training, and evaluation, see the paper and project repository.
Reported performance
Results below are from the revised manuscript (percent; higher is better):
| Model / training | MedAD-38K accuracy | Reasoning–answer consistency | External accuracy |
|---|---|---|---|
| Qwen3.5-0.8B, zero-shot | 76.43 | 92.12 | 68.12 |
| Qwen3.5-0.8B, SFT only | 90.89 | 94.35 | 88.36 |
| MedAD-R1, SFT + Con-GRPO | 95.12 | 97.26 | 90.09 |
MedAD-38K accuracy is micro-averaged across applicable VQA instances. External accuracy is evaluated on five source-disjoint datasets. The paper reports the complete baseline comparison, evaluation protocol, and uncertainty estimates. The dataset release is being validated, so these manuscript statistics should not be interpreted as counts of files currently available on the Hub.
Intended use and limitations
This is a research checkpoint for evaluating medical-image VQA and reasoning consistency. It is not validated for clinical diagnosis or patient care. Performance on new institutions, acquisition settings, and populations should be assessed independently.
Project resources
| Resource | Link | Status |
|---|---|---|
| Paper | arXiv:2602.01081 | Public record; revised version may be pending. |
| Code | GitHub: zhtstar/MedAD-R1 | Implementation in preparation. |
| Dataset | Hugging Face: zhtstar/MedAD-38K | Information page only; validated files forthcoming with manual access review. |
| Model | Hugging Face: zhtstar/MedAD-R1 | Manual access review. |
Citation
Please cite the arXiv paper. Citation details will be updated when the revised version is public.
- Downloads last month
- 16
