Skip to main content
MEDAL

MEDAL Benchmark

Building a benchmark around crowdsourced medical imaging tasks selected for their potential to meaningfully improve clinical practice.

A benchmark is a standardized test for AI models: It defines a task, provides the data, and specifies how researchers measure and compare model performance. Meaningful benchmarks require more than high-quality data. Task choice, data diversity and provenance, validation design, and evaluation metrics all shape what we can learn about model capabilities.

Public benchmarks increasingly face risks of data leakage, dataset contamination, and repeated optimization on the same tasks, which can make it difficult to assess the capabilities of current models reliably (e.g., [Chen et al., 2024; Song et al., 2025; Xu et al., 2026]). Furthermore, data provenance issues have recently led to the retraction of multiple papers [Gibson et al., 2026] - Retraction notes: link 1, link 2, link 3.

In medical imaging, even a technically rigorous benchmark may have limited value if it evaluates tasks that do not reflect important clinical needs. Many biomedical imaging benchmarks focus on narrow, repetitive, or highly specific tasks and datasets. Evidence suggests that the availability of datasets and benchmarks can strongly shape which problems attract research attention [Varoquaux & Cheplygina, 2022]. In other words, data convenience rather than clinical need drives research.

MEDAL introduces a paradigm-shifting, problem-first approach to benchmark design. The clinical community will identify questions and unmet needs where AI-based interventions could have the greatest potential impact on clinical practice, and contribute high-quality clinical data. MEDAL will prioritize these questions using explicit criteria that capture their potential to improve patient outcomes and the experience of care, reduce healthcare costs, and alleviate pressure on the healthcare workforce.

A funding of €1 million for data contributors from Carl-Zeiss-Stiftung will provide a strong incentive for clinical communities and institutions worldwide to contribute diverse, high-quality data.

MEDAL builds on an international network spanning clinical medicine, medical imaging AI, and benchmarking research, with strong links to relevant scientific societies and community initiatives. Insights from COMPASS will help identify clinically important questions, unmet needs, and emerging topics, while DYNAMO, Adversarial Anatomy Bench, and related work will contribute insights into current model capabilities and robust evaluation methodology. The resulting MEDAL Benchmark will bring these clinical and technical perspectives together, anchoring AI research in clinically relevant problems while providing the AI community with rigorous evaluation tasks in areas where improved capabilities could make a meaningful difference in clinical practice.