DYNAMO
Prospective, contamination-resistant benchmarking of clinical multimodal AI
Recent advances in multimodal foundation models have created unprecedented opportunities for biomedical imaging. Reliable evaluation has emerged as one of the central challenges [Gu et al., 2026]. The widely used static benchmark datasets are problematic for several reasons: Public benchmarks face a growing risk of data leakage, dataset contamination, and repeated optimization (e.g., [Chen et al., 2024; Song et al., 2025; Xu et al., 2026]). These issues have motivated the emergence of dynamic benchmarks in general AI [Kiela et al., 2021; Chen et al., 2025; White et al., 2025; Jain et al., 2025] and computer vision [Shabtay et al., 2025; Zhang et al., 2025]. No comparable resource currently exists for biomedical imaging. Additionally, data provenance issues have recently led to the retraction of multiple papers [Gibson et al., 2026] - Retraction notes: link 1, link 2, link 3. Apart from data governance and reliable validation, biomedical imaging also presents unique requirements regarding clinical diversity.
The Medical Image Computing and Computer-Assisted Intervention (MICCAI) Society is the leading international community in medical image analysis, and its annual challenges cover approximately 66% of publicly organized biomedical image analysis challenges worldwide [Reinke et al., 2025]. They also cover an exceptionally broad spectrum of real-world biomedical imaging tasks, such as detecting stroke damage in brain MRI scans, identifying cancer lesions in whole-body PET/CT images, diagnosing tuberculosis from chest X-rays, assessing heart function and predicting treatment-related heart damage from ultrasound videos, generating radiology reports from head CT scans, and tracking surgical sponges, needles and other objects throughout an operation to help prevent items from being left behind.
Building on this unique community, we introduce a prospective, contamination-free, continuously evolving biomedical imaging benchmark for multimodal AI. Rather than relying on a fixed dataset, DYNAMO continuously incorporates new biomedical imaging tasks from successive MICCAI challenges. As each annual edition contributes unseen clinical datasets spanning diverse diseases, imaging modalities, institutions, and task types, the benchmark naturally evolves over time, providing a prospective and contamination-resistant framework for evaluating frontier multimodal foundation models.Participating challenge organizers contribute their unpublished test data, which are transformed into standardized vision-language evaluation tasks enabling systematic evaluation of frontier general-purpose and biomedical foundation models.
The DYNAMO benchmark enables us to characterize current model capabilities across clinical domains, identify important limitations, and establish a scalable framework for longitudinal evaluation as both datasets and models continue to evolve.