Together with
@ucsc, we introduce AutoMedBench, an open benchmark for evaluating medical AI agents across the full workflow.
An agent can produce a final result even when intermediate steps fail. AutoMedBench examines planning, setup, validation, inference, and submission to help researchers identify those failures and compare systems.
6,300+ runs. 96 task combinations. 7 tracks.
See how agents compare on the #
MICCAI2026# benchmark 👉