Introducing WeirdML v3, a fully agentic benchmark featuring 11 complex hand-made tasks.
Models must explore and understand unfamiliar data, develop ML and data analysis pipelines and produce results despite limited data, unspecified goals and/or very limited feedback.
1/8