Ever spend way too long hunting across scattered sources just to find "which benchmark actually fits this task"?
Title: Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation
URL:
โ What is Benchmark Radar?
๐ก It's a search engine that crawls 37 sources daily, arXiv, Hugging Face, GitHub and more, and pulls each benchmark's paper, dataset, repository, and score history into one place, with a CLI for offline queries.
โ How much data are we talking about?
๐ก 1,283 source records from four catalogs, covering 12,916 numeric scores across 790 benchmarks. Daily collection alone logs 11,068 observations across 6,546 distinct artifacts.
โ Why can't scores just be compared directly?
๐ก Only 82 of the scored records use a verified 0-100 percentage scale, the rest use different or unverified scales, and roughly half lack a known release date, so naive comparisons would be misleading.
โ So what's the fix?
๐ก Instead of forcing comparability, it preserves source identity and citations so readers can inspect the evaluation conditions themselves.
#
LLMEval# #
Benchmarks#