Ever spend way too long hunting across scattered sources just to find "which benchmark actually fits this task"?
Title: Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation
URL:
❓ What is Benchmark Radar?
💡 It's a search engine that crawls 37 sources daily, arXiv, Hugging Face, GitHub and more, and pulls each benchmark's paper, dataset, repository, and score history into one place, with a CLI for offline queries.
❓ How much data are we talking about?
💡 1,283 source records from four catalogs, covering 12,916 numeric scores across 790 benchmarks. Daily collection alone logs 11,068 observations across 6,546 distinct artifacts.
❓ Why can't scores just be compared directly?
💡 Only 82 of the scored records use a verified 0-100 percentage scale, the rest use different or unverified scales, and roughly half lack a known release date, so naive comparisons would be misleading.
❓ So what's the fix?
💡 Instead of forcing comparability, it preserves source identity and citations so readers can inspect the evaluation conditions themselves.
#
LLMEval# #
Benchmarks#