Today we’re launching AutoEval: a new evaluation methodology that ranks models using reward models based on millions of real Arena user preferences.
Highlights:
- High-quality evaluation signals calibrated on real preference data
- Strong alignment with live human evaluations
- Evaluations that are several orders of magnitude faster (hours instead of days)
- Support for Text, Vision, Image, and Code Arena
AutoEval enables us to evaluate newly launched models much faster and share results with the community sooner. We’ll now show AutoEval estimated scores for new models directly on the leaderboard.
More details in the thread. 🧵