See how AI models can help you with multi-day engineering workflows with Android Bench 2.0.
The updated benchmark evaluates long-horizon tasks like building apps and features from scratch, migrating cross-platform codebases to Android, and making complex architectural transitions, with continuous completion scoring that shows which tasks models perform well on.
Check out what's new 👇
顯示更多
📢 Introducing Android Bench 2.0.
We've leveled up our AI evaluation framework to handle real-world, multi-day engineering challenges—from building apps from scratch to migrating cross-platform codebases to Android.
Let’s dive into what’s new. 🧵👇🏽
顯示更多