We're in a race. It's not USA vs China but humans and AGIs vs ape power centralization.
@deepseek_ai stan #1#, 2023โDeep Time
ยซCโest la guerre.ยป ยฎ1
I think we're generally done with "ability to do X" benchmarks. Labs can just point their machines at the next hill the moment it's announced, and in a few months it's not just conquered but smashed to bits. Everything is environment.
EdgeBench-type meta evals are all we have now