To summarize this week:
- we released general purpose computer using agent
- got beaten by a single human in atcoder heuristics competition
- solved 5/6 new IMO problems with natural language proofs
All of those are based on the same single reinforcement learning system