FULL INTERVIEW:
@sophiadew sat down with
@poteto and
@roshan_s from Cursor on Grok 4.6, agents that fix their own bugs, and how an engineer's job changes.
Grok 4.6 Extra High scored 70.8% on CursorBench 3.2 at $2.81 a task. Fable 5 Max scored 70.5% at $17.32. Same result, 6x the price.
1:07 what Grok Bot actually is
2:15 the Odyssey IMAX tickets an agent booked for her
3:55 live demo: a coordinator agent running a team of bots
4:29 a marketing agent posting to LinkedIn, live on air
7:46 what Grok 4.6 changes
9:59 what changes when your agent is a colleague, not a tool
11:44 Benny, the Slack bot that fixed bugs while she slept
13:04 how much of Grok Bot was built with Grok Bot
13:29 second demo: dispatching bug fixes to a team of agents
14:36 your bots can kick off Cursor cloud agents
17:46 rewriting your codebase so agents can thrive in it
18:34 refactors that would have taken a team months
20:12 the Michelin kitchen metaphor
22:23 the iMessage support she hacked in early
23:36 the CursorBench numbers
25:17 "if you have a budget of 20 bucks to do all of your AI tasks"
@SpaceXAI @cursor_ai @bot