Register and share your invite link to earn from video plays and referrals.

Greg Kamradt
@GregKamradt
President @arcprize, builder/engineer
1K Following    49.8K Followers
While scores at the top of ARC-AGI v1/2 are topping out at 95% there is still a lot the benchmark tell us about efficiency Not only does GPT-5.6 Sol get 92.5% on ARC-AGI-2 (sota) but it does so at 1 OOM less cost than GPT-5.5 Pro(!) GPT-5.5 Pro came out 3 months ago...
Show more
Great night hanging with the Claude Code folks. 🙌✌️ @trq212 @The_Whole_Daisy
> Last year, I was constrained by tokens. I fixed that by joining OpenAI > Then I was constrained by CPU, now I feel my constraint is actually *attention* .@steipete came to agents/pizza/wine to demo his agentic engineering workflow
Show more
oh, and if you have a feature request, you tell your agent to put it here
as the energy requirement to build new apps drop I find myself creating tools just for me (I used to create them for other people) the most recent one is readside/dev I liked elevenlabs reader, but I wanted to have a google-doc style comment bar I could chat with claude code - specifically for technical articles that I didn't understand Some features I've added: * Highlight and chat with claude code about a specific piece * Save your highlights * Listen to your articles (I do this while I'm in the car) * Use an LLM to improve the readability of the article (remove the cruft that comes from normal scraping) * Use an LLM to increase the text which gets sent to the TTS model (my main gripe with elevenlab reader was it would read the citations) * Save your articles, mark as read * Add "readside .dev/" to the beginning of any article URL and auto import it It's been fun to prompt new features right when I need them Open to feedback if you have it
Show more
Genuine problem for me. I struggle to keep track of all the work I’m doing with agents. Hopefully this is a UI/UX problem!
> Last year, I was constrained by tokens. I fixed that by joining OpenAI > Then I was constrained by CPU, now I feel my constraint is actually *attention* .@steipete came to agents/pizza/wine to demo his agentic engineering workflow
Show more
> At @tufalabs, we're specifically interested in multi-turn, long-context interactive environments and *ARC AGI-3* is exactly the type of problem we wanna take on. > It reduces some of the complexities around safety and setting up environments and focuses on the core problems...so we're super excited
Show more
ARC-AGI-3 is built different, it has dumbfounded almost all regular attempts so far because it's so much harder than anything that came before. It has no rules, it's agentic and has no explicit goals, they need to be discovered. @tufalabs won the first milestone of @arcprize > There is no language built into the benchmark, but these guys "put the language back in", because in their view - it's the best way to climb up the notional "abstraction mountain" and effectively use many of the abstractions which have evolved over millions of years of language evolution. > They built a novel harness "The Duck" around a 27B open weights model (Qwen 3.6) to solve extremely challenging and novel reasoning problems that require abstraction. > This is the launch video of their winning agentic harness, "The Duck". We have also released an exclusive interview with them on MLST, just dropped. > The million dollar question is: what will @fchollet think about how they've done it, and is this a step towards AGI?
Show more
We have an internal slack bot showing leaderboards for ARC Prize 2026 Lots of new leaderboard positions after the 1st place templates were open sourced yesterday
@GregKamradt Evaling harness components is super important for agent engineering. ARC-AGI is a natural benchmark for verifiably testing this capability
.@tufalabs just open sourced their 1st place notebook 👀
We're awarding $37.5K in prizes on June 30th to the top open source ARC-AGI-3 solutions Right now only one team, @tufalabs, sits above the template scores Which means the leaderboard is wide open for prizes Come build and agent and win money
Show more
Very cool to see @sethkarten's work on continual harnesses He used ARC-AGI-3 to study two questions: 1. Can Continual Harness discover hidden rules in games designed to be unknown at test time? 2. Which part of Continual Harness contributes most to its long-horizon progress? They found two things that made CH outperform baselines: 1. Reusable skills turn discovered mechanics into efficient execution routines 2. Reset-free refinements that improve the harness's world model as trajectories grow longer
Show more
@georgepickett is one of the best for this chained workflow style of shipping; he's been chaining skills and execplans to make /goal before goal was a thing. so this is worth a watch
First talk coming out later today with @georgepickett /goal and chill
I want to throw a meet up that is agentic engineering fancy pizza wine small ish, <30 people, high quality recording, sf Who’d be in?
We're awarding $37.5K in prizes on June 30th to the top open source ARC-AGI-3 solutions Right now only one team, @tufalabs, sits above the template scores Which means the leaderboard is wide open for prizes Come build and agent and win money
Show more
I want to throw a meet up that is agentic engineering fancy pizza wine small ish, <30 people, high quality recording, sf Who’d be in?