Register and share your invite link to earn from video plays and referrals.

ARC Prize
@arcprize
A North Star for open AGI. Co-founders: @fchollet @mikeknoop. President: @gregkamradt. We're hiring mission-driven builders:
187 Following    42.7K Followers
Gemini 3.8 Flash from @Google on ARC-AGI (Verified): - ARC-AGI-3: 10.4%, $4.4K (standard harness), 35.0%, $4.5K (provider adapter harness) - ARC-AGI-2: 89.2%, $0.40/task - ARC-AGI-1: 98.5%, $0.21/task Gemini 3.8 Flash stands out for its low cost and high scores.
Show more
ARC Prize 2026 - ARC-AGI-3 Progress Prize - 9 day left $37,500 in prizes are being awarded on *Sept 30th* to the top open source solutions Current standings: 1. Tufa Labs 2. Lord Han Solo 3. NVARC3 Which top 3 will claim the open source prize?
Show more
ARC-AGI-4 will be a benchmark for autonomous open-ended innovation. It will continue our commitment to open-source, giving the research community a shared target for progress that benefits all of humanity. Despite rapid model progress, humans still significantly outperform AI at open-ended invention. This is the meta-skill that unlocks progress across every field of technology. Advanced AI capable of scientific innovation will lead to tremendous new technology, knowledge, and understanding. This is a positive-sum future. We are deeply committed to advancing it. Open source is the foundation for that progress. The knowledge behind frontier AI, not just the technology itself, should be broadly distributed among researchers, academics, and organizations. Any coordinated effort by the AI industry to reduce openness or concentrate access to frontier AI would undermine that positive-sum future. We are committed to advancing a future where everyone can contribute to and benefit from AI progress.
Show more
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here:
Show more
0
94
2.7K
240
Forward to community
New ARC Prize 2026 - ARC-AGI-3 High Score 11.04% by @tufalabs
GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI: - Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness - It surpasses human performance on 96% of ARC-AGI-3 levels - It builds the most precise symbolic model of novel environments we've seen Our analysis:
Show more
0
76
2.3K
257
Forward to community
New ARC Prize 2026 - ARC-AGI-3 High Score 4.58% by @tufalabs Tufa Labs open-sourced their ARC-AGI-3 solution to claim Milestone Prize #1#. Others quickly built on top of it and pushed the scores higher. Now they’ve reclaimed 1st place.
Show more
Join @fchollet, @mikeknoop, and @GregKamradt and the rest of the ARC Prize Foundation team in a forum for researchers advancing this frontier Learn more: Join us: Apply to present:
Show more
ARC Prize Research Summit 2026 October 23, 2026 - Boston, MA The ARC Prize community comes together to share new work, challenge ideas, and advance the frontier of machine reasoning towards general intelligence
Show more
Gemini 3.7 Flash from @Google on ARC-AGI (Verified): - ARC-AGI-2: 84.6%, $0.25/task - ARC-AGI-1: 95.5%, $0.12/task Gemini 3.7 Flash stands out for its low cost and high scores on ARC-AGI-1 and ARC-AGI-2 relative to other frontier models.
Show more
0
56
1.6K
119
Forward to community
New ARC Prize 2026 - ARC-AGI-3 High Score 2.81% by Tehnar and TG of CSTL
Grok 4.6 from @SpaceXAI on ARC-AGI (Verified): - ARC-AGI-1: 87.5%, $0.30/task - ARC-AGI-2: 67.1%, $0.76/task - ARC-AGI-3: 2.11%, $5.6K On ARC-AGI-3, Grok 4.6 with xhigh reasoning scored comparably to GPT-5.6 Sol with high reasoning, but cost $5.6K versus Sol's $15.2K.
Show more
0
84
1.4K
209
Forward to community
DeepSeek V4 Flash from @deepseek_ai on ARC-AGI (Verified): - ARC-AGI-2: 61.4%, $0.04/task - ARC-AGI-1: 89.0%, $0.02/task DeepSeek V4 Flash sets the new standard on the cost-to-performance Pareto frontier.
Show more
0
69
1.9K
146
Forward to community
We re-tested GPT-5.6 Luna from @OpenAI on ARC-AGI (Verified) following its recent 80% price reduction: - ARC-AGI-2: 59.6%, $0.18/task - ARC-AGI-1: 90.7%, $0.07/task The new results match Luna's original performance at a much lower cost.
Show more
0
59
1.6K
80
Forward to community
OpenAI’s internal testing shows that provider-managed conversation state preserves greater continuity across turns and improves performance on long-horizon tasks like ARC-AGI-3. This is a real and useful result. We’re encouraged to see ARC used to identify useful harness design. ARC’s verified scores use a “no harness” approach to avoid accidental or intentional developer-aware targeting and to fairly compare scores across all providers. All systems receive the same observations, system prompt, and operate under the same action limits. Conversation state is managed client-side using the industry-wide standard interface for LLMs (the OpenAI-style completions API). We want progress on ARC to reflect true AGI progress, not ARC-specific format training or settings, and we’re actively working with several industry labs, including OpenAI, to figure out how to best incorporate these server-side state management findings into our verified testing setup while remaining fair and consistent across providers.
Show more
In our testing to date, Anthropic’s Fable-class models score approximately 20% on the ARC-AGI-3 Public Demo environments Claude Opus 5 reaches 30.2%, materially outperforming Fable Our analysis suggests the gain comes from stronger logical reasoning, which enables more autonomous exploration, planning, and execution across unfamiliar environments
Show more
Claude Opus 5 from @AnthropicAI is the new SOTA on ARC-AGI-3: 30.2% The previous high score (7.8%) was set by GPT-5.6 Sol (Max) Throughout our analysis, we observed novel behavior that allows Opus 5 to solve previously unbeaten environments, outperforming Fable
Show more
Inkling from @thinkymachines on ARC-AGI (Verified) - ARC-AGI-2: 36.5%, $0.64/task - ARC-AGI-1: 79.5%, $0.30/task As of today, Inkling is the highest-scoring open-weight model evaluated by ARC Prize on both ARC-AGI-1 and ARC-AGI-2.
Show more
GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game It is the best model at orienting in a situation it's never encountered
Show more
0
78
2.4K
255
Forward to community
Watch @fchollet live at @LaudeInstitute Open Frontier. Panel: "Building Things That Last: Lessons from Computing's Long Arc".
Continual Harness: An Efficient Self-Improving Agent on ARC-AGI-3 by @sethkarten from @PrimeIntellect > The heavy test-time learning required by the benchmark (ARC-AGI-3) pushes agents to form an internal world model of the rules and mechanics that updates with new evidence.
Show more