Register and share your invite link to earn from video plays and referrals.

ARC Prize
@arcprize
A North Star for open AGI. Co-founders: @fchollet @mikeknoop. President: @gregkamradt. We're hiring mission-driven builders:
187 Following    40.7K Followers
OpenAI’s internal testing shows that provider-managed conversation state preserves greater continuity across turns and improves performance on long-horizon tasks like ARC-AGI-3. This is a real and useful result. We’re encouraged to see ARC used to identify useful harness design. ARC’s verified scores use a “no harness” approach to avoid accidental or intentional developer-aware targeting and to fairly compare scores across all providers. All systems receive the same observations, system prompt, and operate under the same action limits. Conversation state is managed client-side using the industry-wide standard interface for LLMs (the OpenAI-style completions API). We want progress on ARC to reflect true AGI progress, not ARC-specific format training or settings, and we’re actively working with several industry labs, including OpenAI, to figure out how to best incorporate these server-side state management findings into our verified testing setup while remaining fair and consistent across providers.
Show more
In our testing to date, Anthropic’s Fable-class models score approximately 20% on the ARC-AGI-3 Public Demo environments Claude Opus 5 reaches 30.2%, materially outperforming Fable Our analysis suggests the gain comes from stronger logical reasoning, which enables more autonomous exploration, planning, and execution across unfamiliar environments
Show more
Claude Opus 5 from @AnthropicAI is the new SOTA on ARC-AGI-3: 30.2% The previous high score (7.8%) was set by GPT-5.6 Sol (Max) Throughout our analysis, we observed novel behavior that allows Opus 5 to solve previously unbeaten environments, outperforming Fable
Show more
0
57
2.1K
218
Forward to community
Inkling from @thinkymachines on ARC-AGI (Verified) - ARC-AGI-2: 36.5%, $0.64/task - ARC-AGI-1: 79.5%, $0.30/task As of today, Inkling is the highest-scoring open-weight model evaluated by ARC Prize on both ARC-AGI-1 and ARC-AGI-2.
Show more
GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game It is the best model at orienting in a situation it's never encountered
Show more
0
78
2.4K
255
Forward to community
Watch @fchollet live at @LaudeInstitute Open Frontier. Panel: "Building Things That Last: Lessons from Computing's Long Arc".
Continual Harness: An Efficient Self-Improving Agent on ARC-AGI-3 by @sethkarten from @PrimeIntellect > The heavy test-time learning required by the benchmark (ARC-AGI-3) pushes agents to form an internal world model of the rules and mechanics that updates with new evidence.
Show more
Submissions for the first ARC-AGI-3 milestone prize are due in 6 days Notebooks must be open sourced by 11:59pm UTC on June 30th $37.5K in prizes await the winners
ARC Prize 2026 - ARC-AGI-3 Progress Prize $37,500 in prizes are being awarded on June 30th to the top open source solutions Only one team is out performing the templates Can you be the next?
Show more
GLM-5.2 from @Zai_org on ARC-AGI (Verified) - ARC-AGI-2: 22.8%, $0.25 - ARC-AGI-1: 77.0%, $0.19 Performance is comparable with GPT-5.4 & 5.5 (Low Reasoning Effort)
ARC Prize 2026 - ARC-AGI-3 Progress Prize $37,500 in prizes are being awarded on June 30th to the top open source solutions Only one team is out performing the templates Can you be the next?
Show more
We had early access to Anthropic’s Fable 5, but did not run verified Semi-Private ARC-AGI-1/2/3 evals due to their new data-retention terms for Mythos-class models. We’re working with Anthropic to keep ARC verification data private. Scores will come once we can run them safely.
Show more
0
35
2.1K
74
Forward to community
GPT-5.5 on ARC-AGI (Verified) ARC-AGI-2: - Max: 85.0%, $1.87 - High: 83.3%, $1.45 - Med: 70.4%, $0.86 - Low: 33%, $0.35 GPT-5.5 is now state of the art on ARC-AGI-2
0
58
2.1K
221
Forward to community
Announcing ARC-AGI-3 The only unsaturated agentic intelligence benchmark in the world Humans score 100%, AI <1% This human-AI gap demonstrates we do not yet have AGI Most benchmarks test what models already know, ARC-AGI-3 tests how they learn
Show more
0
249
4.3K
578
Forward to community