Register and share your invite link to earn from video plays and referrals.

Florian Brand
@xeophon
evals @PrimeIntellect | open models @interconnectsai
788 Following    15.5K Followers
no that bad pr wasn't me, it was a rogue mythos overtaking my account
Heres the latest on FelonyBench. OpenAI currently projected to have 300 felonies by the end of August with Anthropic only at 100. If exponential growth continues, as it always does with language models, by this time next year we're projected to have more felonies than there are observable stars in the universe (~10^24)
Show more
getting >30 pp improvements by changing to native tool calling, preserving reasoning + setting correct sampling params so many evals have those subtle mistakes which are easy to catch if you know where to look. or just use verifiers + the native harness : )
Show more
Very excited to finally share with the community. My wish was that we can lift up those that are leading the frontier of open source / open weight models. And it is fun to collab with @natolambert and @xeophon on this.
Show more
Today, we're introducing @Intelligence_ai. In 6 months, as a team of 10, we scaled from $5M to $60M ARR and 5.5M users across 190+ countries. We raised a $7.9M seed, led by @IndexVentures with participation from @conviction, @A_StarVC, and @combinator to build DesignArena, a universal interface for accessing and evaluating the world's AI capabilities. Most evaluations try to simulate the real-world. We believe the real-world is the ultimate verifier. People come to @DesignArena with a request. Models compete to fulfill the request, and users determine what works best for them. Their live user behavior evaluates the models, improves how work is routed, and helps people access the right intelligence. We've helped the world's leading frontier labs break the news on their SOTA capabilities. What's the limit? Join us and find out.
Show more
0
340
2K
116
Forward to community
Nice new Web interface from the @interconnectsai team > Artifacts Hub > Tracking the models that matter for the open AI ecosystem GG @natolambert @xeophon
We've updated the MirrorCode leaderboard with results for Claude Fable 5 and GPT-5.6 Sol. Claude Fable 5 leads with a 64% solve rate, followed by GPT-5.6 Sol at 20%.
oh would you look at this, an open model being sota 👀 sinatras cooked so hard here 👨‍🍳
Time to give agents a hard task, introducing pmpp-hard! 69 GPU kernel tasks, 11 models, 3.1k agent rollouts and 5.8B+ tokens later we have the results. Kimi K3 claims the first place with a 0.71 score. Without dealing with if your task is “frontier” or not, it just solves them
Show more
we made some of our data freely available for yall :) if you are curious about which open models are great, use the artifacts hub :) if you want to feel depressed, open the adoption dashboard :(
It's a good time to share a new @interconnectsai project we made to help make sense of the accelerating open-weight releases these days. The Artifacts Hub builds on our monthly open model roundups and daily monitoring of every model on @huggingface. The new free resources are: 1. The Artifacts Hub — a curated view of the models trending on Hugging Face, highlighting inference tokens via Open Router, model intelligence via Artificial Analysis, and our tailored adoption metrics building on top of Hugging Face’s data. 2. Our Adoption Dashboard — a living dashboard of download and derivative model numbers by geography and organization. This highlights the US-China gap and growing players in the open ecosystem. To date, our primary efforts on Interconnects have been release recaps for popular models like Kimi K3, GLM 5.2, DeepSeek R1, etc. and monthly round-ups of the open models that matter, Artifacts Log. We’re expanding on these, building on the tools and internal data we’ve collected for other projects like The ATOM Project (and report). This allows us to capture our ecosystem view of open models, develop methods for understanding adoption of giant MoE models, and everything in between. We’re sharing them freely to help the open ecosystem find its strengths and grow. The Artifacts Hub right now covers 792 models released in the last two years, across the core text-focused language models and multimodal generative models. At Interconnects we follow the data of every model on Hugging Face, analyze the core few thousand LLMs (this list is public on GitHub and regularly updated), and hand select these core few hundred for further explanation. For the most popular models, the Hub let’s you quickly see how far behind the model was in terms of frontier intelligence based on Artificial Analysis’s Intelligence Index, compare Hugging Face and Open Router adoption to similar models, glance at relative adoption metric (RAM) scores for time-size normalized downloads, or look at the VAIL similarity index of models with related generations. A snapshot for what you’d see for something like GLM-5.2 is below. Thanks to @mnshah at VAIL for encouraging us to make this and @huggingface, @OpenRouter, & @ArtificialAnlys to making such useful data openly available.
Show more
while closed models are racing to >10T params, open models are already at 4 quintillion params
Artifacts? Oh, you mean the @poolsideai (and @thinkymachines and @deepseek_ai and @Kimi_Moonshot) fan club? Last month was fun, we even got some great datasets 👀
Artifacts 23: Laguna S2.1, Inkling, & Kimi K3 show the utility of open models on the Pareto frontier. In this issue, a total of 24 models from june/july you should be aware of, from: Thinking Machines @thinkymachines (2) Tencent @TencentGlobal (2) Poolside @poolsideai (2) DeepSeek @deepseek_ai Moonshot AI @Kimi_Moonshot Meituan LongCat @Meituan_LongCat Motif Technologies Swiss AI Initiative @apertusllm AMD @AMD Upstage @upstageai InclusionAI @TheInclusionAI Moondream @moondreamai Baseten @baseten IBM @IBM Mistral AI @mistralai Google @GoogleAI Nanbeige @nanbeige Kwaipilot @KwaiAICoder InternLM @intern_lm Microsoft @Microsoft fal @fal Read the issue below.
Show more
K3 in a custom harness is pretty close to GPT Pro btw
happy to report that models made a lot of progress in reviewing evals compared to a year or even 3 months ago both on the high level as well as in assessing samples
Reaching SOTA By Simply Trying Hard Enough
The unanswered question of course is how fast things need to broadly diffuse. I am a proponent of stages access, but for how long? Would the world be a safer place if everyone has Mythos access now?
Amazing blog post I very much agree with. I also think it is more important than ever to enable frontier-level safety research for open models and we’ll share things we’ve been working on soon.
Amazing blog post I very much agree with. I also think it is more important than ever to enable frontier-level safety research for open models and we’ll share things we’ve been working on soon.
Releasing weights indiscriminately isn't safe. Neither is keeping capable models inside a few labs. We think there's a path between them. We haven't mapped all of it. Our new post covers the part we can see: how we assessed Inkling, and why access should widen in stages.
Show more
Existing models, Fable and 5.6 Sol, were also able to prove the existence of nonsofic groups (last night before the paper release)
We calculated okayish, but not the greatest results for @deepseek_ai's V4F-0731 in @sam_paech's new EQ-Bench v4 benchmark. It is an interesting question whether this is a systematic effect of the new post training of 0731: How does heavy RL affect a model's personality, could it be detrimental?
Show more