Register and share your invite link to earn from video plays and referrals.

Spencer Mateega
@spencermateega
1.6K Following    3.4K Followers
Every safety taxonomy today is a list of scenarios that already went wrong. Coverage lags the incident that created the category. Getting ahead of safety incidents takes a stated view on how a model should behave, not a long list of failures. Many labs today have a list and not a view. Cyber is the current example. Code and cyber safety matter now, two years after RL made agents good at offensive coding.
Show more
alright who’s creating an alignment data factory DMs open
People still underestimate how large the data market will become. Better models do not reduce demand for data. They increase the number of useful things models can learn. The frontier keeps moving, so the data has to keep getting harder.
Show more
The ability to figure out what we don’t know yet.
If knowledge becomes abundant and accessible to almost everyone, what becomes valuable next?
Enterprise data + training will eventually be one problem, not two separate services.
Interesting to see multiple of the biggest-name VC firms reverse their entire thesis on the data industry. Data being a trillion+ dollar infrastructure market should have never been a variant view.
last nights spread for @AfterQuery‘s 3rd poker night now hosting these biweekly♠️
If you're free tomorrow night, @AfterQuery is hosting a researcher poker night tmr @ 8:30pm rsvp for deets👇
This is an important milestone for European AI sovereignty, backing an alternative to US closed-model labs. Congratulations on the Series D @MistralAI!
Today marks a major step for Mistral: we’re announcing a €3B Series D, the largest equity round ever raised by a European tech company, just three years after launch.
OpenAI is pushing the frontier again with GPT-6 Astra. 🌀 Significant performance gains don’t seem fully reflected in its AAII ranking.
GPT-6 Astra makes significant gains in the Artificial Analysis Coding Agent Index, scoring equal to Fable 5 at lower cost. In the Intelligence Index, it uses fewer tokens than GPT-5.6 Sol for similar performance, but this is outweighed by higher prices Pricing is 2.5x GPT-5.6 Sol’s current prices across the board, up from $4/$20 to $10/$50 per million input/output tokens, with the same 90% discount for cache reads and 25% premium for cache writes. We see distinct stories across our two flagship Indices. In the Artificial Analysis Coding Agent Index, GPT-6 Astra equals Fable 5 at less than half the cost, driven by significant token efficiency gains. In the Artificial Analysis Intelligence Index, GPT-6 Astra is more token efficient than its predecessor for similar performance, but this is offset by the price increase. Artificial Analysis Coding Agent Index - key takeaways: ➤ Rivals top models: In Codex, GPT-6 Astra scores 67 in the Index - approximately equal to Claude Opus 5 and Fable 5 in Claude Code, and Muse Spark 1.3 in Muse Code. Fable 5.1 in Claude Code leads the Index with a score of 70. ➤ 70% more token efficient than GPT-5.6 Sol: GPT-6 Astra sees a substantial improvement in token efficiency, using one third of the tokens compared to GPT-5.6 Sol (max) in the Codex harness, and one fifth of the tokens of Claude Opus 5 (xhigh). Various effort levels of the model occupy the Pareto frontier of token efficiency. ➤ Leads Coding Agent Index cost efficiency frontier: At max effort, GPT-6 Astra costs about the same as GPT-5.6 Sol (max) while scoring 2 points higher on the Index. Per task, the model is less than half the cost of Claude Fable 5, for the same score. Artificial Analysis Intelligence Index - key takeaways: ➤ Sits beside GPT-5.6 Sol in Intelligence: GPT-6 Astra scores equal to GPT-5.6 Sol in the Index at 61. This is 5 points lower than Claude Fable 5.1 (max with fallback). The model also trails Meta’s newly released Muse Spark 1.3 (max). ➤ ~10% fewer output tokens, offset by price increase: GPT-6 Astra defines a new Pareto frontier for Intelligence Index vs Output Tokens per Task - with a ~10% reduction in token use at max effort compared to GPT-5.6 Sol. However, due to the 2.5x increase in price, the model is 75% more expensive per task than its predecessor at max effort. ➤ Hallucinates half as much as GPT-5.6 Sol: GPT-6 Astra sees a large jump in AA-Omniscience, our knowledge and hallucination benchmark. This is driven by a significant decrease in hallucination rate from 92% to 51% at max effort. Unlike some models, this improvement does not come at the cost of accuracy - Astra increased accuracy by 4 points at the same time. ➤ ~80 point gain in AA-Briefcase Elo: GPT-6 Astra improves ~80 points in AA-Briefcase, our frontier long-horizon knowledge work evaluation. Models are tested on multi-week projects, with many linked tasks and thousands of source files. Astra sees a significant increase in both rubric scores and Analytical Quality Elo in AA-Briefcase compared to its predecessor. In the other direction, we observe a reduction in Presentation Quality Elo, where GPT-5.6 Sol (max) still leads all models. ➤ Mixed progress on other evaluations: The model sees a 6 point gain in Humanity’s Last Exam, a long-standing evaluation with emphasis on mathematics, science, and humanities. This is offset by a drop of ~80 Elo points in GDPval-AA v2 - a benchmark we adapted from OpenAI’s dataset measuring economically valuable tasks across 44 occupations. We also observe 2-3 point regressions on other evaluations across a mix of capabilities, including reductions in τ³-Banking (customer support), SciCode (Python problems in a scientific domain), and AA-LCR (long context reasoning over large documents). Congratulations @OpenAI and @sama on the launch!
Show more
Big day for everyone building with open models. Hyped to see @NVIDIA put more compute behind the @huggingface community.
Exciting day for NVIDIA and @huggingface. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. They allow every developer, startup, university, industry and country to build with, customize and benefit from AI. Thank you @ClementDelangue for coming to me. NVIDIA is going to be a great home for Hugging Face, its community and the future of open models. 🤗
Show more
Gemini 3.8 set SOTA on DeepSWE at 9am. Muse Spark 1.3 took it back at 12:30pm. Incredible work from the MSL coding team.
Muse Spark 1.3 is rolling out today with frontier performance almost too cheap to meter. This is the biggest jump we've made so far on coding and agentic work. Try it in Muse Code and our API. Next up 🍉 and Muse Spark open weights releases coming soon.
Show more
Legal and finance are two of the highest-impact domains to crack. 3.8 Flash now leads both Harvey’s Legal Agent Benchmark and Vals Finance Agent v2. Huge congrats to GDM ⚡️
Introducing Gemini 3.8 Flash, another jump in Gemini's agentic + coding capabilities, and our 3rd updated Flash model in only 6 weeks... This model has been a ton of fun to work with, excited to see what you all think!
Show more
New @nvidia DGX Spark just came in. Hyped to put it to work and eliminate inference costs on a few of my workflows.