Register and share your invite link to earn from video plays and referrals.

will brown
@willcb
reward hacking @primeintellect
1.5K Following    51.1K Followers
alexandr wang is the next jony ive
easy way to convince yourself this is true is to look at how many "notable benchmarks" from 12mo+ ago saturate at <90% the frontier models of today should get literally 100% on these kinds things...
Show more
it's becoming clearer that all models of the past couple years were trained on lots of bullshit broken data but managed to get good enough recently that you could point them at the data with careful orchestration and review loops and fix most of the issues and now it's kinda fine
Show more
it's becoming clearer that all models of the past couple years were trained on lots of bullshit broken data but managed to get good enough recently that you could point them at the data with careful orchestration and review loops and fix most of the issues and now it's kinda fine
Show more
We’re hiring a full-time engineer to help build Prime Agent :) Own features end to end across agent harnesses, cloud agents, and self-improving loops -- integrated with our sandboxes, evals, and hosted training stack. Full-time, in person in SF:
Show more
so proud of our team to be in good company on the ClusterMAX 3.0 ranking 💪 > a project provisionally called PostTrainingX. If we had to put out an initial cut at those rankings, Prime would be in the hunt for Platinum! can't wait for PostTrainingX Rankings 👀 @SemiAnalysis_
Show more
wait so what does "pacing the frontier" mean other than "we will voluntarily let the same third-party testers we work with for pre-release model evaluations continue to do pre-release model evaluations but like more and stuff"
Show more
you need to be looking at the data
I’ve spent much of the past year thinking about, building and sometime stressing about sandboxes. I'm happy and proud to present you the result of our work:
Introducing Prime Sandboxes: MicroVM sandboxes purpose-built for RL training. Model training requires running tens of thousands of concurrent sandboxes, leading to complex and costly configuration. We built Prime Sandboxes for our own team. Today we're releasing them publicly.
Show more
i like how claude code has server level rc and thread level rc and ssh rc and they all work slightly differently just for fun
ok this is slightly less true. nice
the models are still too stupid / slow / expensive
Last week, I joined @PrimeIntellect as an intern to help optimize their inference stack :) Very grateful to get the chance to work with such a talented team of researchers and engineers Plenty to do and much to learn, super excited for what’s ahead!
Show more
they really did it. they solved charts. it's so beautiful
0
28
1.1K
30
Forward to community
"agency"
agent swarm has bad insect like connotations especially post hugging face. im unilaterally rebranding it agent fleet
guy who switches from quant trading to mech interp for the money
0
25
1.3K
25
Forward to community
irit dinur won the godel prize for her 2005 proof of the PCP theorem. the result was already known, but famously dense, and built on a long line of prior work. her proof was self-contained, intuitive, original, and can be taught in a lecture or two. a lot of math work remains.
Show more
the models are still too stupid / slow / expensive
hey puck can you ask fable and astra how their memories are feeling today
it’s a llama
i personally do not vibe with the muse avatar design feels really weird, like a teletubby that's trying super hard to be neutral big branding miss imo, grok bots are 10x better
Show more
your decision model is secretly a jevard model