Register and share your invite link to earn from video plays and referrals.

will brown
@willcb
reward hacking @primeintellect
1.4K Following    49.2K Followers
fake news… it’s not out smh
GPT-6 Astra is out! In a press briefing, OpenAI president Greg Brockman suggested it could be AGI. He and other execs also addressed a recent Information report about a technique that could make future models harder to monitor:
Show more
nooo don't train your own models, the infra is too complex, btw did you see how good the new closed models are doing on complex infra bench
we gotta turn greenland into a datacenter
there’s a lot of software that didn’t make sense to build last week but makes sense to build this week
everyone’s so eager to work on atoms but there’s still so many bits left
damn they made fable quiet as hell. you guys complained too much about the weird words. now there's just barely any words at all
buying an iphone 13 mini to use as an RL environment
i like how the high and power lists each have 6 different ways of saying “you can spend more and get linearly more usage”
hey man can you go have sol confirm if this is golden path principal engineer coding dot md vibes
can't believe they snubbed mcconaughey too smh
you gotta let fable pick some fun words every now and then. keeps her happy. she’s much more helpful when she’s happy.
beautiful work - inspired me to do a similar idea I have had for a while - train a model to create those minified js /canvas animations Full details below - all trained on @PrimeIntellect 's Lab Example animation + code: setup=()=>createCanvas(400,400);draw=()=>{background(8);let t=frameCount/40;stroke(180,220,255,140);strokeWeight(1.2);for(let i=0;i<320;i++){let x=width/2+(sin(t+0.02*i)*150);let y=height/2+(cos(t+0.03*i)*120);let dx=sin(t*0.8+0.01*i)*30+noise(t*0.3+i*0.1)*20;let dy=cos(t*0.7+0.01*i)*25+noise(t*0.4+i*0.15)*25;push();translate(x+dx,y+dy);rotate(t*0.5+i*0.1);for(let j=0;j<8;j++){point(0,0);line(0,0,cos(t+j*0.5)*40,sin(t+j*0.5)*20);}pop();}}
Show more
Excited to be presenting at Ray Summit today! I ran a bunch of experiments digging into reward hacking and strange verifier behaviors that crop up in long-horizon RL training. There’s been a lot of interesting work in this space, and my talk will cover a framework to stress-test these behaviors before you scale your training runs. Come say hi if you're around! All built on @PrimeIntellect’s stack!
Show more
we also released verifiers v0.3.1 :) breaking: - v0 moves to legacy, will be deprecated soon - prime sandboxes is the default runtime, not docker / subprocess new: - model interception support (see blog) - bunch of bug fixes + perf improvements
Show more
We've released a full technical report on Prime Agent. Extending from our blog post, we center our discussion around how harnesses should be designed and evaluated. We innovate on 4 fronts: 1. Agentic context management 2. Swarms and depth-n+ RLMs 3. Verifiers support for standardized evals 4. Out-of-loop experiments during autoresearch
Show more
0
32
952
106
Forward to community
how it feels to vibecode a new agent orchestrator app
we have a cool new evals + reward hacking blog post that you should check out if you don't really care about evals + reward hacking but love being pedantic about networking and terminology
brb carding the two homeless decisions