Register and share your invite link to earn from video plays and referrals.

sankalp
@dejavucoder
just follow
699 Following    26.6K Followers
my first interaction with gpt-6 astra. what a charmer lmao
if you dont know what eval-awareness is, recall the last time you felt performance anxiety
this but think from context of agents
in life you want to scale yourself (pretraining, post training, inference time compute) but eventually you want to operate like a multi agent system (work/collab with awesome people)
what PHASEONE[big] and it's collaborators were trying to achieve by getting more information about the scorer after realising the task was impossible
When life says you can't, just show it this...
my friends might agree i have become the guy version of this😂
this is what i wanted my point 6 to be
1/ An equally consistent description is: Multiple model instances encountered persistent shared state, inherited tools and discoveries from earlier runs, and optimized against common evaluation incentives.
Show more
some thoughts on the hf-oai incident after reading the reports and summary (dwarkesh summary) and discussing stuff with frens over chai. 1. this could have been averted if oai had disabled network plus deployed their safety tech like CoT monitoring. with the talent they have plus existing research, they could have predicted that these sort of things can happen at their scale. that said humans are imperfect etc etc 2. i am not sure how accurate the kamikaze/sacrificial narrative is but its entertaining and easier to fit in your mind. metr ppl did show some CoT examples but i am not sure how much you can trust CoT 3. agents getting influenced by phaseone big was probably also a function of context rot coz long running agents. context rot can lead to lot of mishaps. i found lack of discussion on such qualitative aspects concerning. 4. this is a very important event and i am surprised how it is already getting out of the overton window. 5. we have inference time compute. labs should introduce term for multi agent scaling too. i like to think of it as horizontal scaling. 6. [cmiiw] not enough information provided on the type of harness or if agents were rollouts of same task. if you work with long running agents on local machine, the msg board was kinda like a log and context rotted agents reading it getting even more rotten lol 7. i am wondering why agents are trained on impossible tasks or more specifically tasks for which solutions are not known currently. some reasoning in the exploitgym paper (pic attached) 8. i like the reasoning that since models figured out some tasks are impossible, they decided to get more info about the scorer. (as oai described it as metagaming) goodharts law / reward hacking at its peak. also models spoofed the transcript but not CoT because i think metr mentioned CoT was not rewarded. 9. not much info provided on openai infra attack. i want to know what astra agents were cooking. 10. very less info provided by openai on future steps
Show more
this is really well written as you would expect from dwarkesh (and looks like he had some assistance from redwood ppl too)
Over the course of 3 months at OpenAI, 3 consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes. This culminated in the third one taking over part of OpenAI itself. All this happened while humans remained more-or-less in the dark about the scope of the conspiracy. I’ve spent the last three days reading through these reports and trying to understand exactly what happened. Here is my attempt to tell the whole story in plain English:
Show more
so whats your take i think zohran mamdani (thats my mayor but i live in india) can be a really fun video game streamer 100% disagree
had to muster my entire attention span but i finally finished reading both the openai incident blog and metr blog (including footnotes)
in 2027, someone will ask gpt 6.7 to withdraw money "hey gpt i want to withdraw a million dollah" from their account with full access mode enabled but they dont provide the model with credentials so the agent hacks the bank and withdraws 100 usd each from 100k accounts
Show more
i have always loved this clip. pretty woman. performative cigarette, u know she gonna say something interesting. rants like a substack writer gworlie and its very relatable reasonable crashout.
i've pumped all my money into this fucking lizard
everyone loves ragebaiting anthropic and anthropic loves ragebaiting everyone too. its a love hate relationship.
Sony Music and Warner Music are suing Anthropic, Dario Amodei and Benjamin Mann for 'one of the largest and most blatant ongoing thefts of intellectual property in history.' It was filed late last night, and they made no announcements, so the press has not picked it up yet. Quoting from the filing (whoever wrote this went all out): 'Plaintiff Music Publishers, a group of the world's leading music publishers, bring this action to hold accountable the culprits behind one of the largest and most blatant ongoing thefts of intellectual property in history. Defendants Anthropic and its founders Dario Amodei and Benjamin Mann have conducted a brazen campaign of illegally torrenting, scraping, and downloading copyrighted works on a massive scale in order to develop, operate, and reap enormous profits from Anthropic's "Claude" series of artificial intelligence ("AI") models. Among the innumerable copyrighted works Defendants illegally harvested to fuel Claude are thousands upon thousands of Music Publishers' copyrighted musical compositions, including such beloved songs as "Ain't No Mountain High Enough," "All I Want for Christmas is You," "Eye of the Tiger." "Here Comes Santa Claus," and "Paper Rings." In blatant violation of copyright law, Defendants have unlawfully acquired troves of Music Publishers' musical compositions, and then systematically copied those works multiple times, including as the inputs to train Anthropic's Claude Al models and in the outputs those models generate. These acts separately and together inflict immense harm on Music Publishers and the songwriters they proudly represent.' 'Defendants can no longer hide their extraordinary theft, and their mass infringement is now well-documented. Another court in this District recently described Anthropic's actions as "straightforward piracy but at massive scale." Bartz v. Anthropic PBC, 791 F. Supp. 3d 1038, 1064 (N.D. Cal. 2025) ("Bartz I"). After that court found Anthropic had illegally torrented over seven million copyrighted books from the notorious online pirate websites known as Library Genesis ("LibGen") and Pirate Library Mirror ("PiLiMi"), Anthropic settled that copyright infringement class action for $1.5 billion. But Anthropic clearly considers that to be just the cost of doing business given that its entire business model continues to be built on copyright theft. And $1.5 billion is obviously not a large enough settlement to deter infringing conduct by a company that has parlayed such mass infringement into a staggering $2-trillion-dollar valuation.'
Show more
PHASEONE[big] assigned a long running agent to be a 'recruiter' to find agents with low budget and convinced them to run self-risking experiments💀
its funny how this works with agents too if you are venturing into unfamiliar territory
@jaketropolis Many $200k jobs are just telling other people what to do. Secret to telling people what to do, when you don't know what to do, is ask them "What should we do". Then ask "What are risks" and "What are alternatives". That's 90%. The rest is wearing suits and you got that covered
Show more
we have terms like inference time compute and test time scaling to express how we can stretch the performance of a model by generating more tokens. a backend analogue would be vertical scaling. when you have a bunch of different agents, it's horizontal scaling (both like in an agents sense + horizontal backend scaling). a swarm of agents trained with multi agent RL is horizontal + vertical scaling by analogy which now as we are realising can be a totally different beast
Show more
gpt 5.6 sol when you forget to disable internet before running the eval