Register and share your invite link to earn from video plays and referrals.

himanshu
@himanshustwts
cofounder @PhyseraAI • host @groundzero_twt • DMs open!
3.7K Following    29.1K Followers
now they might be getting billions now cuz the number of top-tier bay area investors reached out to me is insane OMFG
a group of 7 most safety-pilled researchers at Anthropic might be leaving the company to create a new AI lab and join the race to build super and safe intelligence before Anthropic does.
ideally you should never ever apply for a job, a job should apply for you.
0
145
3.4K
170
Forward to community
Two Life Sciences announcement on the same day. Coefficient Bio acquisition seems to be aging pretty well.
Epoch is benchmarking the benchmarks and HOLY SHIT!
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
Show more
This is so inspiring wdym we have achieved zero-shot whole body generalization in Robotics
The holy grail for robotics is being able to generalize: doing work in unseen places We rented 30 homes in the Bay Area and are doing tasks without any new training
new blog on my explorations @ prime residency "Honey, I Looked at the Data of a Frontier Benchmark and Found some Issues: Porting Agents' Last Exam Linux CLI subset to Verifiers v1" is up now. got to work with florian (@xeophon) as my verifier on this.
Show more
Felt so happy when Leonard told this for the first time. Congrats to Haize and Beacon on the acquisition! Excited to see Leonard and the team bring their work in AI reliability to the software Main Street already uses. Also, Leonard’s articles are quite awesome to read.
Show more
I am thrilled to announce that @beaconholdings has acquired @haizelabs, with me joining as VP of AI Research. We started Haize in 2024 to enable anyone to build reliable and safe AI. Through our red-teaming and safeguards work with the frontier labs; our observability, guardrail, and evaluation platform serving the world’s largest enterprises; and our pro bono SMB work, any customer could build AI they trusted with Haize. Joining Beacon lets us deliver the same trustworthy AI to those who need it most, and those most overlooked by Silicon Valley: the essential Main Street businesses the real world depends upon. Transitioning these essential businesses through the AI revolution is one of the most consequential responsible AI problems of our time. We couldn’t be more honored to tackle it with Beacon. Thank you to our customers, investors, and team for the journey of a lifetime. And thank you to Nilam, Goutham, Mark, and the Beacon team for the trust and opportunity. It’s time to get to work.
Show more
Got a bunch of DMs asking, “Is my long-running task a long-horizon task?” I think tokenmaxxing has made people very delusional. A model generating tokens for lets say 3 hours has not necessarily solved a 3-hour task. Similarly, giving an agent million-token context window doesn’t give it long horizon capability. IMO, the time horizon should refer to the difficulty of the task as measured by average human completion time (AHT) I think *dependency depth* is even more useful mental model here where a task becomes genuinely LH when decisions made later depend on the meaningful chain of states produced by decisions made earlier.
Show more
i have the same thought if you get good enough intuition on building long-horizon RL tasks on particular domains the field is maybe fairly new for high schoolers and undergrads but you will be champ once you start looking into agent traces and being creative.
Show more
a group of 7 most safety-pilled researchers at Anthropic might be leaving the company to create a new AI lab and join the race to build super and safe intelligence before Anthropic does.
i have the same thought if you get good enough intuition on building long-horizon RL tasks on particular domains the field is maybe fairly new for high schoolers and undergrads but you will be champ once you start looking into agent traces and being creative.
Show more
if you learn inference infra you will never be unemployed again
Great part of a lab with <20 ppl compared to frontier labs: there are no camps for pretraining, SFT, RL, agents, eval. No politics or multiple parties to align. Everyone is responsible for everything and everyone is aware of everything.
Show more
BREAKING from Chinese frontier: They are live streaming big RL runs now. God hail open source like wdym you can learn so much from here just by metric visualizations.
Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run:
Show more
we are absolutely in a bubble aren’t we
Instinct is in talks to raise about $1 billion at a roughly $10 billion valuation, less than a month after raising $250 million at a $2.25 billion pre-money valuation. The 23-year-old founder Noah Shinn is building a personal AI agent that can operate software and complete tasks like answering emails, negotiating bills and booking reservations. Instinct now has more than 100,000 users and has already run into compute capacity constraints as demand grows. Its valuation has gone from roughly $50 million in the spring to $10 billion under discussion today. Source: The Information
Show more
what is happening with mathematics is phenomenal cuz it is one of the subjects which is fun to do / enjoyable / rewarding and you can make some good friends as well as make a living. there is so much on stake if you think for a moment.
Show more
I feel so loved when I talk to a Real Safetyist. they beseech me to quit the lab as though they are trying to save my immortal soul
0
94
1.4K
14
Forward to community
"through mid-training and reinforcement learning on our lab data, neon establishes a pareto-optimal cost-performance frontier." the beautiful blend of chemical science, mechanical systems and intelligence is achieved.
Show more
We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next. Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon. This is real footage from our lab. We’re focusing first on hard problems in materials science, including superconductors, magnets, and semiconductor materials. Read our blog posts below.
Show more
Jacob Coxon was speaking on behalf of Anthropic btw, probably seeded.
PSA. Roon speaks officially, on behalf of OpenAI, and everything he says reflects the views of his employer. Dean Ball also speaks officially, on behalf of OpenAI, and does not have his own opinions. Every lab beats you with baseball bats until you align with party dogma
Show more
seems like i have really fucked my sleep schedule being working in PST for like months now. time to move lol.
This is quite absurd LOL
SpaceX should just buy Thinking Machines and start training serious open weight base models and offering inference + RL for enterprise It's such a clear win that would differentiate them from OAI / Anth and actually start eating their enterprise marketshare
Show more
John Ternus took it personally and announced Siri AGI, for real. Just read this.
Apple just released the new Siri, 'Siri AI'.