Register and share your invite link to earn from video plays and referrals.

Nikil Ravi
@nikilravi
4.3K Following    644 Followers
Want to watch a 535B parameter (23B active) LLM get trained live? Follow along here
At Marin we’ve found that scaling laws let us simulate the entire training trajectory. This became clear in our last 67B MoE run. The run hit the loss target within 1%, but perhaps more interestingly, arbitrary points during training could also be predicted to within 1%. This model is finishing up long context and will soon enter RL. Based on this finding, we’ve used a scaling ladder to simulate the training trajectory of the 535B MoE we kicked off yesterday. Will see how it goes. Now if someone can figure out how to fit a scaling law to each parameter we can just predict our final model and skip this training business. Something else I've learned from this process is that getting a training run off the ground is way more than just fiddling with the architecture. Lot of herculean efforts from teammates to get trillions of tokens curated, get hardware running on a new GPU stack, improve MFU, and propose great ideas that were tested and integrated. Given that this is our largest run to date, I anticipate lots of learning and adapting along the way. But thats part of the fun. Details:
Show more
I've always felt that the @OpenRouter numbers on token volume was mostly just due to a bunch of people running evals and stuff adjacent to evals, at least for a few weeks after release
What % of global AI token spend do you think is people running evals?
Shouldn’t need saying, but it’s rather tasteless to compliment your own taste
I'm not 100% happy with this, but I am moving on and I think publishing is better than not publishing. So, with Fable and Sol: the interactive "How to Parallelize a Transformer for Training" , based on the hit classic
Show more
Excited to announce the launch of the CSLib Initiative! This will do for CS what Mathlib is doing for math, and enable formal verification of software. Thanks to @coeff_giving for their support. I hope other funders will also support it, and help us formalize other fields.
Show more
A very well-written and timely essay that succinctly summarizes the state of the field of AI security and what needs to be done:
𝗧𝗵𝗲 𝗘𝗻𝗱-𝗦𝘁𝗮𝘁𝗲 𝗙𝗮𝗹𝗹𝗮𝗰𝘆: 𝗪𝗵𝗲𝗿𝗲 𝗜𝘀 𝗔𝗜 𝗦𝗲𝗰𝘂𝗿𝗶𝘁𝘆 𝗚𝗼𝗶𝗻𝗴? Frontier AI models had a giant performance gain in coding in the Fall of 2025 Then with cybersecurity in April This is now happening with open-weight models We are optimistic about the long-term But outside of a few players, we believe the world is not ready in the short-term
Show more
Claim your research profile here! Much more comprehensive than Google Scholar Here’s mine!
Chicago's architecture leaves me in awe every single time I visit
Vals AI is growing! We are hiring for the following positions: - Evaluations Engineer - Head of Research - Member of Technical Staff, Platform - Member of Technical Staff, Research - Clinical Research Scientist, Mental Health AI - Adolescent Mental Health Clinical Expert (contract) Perks: - Collaborate with leading AI labs and work on the front lines of AI development and evaluations - Opportunity for your work to be published and presented at conferences - Unlimited PTO and health insurance - Lunch and dinner covered, plus snacks and coffee - Team socials (pasta night, game night, movies, hikes, pottery) 🔗Application link in the comments
Show more
True, and it seems like much more of the govt/policy world will soon be driven by evals as well
It’s funny that a large chunk of the industry is driven by evals now. Many modeling efforts across the industry are pursuing SoTA eval numbers as their primary goal, well done eval teams! Well done artificial analysis and swebench etc.!
Show more
Introducing People Search for arXiv 🚀 Find collaborators, hire new grads, and see who’s working at the frontier, all with a single query Ask questions like “Which embodied AI researchers are in the Bay Area?” or “How many times has Barret Zoph switched between frontier labs?"
Show more
Something to share: we built a benchmark for evaluating agentic reverse engineering. - instead of grading intermediate outputs, e.g., recovered types/names - we evaluate agents e2e on deterministic goals, e.g., JTAG a firmware to enable its debug mode.
Show more
"Any sufficiently good eval is indistinguishable from reality" is a principle that will probably be very useful going forward. Ender's Law?
today we announced 2 new rounds led by @a16z: - @RayanKrishnan (prev palantir, stanford) and @langstonnashold (prev meta, HRT, NVIDIA) just raised $40M for @ValsAI to scale their independent evaluation layer for LLMs. - @cnitschelm (prev SpaceX) & @trevoroleary (prev Tesla) just came out of stealth with $7M for @UpliftCorp on a mission to "build the foundation a dignified life depends on, starting with the home." and they're hiring! - apply to vals here: - apply to uplift here:
Show more
To avoid all doubt about whether this is related to Site Reliability Engineering, we have changed the name to ReverseEngBench:
New cybersecurity benchmark: SRE-Bench🧵 Models are starting to become competent at finding and patching security vulnerabilities in source code, but can they reverse engineer a binary and understand its behavior? Many existing cybersecurity benchmarks test model capabilities given a source codebase. But most of the software that actually matters for security- the systems defenders protect and the malware they inspect- only exists as binaries. This is true on both sides of the threat landscape. Proprietary enterprise software, security appliances, and firmware are frequent attack targets, yet are almost always shipped as binaries. Malware, meanwhile, is deliberately obfuscated to resist inspection. As agents get more capable, binary reverse engineering becomes a critical, under-tested skill. We worked with collaborators at @Columbia, @ucla, @ucberkeley and @tufts to test this, and we're excited to share SRE Bench: a realistic, contamination-free software reverse engineering benchmark for AI agents. It measures whether an agent can RE a binary well enough to actually understand the underlying code's behavior. Initial results show real separation between frontier models, and how far there is to go:
Show more
Super pumped to partner with Vals AI team on this journey to build the ratings agency for the AI era. @RaghuRaghuram @stuffyokodraws @shangdaxu and the entire AI Infra team are very excited to be working with the Vals team!!
Show more
New benchmark testing a different cybersecurity capability: reverse-engineering a binary!