Register and share your invite link to earn from video plays and referrals.

Rayan Krishnan
@RayanKrishnan
ceo @ValsAI | solve evals, solve intelligence prev @stanford @PalantirTech
426 Following    2.2K Followers
People underestimate how much better Astra is than Sol in its ability to have novel discoveries. Would not expect this trend to slow.
Joined @EdLudlow to talk about AI progress, independent evaluation, and why I’m optimistic. Some quick takes: - We’re better at building AI than understanding it. Attention towards testing/evaluation matters more than slowing down. - Our RSI Index projects models could match human researchers on the tasks we test by August 2027. Embedded evaluators can produce more accurate estimates based on internal systems. - Public conflict masks cooperation. The labs, policymakers, and enterprises we work with want better evidence. I’ve seen enough to believe coordination is possible. - Independence comes at a cost. We’ve rejected contracts that would compromise ours. The same group doing the testing shouldn’t also sell the solution. - Evaluation should scale through better technology. If it becomes a bureaucratic moat for incumbent labs, we’ve failed. - Market-based evaluation has a role with or without regulation. Competition pushes us to build better technology and keep up with the frontier.
Show more
With AI safety topics going mainstream, the public seems anxious and disconnected from the reality of the problem. From what I’ve seen at Vals AI, I’m optimistic we’ll coordinate toward an optimal future for AI. Our study on RSI shows that, at their current pace, Anthropic’s models could match human researchers by August 2027. That creates urgency but gives us time to prepare. It’s hard for the public to know whom to trust when everyone debating has their own incentives. This was the concern I had when I started Vals AI: that a multipolar paradox would emerge, where the actions of self-interested parties lead to a non-optimal outcome for the system. To overcome this, we independently evaluate models for their real-world impact. This mirrors the role of auditing firms. I’m optimistic because we have found rational and willing partners across the industry. Every major lab has been a great collaborator, providing us with early access to models for testing on our public benchmarks. Every member of Congress and government agency we’ve briefed has been eager to learn. Enterprises are becoming more sophisticated about adopting models based on evaluated capabilities/risks. It hasn't been easy. We have had to earn the trust of competing groups and reject significant contracts that would have compromised our independence. But done right, evaluation can scale with the frontier through automated infrastructure. Embedded evaluators can understand systems during development while maintaining independence. This makes it easier for new entrants to compete and promotes transparency that builds trust with the public. We’re eager to see a diverse ecosystem emerge. We’ve open-sourced our core infrastructure and published our methods, supporting peers. This is a time of high variance. The decisions we make now will have an outsized impact on the future we arrive at. I remain optimistic we will get this right in the year ahead.
Show more
Open source and self reported benchmarks are done. The industry desperately needs a scrupulous third party evaluator before AI sees an Enron, ‘08 crisis or worse.
Meta made a “minor” release to Muse Spark, there’s nothing minor about it. Lots to parse here: - This model is so fucking cheap I almost don’t believe it. In practice we see it’s 1/10 the cost of both Fable and GPT 5.5. If you thought OS models would compete away margins, just wait till you see this. It’s somehow cheaper to use MS 1.1 than host your own OS model… - Coding improvements are significant. This was a real shortcoming in 1.0. But 1.1 sees a ~50% improvement in VibeCodeBench and ~10% improvement in SWE Bench. Not quite SOTA, but at this cost/latency it is still incredibly compelling. - Speaking of latency, wow this model is fast. Across our benchmarks, we find it to be 1/4 the latency of Opus 4.8 and 1/2 the latency of GPT 5.5. I would expect Meta to have incredible web infra, but really don’t know what witchcraft they’re pulling to host the model for such fast inference at high rate limits. - There is a public API. This is the first time Meta has released a model through a hosted API. I’m expecting lots of AI natives to hot-swap and rapidly test this model as a replacement. We’ll soon see if it's performant enough for those production uses. - Speaking of AI natives, it’s been a wild week for Harvey’s legal benchmark. Grok 4.5 held the SOTA position for ~24 hours at 12% before MS 1.1 unseated it with a big jump up to ~20%. I suspect many internal evals will see surprising results like this. Glad to collab with @harvey @gabepereyra @nikogrupen @ItsJulioPereyra on this eval. - Intelligence is more jagged than ever, even within individual domains like legal and coding. Every application and user benefits from staying dynamic. There an edge in picking the right model/system for each task.
Show more
0
104
964
89
Forward to community