Register and share your invite link to earn from video plays and referrals.

Feitong Yang
@feitong_yang
Building products for scientists @OpenAI · AI for science · opinions my own
234 Following    1.2K Followers
watching what people are posting on X after new models are launched: I am pretty optimistic that frontier models release human creativity and expand human imagination, and will continue to do so.
Oh wow, it only needed 6 months for model to go from <1% to >99% for solving Arc3. Maybe with RSI, it would only need 3? Making good eval is harder and harder, but making good eval is more and more important.
Show more
Side note: when we released ARC-AGI-3 in March, and frontier models scored <1% on it, a few Singularitarian poasters took it as a personal insult, and got very worked up about it. They argued the benchmark was fundamentally broken, that it could not even be solved by the smartest humans, that the max reachable score was actually 40%, etc. We had to deal with a torrent of insults and hate poasts since because we had released an unsaturated benchmark. As it turns out, the benchmark is perfectly calibrated. It is straightforward for a human to score 100% if they do better than average people – all you need is to use fewer actions than our human baseline (which is not a strong baseline, as we used unfiltered human testers). And naturally as a result it's also very feasible for AI to score 100% once real progress towards agentic general intelligence has been made. The trajectory of AI from <1% to 100% over the course of 6 months shows that the benchmark was able to snapshot the recent rise of agentic capabilities. And that rise has happened faster than most people expected, including us.
Show more
This is very true. It can be a better startup experience than most startups out there.
I think the best way to describe OpenAI's culture is as a mega startup. It's hard to believe unless you see it from the inside. Extreme ownership, care and pace.
There's a lot of space to explore and innovate in building plugins. Also, is shipped as a collaboration across multiple harness. It should be fun to join the team.
AGI is nothing without its plugins AGI is happening. The models are getting smarter and smarter. We're seeing early signs of recursive self-improvement. Whatever your definition of AGI, it either has happened already or will happen shortly. But, to be maximally useful for people models need to be able to interact with the world. They need to be able to speak to the systems and data that people already use every day. OpenAI can’t build every connection to every system in the world. The people who know those systems best need to build them. That's why I'm at OpenAI working on plugins. Come join me?
Show more
Another Update: in openai, we ARE continuously working on Prism, the scientific/technical writing surface. It is owned by a small team; improvements are coming, although we wish we can ship more. The best way to get update is our discord channel. @vicapow and I are there for you.
Show more
I’ve joined OpenAI recently to build products for the science community. Our first step is Rosalind Workbench — a new environment that brings AI, scientific tools, data, and workflows together, so scientists can follow a question from data to evidence and discovery. Let's go
Show more
Rosalind Workbench connects scientific questions to specialized models, tools, and reviewable outputs in one workflow, from protein structure and sequence analysis to sequencing pipelines.
Show more
This is indeed very inspiring
Today, we're kicking off the first phase of the research preview for Model Hardware Standard (MHS): a new standard for AI agents to safely operate physical equipment in scientific research and advanced manufacturing. Read more:
Show more
This is too California. Doesn't even allow not being sunny at all
Today, we’re launching Meteoric (YC S26). Our drones clear clouds over solar farms to increase their annual output by 10-30% without using any chemicals. Ultimate target: weaken severe storms and hurricanes.
Show more
being good at problem solving != being good at software engineer. even latest LLM follows this statement.
Got to say @elonmusk made this messaging platform incredible.
1/2 Thanks Gavin for an especially thoughtful exchange. I don't usually spend much time on social media but I wanted to engage here because it really brings out the heart of an important conversation. First, on regulation, I think that “either concentrate it in the hands of a chosen few companies and politicians via regulation or distribute it widely” is a false choice.  I know that there’s a sort of Silicon Valley shorthand where regulation = regulatory capture = concentration of power, but I’ve always found this to be an overly simplified picture of the world.  Many people outside this bubble think of regulation as something that constrains corporate power and benefits ordinary people.  I don’t necessarily agree with that perspective either, rather I think it’s complicated and really depends on what the “regulation” consists of.  But in particular I think that those in the “regulation = regulatory capture = concentration of power” frame often underrate the decentralizing power of objective and fair institutional processes.  A crude analogy is that the formal court system can sometimes feel stuffy and elitist, but it does a much better job of defending the rights of vulnerable individuals than the alternative, mob justice.  At their best, institutions can vest power in ideas rather than people, and thereby decentralize that power. This is why Anthropic has always made its policy proposals very carefully.  We try very hard to make proposals that disadvantage (slow down) frontier AI companies while *advantaging* smaller competitors.  California’s SB53 (which we supported), and even the much-maligned SB 1047 (which we were ambivalent on), completely exempt any company below a certain amount of revenue or model training costs from being covered at all (it was $500M for SB 53, lower for 1047 but we objected to that).  More recently the testing process we’ve advocated for at CAISI and the White House involves more rigorous tests for frontier models than off-frontier models — something that differentially advantages challengers.  Similarly, the “Pacing the Frontier” letter envisions (or at least Anthropic’s preferred implementation of it envisions) modulating the pace of the very best models while not constraining those who are catching up.  This hurts the business interests of the frontier labs and helps challengers, including open-weights! Overall my view is that AI is *structurally* a technology that tends to concentrate power, for reasons that have nothing to do with regulation (more to do with the extreme implications of the scaling laws).  Open-weights do help some with this but are nowhere near a sufficient solution because they simply shift the concentration somewhat to those with the most compute and chips (which are roughly the frontier labs plus maybe hardware providers).  By contrast I think the right “rules of the road” can simultaneously (a) address AI’s cyber/bio/alignment risks, (b) institutionally constrain the power of the frontier AI companies, and (c) leave room for open-weights models while also addressing the specific risks that they bring. BTW I do not think that the events of the last few months have “failed to result in [my] preferred regulatory path”.  The approach that the Trump administration is reported to be taking — pre-deployment testing for frontier models, and also testing of open-weights models when they get closer to the frontier — is one that I am very supportive of, though of course I have to see the details to be sure.  I am also supportive of Demis Hassabis’ ideas around a FINRA-like entity.  This contrasts with six months ago when most of the industry was still pushing for preemption of all state regulation and no apparent federal approach either.
Show more
Now I start to feel that the joy of programming is not having some digital product working, but just to generate the beautiful structure of program per se. And letting models to do that is killing such joy... So i finally know why some writers won't enjoy LLM's help
Show more
The models are over engineering so much that I am starting to ask it to go through plan mode again...
Writing skills are good ways to help yourself and your team, but writing agent plugins is going to help other people and I believe will become a business industry soon.
four key people
Announcing Discovery Loop! I am very excited to announce that, along with my longtime friends and collaborators @Sanjay_Ghemawat, @OriolVinyalsML and @quocleix, we are founding Discovery Loop (@DiscoLoopAI), a Public Benefit Corporation whose mission is to automate machine learning, science, and engineering to accelerate discoveries and progress. The four of us have worked together for 14 to 30 years, and have helped build some of the world’s most used products, infrastructure and AI models, and we’re excited to turn our attention to this ambitious endeavor. ♾ Learn more at:
Show more
I can't really share or discuss airtable acquisition news broadly, because people around me don't even know what it is...
This browser is different.
We just raised $5.7M for @PolarBrowser, the AI browser that beats Anthropic and OpenAI on every major web agent benchmark. - 4,500,000+ actions taken for users, automating sales, recruiting, and ops - One company cancelled Clay and saves 25+ hrs/wk per person - Team is from MIT, YC, Prod, Citadel, Jane Street, Perplexity, Modal, and Apple 100 hours of work. From 30 seconds of typing. Download the world's most powerful AI browser:
Show more