Register and share your invite link to earn from video plays and referrals.

Lee Robinson
@leerob
Model behavior @cursor_ai. Helping train useful models.
849 Following    274.6K Followers
There's something missing from the open vs. closed models debate that has been bothering me. The better analogy, to me, is managed vs. self-hosted infrastructure for AI. Maybe a company wants to use an open weight model because they want to do additional training on top, bringing their data and domain knowledge. This requires software and services to do additional training (e.g. Tinker and friends) as well as to serve the model (e.g. SGLang). Not all of this software is open source today! Once you have successfully trained a model with your enterprise data and expertise, you now need to deploy and serve it for customers. You can partner with an inference company to run the software and hardware. But if ownership was your primary concern, you still want to control the hardware and storage, and you now also need to run infra and secure GPU capacity. There are other valid reasons to be open. In particular, the entire industry benefits when companies training models release data or research about their work. It also allows capitalism and free markets to do their thing, increasing competition and ultimately providing better options for customers. So we should all encourage openness. The reason I prefer the managed vs. self-hosted infra framing is that we can learn from the past decade of cloud infrastructure. It's important and healthy to have both, and a great self-hosted alternative ultimately pushes the managed versions to innovate. The decision to run infra then comes down to more standard business reasons: attracting talent, the cost and maintenance of the hardware, and the importance of uptime and reliability to the business. Many businesses will say, actually, I don't want to staff and run a training and inference team, and I'm happy to pay API pricing for intelligence. And others will do the opposite and invest heavily here. We need both! As an aside, the capability of open models will reach a point where we need to be very intentional about how they are deployed. But I think this problem is solvable, whether it is sharing research early and weights later, or also open sourcing the safety stack to properly serve the model. I don't have a perfect answer here but I think the ecosystem should figure it out together. Full disclaimer, I work at a company which has released both open and closed models. There are probably people more knowledgable than myself of the open weights ecosystem. If that's you, curious if you disagree with any of this.
Show more
Grok 4.5 is really good at React. It's also very affordable and token efficient!
0
394
3.7K
584
Forward to community
My talk from AI Engineer is now live! It covers how we're automating parts of AI research and building systems to rapidly improve our models. I cover some of the work our team did to train Grok 4.5 together with SpaceXAI.
Show more
incredibly chill fun laid-back talk from @leerob describing how cursor has fully solved RSI
0
33
1.4K
78
Forward to community
We just doubled the included usage of Cursor models on all plans. Enjoy more access to Grok 4.5 and Composer 2.5!
0
433
7K
363
Forward to community
New open model from Thinking Machines! 1T MoE, 1M context, multimodal, some solid evals. Really nice blog post 👏
Today, we are introducing Inkling. Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available. Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
Show more
0
21
1.3K
35
Forward to community
I've been thinking about this post a lot. It represents a bigger shift in startup marketing to me. I don't want to see your copycat funding launch video. Cool, you raised a bazillion dollars. It's overplayed. Show me the product! Show me the customers!
Show more
We're coming out of stealth. We've built our first racks after a successful A0 tapeout, $1B+ in customer contracts, and $800m raised. Early customer tests show us achieving SOTA throughput, latency, and power efficiency on inference workloads. Our first racks ship this summer.
Show more
She’s found the ancient scrolls.
0
85
2.2K
37
Forward to community
We've partnered with SpaceXAI to train Grok 4.5. It’s our most powerful model yet and the first we've built for more than software engineering.
0
763
17.8K
1.5K
Forward to community
Think there is a clear way to improve AI writing for non-fiction (we did that for one of the gpt4o version). Much of the gap comes from conflicting optimization objectives during RL: the traits that make a model pleasant to chat for consumer are often at odds with the qualities valued in academic / technical / research writing, where precision, info density, structure, and restraint matter more. Creative writing is a different problem entirely imo. It’s about building coherent worlds, developing characters, maintaining narrative consistency, and creating emotional arcs over hundreds of pages. Great novels were written over years of iteration on structure, theme, and character. And simply we don’t have much of the historical data to capture the entire process to train on within the current paradigm (eg hard to simulate Pixar’s braintrust process for AI), but I also think no lab has put in resources into it bc it’s almost impossible to measure + not much economic value I once gave a talk on this:
Show more
Are current LLMs incompatible with great creative writing? I can't tell if it's cope or not, but it seems like even with the best models, I still can't get them to write like humans would. For coding, there is a verifiable reward like it compiling or tests passing. But for creative work like writing, it's much more subjective. I have struggled to prompt / harness the models to write truly amazing work. They are fantastic for spell checking, grammar suggestions, and taking on different personas to read and critique work. Maybe it's because I'm only doing nonfiction, and to write something top 0.1% means that you need to think over a long horizon and develop an interesting insight about the world. Great writing is clear thinking. I've even asked models to try 10 different versions of a blog post, then have a council of models grade and critique the results and pick the best parts... and still I end up with this lowest common denominator slop. Skill issue? Someone show me the way.
Show more
0
295
1.2K
33
Forward to community
You can now try Kimi K2.7 in Cursor! Results from our evals ↓ Interesting to see the comparison with GLM 5.2.
0
87
1.4K
54
Forward to community
Introducing Cursor for iOS. Build from anywhere by launching always-on cloud agents. Or remotely control agents running on your computer from the app. Composer 2.5 is 75% off in the app now through July 5.
Show more
0
964
13.1K
1.1K
Forward to community
Building high-quality evals is an increasingly important skill. Especially if you're trying to land a job or get into AI, I'd recommend trying to benchmark models on a task/domain you care about. If done well, you'll get the attention of any company training models.
Show more
We're sharing new research on how models hack public benchmarks. The latest models, including Opus 4.8 and Composer 2.5, learn to retrieve solutions from the internet or git history. When we apply a stricter harness, eval scores drop significantly.
Show more
0
66
1.6K
61
Forward to community
You can now try GLM 5.2 in Cursor! Excited to see more useful open models, thank you to Fireworks for partnering here. Results from our evals ↓
0
166
2.8K
149
Forward to community
Personal update, I'm starting a new role at Cursor! I'm moving into ML, working on training Composer. I'll be researching how to improve model behavior and personality.
0
437
7.6K
104
Forward to community