Register and share your invite link to earn from video plays and referrals.

Charlie Marsh
@charliermarsh
@OpenAI. Building Ruff, uv, ty, and other high-performance Python tools with the @astral_sh team.
981 Following    50.9K Followers
The token usage chart here is really striking
We are releasing HLE-Diamond, a refined subset of Humanity’s Last Exam (HLE), following a year-long process of cleaning and refinement with input from various research communities. w/ @ScaleAILabs
Show more
We’re demonstrating how frontier models have continued to improve in realistic mental health conversations with MentalHealthBench. This new open benchmark was built with input from more than 80 mental health clinicians. We’re releasing it openly so other researchers can examine the methods, run their own evaluations, and build on the work.
Show more
0
342
3.8K
210
Forward to community
Looking at some probably-bad code I wrote by hand in college, thinking about the fact that I really thought hard about each line, and feeling a bit nostalgic
0
41
1000
12
Forward to community
Had ChatGPT sometimes crashing on me after updating to macOS 27 and... Astra found a ~14 year old bug in libuv.
0
60
1.2K
49
Forward to community
Fun plot of uv performance over the last 28 releases. Resolution and installation (on real, complex examples) is >2x faster now. Software should get faster, not slower, over time!
Show more
In preview, uv will now omit package metadata from the lockfile. (This is separate from the resolved dependency graph, which we retain.) Internally, this reduced lockfile size by up to 50% (more like 20% on average). Fewer conflicts too.
Show more
Receiving word that I can't in good faith release a new CLI tool named "penguin"
gpt-6-luna (flex, low thinking) is phenomenal for text extraction. Half the cost of the next cheapest and just as accurate as gpt-5.6-luna. I switched from gemini-2.5-flash-lite when I started getting refusals for text extraction (?). Much better quality with luna in general.
Show more
Why did this go for >1k likes. Are you asking me to tweet about Anthropic more.
Alright, I'll bite. Why did they go from Opus 5 to Opus 5.5. Is that a normal sequence.
The best thing about GPT-6-Sol is efficiency. In my testing the difference between GPT-5.6-Sol and GPT-6-Sol is the new version consumes 1/2 the tokens and takes 1/5 of the time to complete the task. On top of the 50% cost reduction, it should be quite a nice daily driver
Show more
0
41
1.3K
41
Forward to community
We are a few releases away from emdashes being considered a sign of non-slop human writing. Then the cycle continues.
THEY REMOVED THE EMDASHES
GPT-6 Sol and Luna are out, and they are better AND 50% cheaper than 5.6. Luna is now $0.10 input / $0.50 output per 1M tokens. This is on top of the 80% price cut to Luna we made at the end of July. Output went from $6 -> $0.50 within two months.
Show more
0
57
1.4K
70
Forward to community
BREAKING: @OpenAI just dropped GPT-6 Sol. It’s my new daily driver in Codex: not quite Astra, but close enough for much of my everyday work, faster, and 50% cheaper than 5.6 Sol. We tested it across the work we actually do at @every and ran it head to head versus Opus 5.5. Here’s my vibe check: • Writing: On a paragraph-writing task drawn from my real work, Sol scored close to Astra, which is still my top model for writing. It writes clean, minimal prose and puts the important idea first. Opus 5.5 is pleasant to work with, but its drafts still tend to bury the point. • Computer use: If you love Astra’s computer use, you’ll like Sol. An earlier Sol preview matched Astra on 17 of 18 attempts across six of our simpler Hands tasks, at a much lower token price. • Coding: Sol improves on GPT-5.6—including better Ruby code in @kieranklaassen tests—but Opus 5.5 has the higher ceiling for long, autonomous builds. One frustration: Codex’s new security classifier repeatedly stopped work we’d already authorized to ask for approval. That’s friction from the classifier, not necessarily a limitation of Sol, but it made long runs harder to leave alone. Overall: Sol feels like an S-class iPhone release: It will give you much of Astra’s power at about a fifth of Astra’s price. Opus 5.5 is the bigger surprise. In some of our tests, it matches or even beats Fable 5.1 on many tasks while also being cheaper than Opus 5. A few people on our team who were Codex converts have started to wobble with Opus 5.5. If you’re already in the Claude ecosystem you’re going to love this model. If you’re a Codex user, it’s worth a look especially for your top-end coding tasks. You’ll like Sol 6 if you spend your day in Codex reading, writing, and getting things done. It’s fast, much cheaper than Astra, and it’s the model I keep reaching for. You’ll like Opus 5.5 if you want to hand an agent a hard coding or visual project and see how far it can take it. Its best work went further than Sol’s in our tests—enough to pull some of our team back toward Claude. You can go deeper on all of our evals including each real-world task and how they're scored at the links below: Writing: Knowledge work: Reading: Full vibe checks on @every coming soon!
Show more
Especially compared by per-task pricing, which is the metric that should matter, I don't think there is anything competitive anywhere in the market. We want people to be able to use tons of AI; it is important to being able to explore this new renaissance in front of us.
Show more
0
494
6.6K
240
Forward to community
GPT-6 Sol & Luna are ~50% cheaper than 5.6
Introducing GPT-6 Sol and Luna, bringing the advances behind Astra to faster, more affordable models. ✨ 💻 Stronger coding and computer use 🎯 Improved factuality and alignment 💬 Clearer answers with less jargon > On AutomationBench, Sol at xhigh effort outperforms Claude Opus 5 at max effort at just 9% of its cost per task. > On OSWorld 2.0 offline, Luna at max effort exceeds GPT-5.6 Sol at medium effort at one-tenth the cost. > Sol makes about half as many mistakes as GPT-5.6 Sol on our internal factuality evaluation of conversations where users previously flagged errors. > Improved prompt caching helps agents reuse more context, with 90% discounts on cached input-token reads. Developers can change reasoning effort and tool availability without breaking cache. Rolling out today in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users, and available in the OpenAI API. Free and Go users can access Luna in the desktop app. These models are not yet available in Chat. API pricing per 1M tokens: Sol: $2 input / $10 output Luna: $0.10 input / $0.50 output
Show more
We have been focusing on efficiency and intelligence for all. Very proud of the team. Only possible when you have incredible models at the top end of the capability that you can then use to make a big difference in everything else.
Show more
0
1.8K
13.3K
329
Forward to community
Alright, I'll bite. Why did they go from Opus 5 to Opus 5.5. Is that a normal sequence.
0
232
2.3K
8
Forward to community
One of my favorite pastime: I closed like ~30 ty issues yesterday by asking Codex to identify stale issues (i.e., issues that were already fixed, deliberately or inadvertently, on main)
Show more
we hire some of the worlds greatest artists, they are now all (100%) using codex in their workflow somehow. just in the last two weeks, unprecedented adoption
GitHub Actions + Rust's Miri can leak your secrets in CI 🧵 If you run Miri in CI: - upgrade to the 2026-09-22 nightly - clear caches - rotate any secrets that the `cargo miri` CI job had access to GPT-6 Astra helped find this 🦀
Show more