Register and share your invite link to earn from video plays and referrals.

CHOI
@arrakis_ai
AGI is Here
1.4K Following    12.6K Followers
Build your benchmark,
Introducing Vals-Smith: turn your code base into a customized benchmark. Public benchmarks tell you which model is strongest overall, not which model is the best on your code. Vals-Smith turns your merged pull requests into real coding tasks and measures the percentage a model can actually resolve. New models ship every week. Vals-Smith tells you which one to trust with your code.
Show more
Introducing Vals-Smith: turn your code base into a customized benchmark. Public benchmarks tell you which model is strongest overall, not which model is the best on your code. Vals-Smith turns your merged pull requests into real coding tasks and measures the percentage a model can actually resolve. New models ship every week. Vals-Smith tells you which one to trust with your code.
Show more
A lot of people are saying Google is falling behind after Gemini 3.6 Flash. I think they're reading it the wrong way. To me, Google has changed its strategy. Yes, Gemini is behind GPT-5.6 Luna, Grok 4.5, and Claude Sonnet 5 in coding. But it leads in computer use, visual understanding, and long context. At 1 million tokens, it scores more than twice as high as Gemini 3.5 Flash. That doesn't look like a company that is losing. It looks like a company building for real work. Frontier models are already smart enough for most thinking tasks. Now the question is not who gets another benchmark record. The question is who helps people do their jobs every day. People on X talk about agents and hard benchmarks. Most companies are still trying to figure out where AI fits. Most workers are not running agent systems. They need a model that can read documents, understand charts, keep track of long conversations, and work inside the tools they already use. That is exactly where Gemini is strong. If AGI is about doing every kind of knowledge work, then vision, long context, and real world understanding matter just as much as coding. As Demis Hassabis has said, intelligence has to bring all of these things together. Many developers think Google is losing the coding race. I think Google has stopped chasing benchmark wins and started focusing on where the money is. A fast, low-cost model that fits into everyday work may end up being the better strategy.
Show more
Gemini 3.6 Flash benchmarks are out, and it's... beaten by other models on code tasks, and is only really consistently SoTA on vision and context benchmarks. But hey, 3.1 Pro is now so old 3.6 Flash outperforms it across the board 😭
Show more
0
149
958
60
Forward to community
hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final ((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3, has jacobian determinant -2, and sends (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) to (-1/4, 0, 0)
Show more
0
1.7K
43.7K
5.4K
Forward to community
Vibe Codexer 🚨
Meet kbd-1.0-codex-micro, built with @work_louder. Map the buttons and joystick to your workflow, and keep your pinned chats in view. Get yours before stock returns 410.
This is a crazy update. I’ve been using ChatGPT since the GPT-3.5 days, but GPT-5.6 and ChatGPT Work have given me a genuine wow moment. The difference isn’t just that I’m finding new ways to use it. It’s seeing completely different people experience their own “ChatGPT moment” all over again. Now that everyone is trying Codex, it feels like we’re reliving November 2022. You can almost smell that same excitement in the air—the feeling that something fundamental has changed. This isn’t just another model upgrade. It feels like the beginning of another ChatGPT moment.
Show more
Or… what if we gave you $100 in Codex credits if you tell us what you love about GPT-5.6 Sol or why you switched? Tweet it, claim your gift, enjoy more usage. First 10k get the free tokens!
Show more
I’m seeing a lot of “ChatGPT moments” with Codex (now inside the ChatGPT app). Over the past few weeks, I’ve been visiting manufacturing companies about once a week to introduce Codex. I run one of Korea’s largest AI communities, and I used to think many people weren’t using AI because it was too difficult. What I’ve realized on-site is something different. Most people weren’t struggling because it was hard—they simply didn’t know what was possible. Every visit, I’m surprised by how quickly teams start producing real results once they’re shown the right workflows. Looking at the outcomes and adoption speed, it’s hard to even imagine how much productivity companies will unlock as they embrace AI more aggressively. One more observation that’s surprisingly consistent: Everyone says the same thing at the end… “I’m running out of tokens.” 😂
Show more
I think Pro users should at least get enough Sol Ultra usage to keep working continuously for 5 hours without hitting the limit, regardless of the separate weekly quota. @thsottiaux please...
> According to OpenAI's Borys, the model behind it is essentially the same one as GPT-5.6, which is set to launch tomorrow, with only a thin harness layered on top. In other words, he said that anyone who builds their own harness should be able to achieve similar results with GPT-5.6. THE SUN IS RISING 🌞
Show more
Borys from OpenAI is currently on the livestream for AWTF algorithm track A few highlights from his comments: - "What are your current thoughts on performance of the model?" - "Actually, I think it's quite unexpected. I'd personally expect that it would solve everything. But obviously you have like problem E, so you did a really good job creating hard problems. At least before the competition, we obviously tested our system on previous competitions (...) and the system was able to solve everything and mostly under one hour so similar to today's A, B, C. (...) Today's D and E are actually much harder than any atcoder problem that we saw before." - "There's a model inside and a little bit of harness to make sure that we can extend the test time compute. And the model itself is similar to 5.6. Everyone could write their own harness to basically increase test time compute and get similar results 5.6. - "The progress is huge. I'm pretty sure that half a year ago, we couldn't solve most of the problems here" - "AI doesn't internet access. (...) I think that our models already know a big art of the internet so it doesn't need to google anything"
Show more
WTF... OpenAI just solved all five problems in the Algorithm division of the AtCoder World Tour Finals. Using E869120's well-known difficulty scale, a 1500-rated problem is the kind of challenge researchers can spend months working on. The two 2500-rated problems go far beyond AtCoder's usual upper limit of around 1800. Even the "easiest" 900-rated problem—one of the first three solved in just 38 minutes—is difficult enough that even world-class competitive programmers have roughly a 50-50 chance of solving it within the contest time. The amount of time that the hardest problems humans can create are able to resist AI is getting shorter and shorter. Congrat @OpenAI
Show more
I also had the chance to test the early access version of GPT 5.6 Sol. After spending time with both models, I think they excel in very different ways. Fable 5 stands out because of the intelligence that comes from its sheer model scale. It constantly surprises you with unexpected directions, proposes alternative approaches, and even refines your ideas before you ask. It feels highly metacognitive, with rich language, expressive writing, and a broad way of thinking. In many ways, it reminds me of the feeling I had when GPT-4.5 first arrived. GPT 5.6 Sol feels completely different. I'm not sure I'd describe it as having the same level of metacognition, but when you give it a concrete problem, it becomes incredibly persistent about finding the best possible solution. It doesn't wander—it stays focused. Where Sol really shines is structured execution. If the workflow is clear, it follows it almost perfectly, breaking complex tasks into logical steps, creating sub-agents when appropriate, and systematically solving the problem from start to finish. It feels less like a creative collaborator and more like a dedicated research engineer. Because of that, I think GPT 5.6 Sol is the ideal model when you already know the problem you're trying to solve. Give it a research paper, a technical specification, or a difficult engineering task, and it feels like it'll keep working until it's implemented correctly. The persistence, execution quality, and overall performance are genuinely impressive. The main reason I kept using Fable was its expansive thinking, writing style, and frontend generation capabilities. But if my goal is coding or solving difficult engineering problems, GPT 5.6 Sol would very naturally become my primary model. And at this point, with AGENTS.md and Skills, we already have the tools to teach GPT the workflows and reasoning patterns we want. Sol follows those instructions with remarkable consistency, which makes it an exceptionally strong foundation for agentic coding.
Show more
🚨 Holy sh*t... The movie Her just became reality.
Introducing GPT-Live, a new generation of voice models for natural human-AI interaction. Rolling out in ChatGPT starting today. You’ll want to turn the sound on for this one.
Show more
Holy moly… psyho just became humanity’s last programmer.
Humanity has not prevailed ( • ᴖ • 。) AWTF Heuristic is now over and OpenAI has completely demolished human competitors. In the end, humans performed quite well, but the performance of OpenAI was outstanding. I'll post my detailed thoughts about this in a few days, since this deserves a longer commentary.
Show more
Codex can now hand off threads between local and remote hosts. Start work on your laptop, send it to a remote box before you close the lid, bring it back later. And yes, Codex can orchestrate the handoff for you.
Show more
A bit sad we didn’t get a chance to talk, but I was really happy to see you in person, Mark Chen! Hope you have a great time in Korea and enjoy your stay. @markchen90
Ironically, just days after Dario Amodei argued in his essay Policy on the AI Exponential that the U.S. government should have the legal authority to halt or roll back frontier AI models that fail safety evaluations, Anthropic became one of the first companies to find itself on the receiving end of exactly that kind of intervention. In a sense, the very switch he proposed was flipped on his own company almost immediately.
Show more
The US government, citing national security authorities, has issued an export control directive to suspend all access to Fable 5 and Mythos 5 by any foreign national, whether inside or outside the United States, including foreign national Anthropic employees. The net effect of this order is that we must abruptly disable Fable 5 and Mythos 5 for all our customers to ensure compliance. Access to all other Claude models is not affected. We apologize for this disruption to our customers. We believe this is a misunderstanding and are working to restore access as soon as possible. Read our full statement:
Show more
Holy Shxt...
NEWS: The Trump administration is blocking foreign governments, companies and individuals from accessing Anthropic's most advanced AI models W/@m_ccuri