Register and share your invite link to earn from video plays and referrals.

Mario Zechner
@badlogicgames
Armin's handler at Old man yelling at Claudes.
1.5K Following    73.4K Followers
that's hilarious
i was curious as to why the claude code users were never reporting the "claude uses bash to edit files" issue. apparently they take a git snapshot before every bash call and produce a fake edit view? fix the RL man 😭
Show more
team decided to go all in on react. i never touched frontend stuff except for lit (don't @ me). perfect timing. thanks @odysseus0z !
and all it took was to add a proper extension system to claude code (sorry thariq, could not resist :D)
We're adding support for AGENTS.md to Claude Code. Starting today in version 2.1.277, if there is no CLAUDE.md in a folder, Claude will check for and use AGENTS.md. You can toggle this behavior in /config.
Show more
recommended reading.
Two weeks. That's how long it took to go from GLM-5.3-Flash's first run on domestic accelerators to serving all of its production traffic, with 3.2× end-to-end throughput along the way. What I keep thinking about is who did much of the work: an Infra Agent powered by GLM-5.3. A model helping optimize the system that serves it. The conditions were hard. Limited memory and interconnect bandwidth. 1M-token context. Multimodal requests. An immature software stack where kernels were missing and documentation was often guesswork. Every optimization was a trade: compute for memory (ReplaySSM), communication for memory (intra-node tensor parallelism), precision for capacity (mixed INT8/FP8/BF16 caching), and disaggregation for scheduling freedom (Encode–Prefill–Decode). But the most important lesson wasn't about any single optimization. When the agent got stuck, it was rarely because it couldn't write the code. It was because it didn't know *why* things got worse. "Throughput down 20%" tells you something broke. It doesn't tell you which layer, which hypothesis, or what to test next. In RL terms, it's a sparse reward with a credit assignment problem. And an end-to-end benchmark that takes hours makes exploration painfully slow. Senior engineers solve this with an implicit process reward in their heads. They know when to check the timeline, when to run a microbenchmark, and which layer's output to compare. So we made that explicit. We call it dense feedback: layered verification interfaces the agent can call directly. Correctness feedback: did it compute right? System behavior feedback: where did the time go? Performance feedback: which option wins, under which conditions? Each signal has to be local, cheap, and objectively verifiable. Three things the agent found: First, precision drift in KDA's context-parallel path that grew with sequence length. The cause was TF32 rounding error compounding through chained state-matrix merges. The fix is now merged upstream in Flash Linear Attention (PR #1180#). Second, KV transfer never overlapped with DeepEP dispatch. The agent followed the call chain across the Python/C++ boundary and found that the intranode path never released the GIL. After the fix, transfer overhead fell from over 30% to under 1%. Third, a decode kernel recomputing the same normalization four times because of how it was chunked. The agent restructured it and got a 1.71× speedup. The idea came from "optimization skeletons" it had distilled by reading existing kernels across SGLang, FLA, and DeepGEMM. To be clear about the boundaries: humans still defined the goals, built the feedback environment, and reviewed every high-risk change. But the engineer's role is changing, from the person who solves the problem to the person who designs the feedback. There's a deeper implication too. A layered, verifiable feedback environment built on real infrastructure tasks is exactly what training the next generation of models needs most. Every task the agent completes can become training ground for its successor. We are still far from recursive self-improvement. But the smallest loop now exists. The model optimizes the system. The system serves the model.
Show more
i'm tending towards agreeing with this (ie we know for a fact openai didn't release whatever swarm harness they used for NS) but i don't understand *why.* is it really just to mislead the public about model capabilities? i'm biased against such arguments due to my gary marcus / ed zitron antibodies but i can't explain it any other way besides even worse pure laziness and disrespect
Show more
>be me >longtermist >future people matter just as much as people alive today >completely agree >there could be trillions of them >also agree >therefore we must decide what they need now >hang on >born in 1990 >making plans for people born in 2190 >they will have 164 more years of science >164 more years of political experience >164 more years of technological progress >possibly machine superintelligence >fortunately they also have me >guy with a pdf >explain that future generations may be unimaginably capable >they could cure diseases >build space habitats >engineer new forms of intelligence >solve problems we cannot even formulate >incredible >anyway i've worked out their priorities >friend asks why they can't decide for themselves >because by then it may be too late >too late for what >for us to tell them what to do >imagine 5000 BC >very serious longtermist meeting >our descendants will face unimaginable dangers >we must prepare >manufacture 80 million stone axes >bury them in strategic caches >future-proof civilization >bronze age begins >lol >lmao >committee publishes report blaming insufficient stone-axis funding >obvious lesson >don't just leave descendants tools >leave them the ability to make better tools >education >institutions >science >wealth >optionality >capacity to change course >sounds suspiciously like helping people who are alive >quickly return to asteroid scenarios >friend says maybe the best thing for 2200 is to help 2050 become competent enough to help 2100 >then 2100 helps 2150 >then 2150 helps 2200 >sounds inefficient >why use succession when i can personally optimize the year 2200 from my laptop >explain moral uncertainty >we should expect future people to know more than us >possibly be morally wiser than us >this is why we must be very careful not to lock in today's values >friend nods >asks whether they should be allowed to reject our values >that's value drift >very important distinction >future people are morally equal to us >they are not epistemically equal to us >obviously >they know less >friend points out they literally live in the future >ignore heckler >write grant proposal >goal: preserve humanity's option value >excellent >preserve options by building institution dedicated to one worldview >hire people who agree the worldview should remain influential for centuries >call this epistemic diversity >someone proposes spending money on children today >ridiculous >their effect on the distant future is impossible to calculate >someone proposes funding my speculative intervention >its effect on the distant future is impossible to calculate >exactly >that's why expected value is so large >imagine ancestor from 1720 >leaves me a letter >dear descendant >i have carefully considered your interests >you must invest in sailmaking >avoid electricity >maintain the divine right of kings >and under no circumstances revise these instructions >laugh at primitive idiot >open document titled >"Recommendations for the Long-Term Future of Humanity" >begin to understand the actual duty to descendants >not "solve their problems for them" >more like "don't destroy the world and leave them a functioning civilization" >this seems disappointingly normal >where are the exponents >try new framing >humanity is an intergenerational relay race >our job is to hand over the baton in good condition >next runner may be faster >may choose a better route >may discover the race is actually swimming >important thing is not welding the baton to our hand >colleague objects >some problems really are irreversible >asteroids >nuclear war >engineered pandemics >AI >fair enough >then show that this problem actually requires action now >show that delay destroys future options >show that your intervention helps >this sounds like a lot of work >can't we just say 10^50 people >realize the funniest part >entire argument for longtermism depends on future civilization becoming vastly more capable than us >entire argument for present control depends on future civilization being unable to reconsider our decisions >Schrödinger's descendant >smart enough to colonize the lightcone >too dumb to edit the roadmap >future humans convert stars into computronium >build minds beyond our comprehension >master matter and energy >open ancient archive from 2026 >"hello descendants, we anticipated your needs" >everyone stops >thank god >the stone-tool people left instructions -- this part is inspired by some @visakanv writings
Show more
I made @typesafeai 's new ultra fast model, Jev, generate text, even though it shouldn't, that's fine because I can't read, and it can't write** ** up to 20 words for 0.5$, what a steal
Show more
brb, quickly grabbing
We are going to lean into making Hermes more like Pi, and less like OpenClaw
You might find our ProgramAsWeights work interesting! We train a neural compiler on (English function description, input, output) examples to produce small neural programs that run locally on CPU. The paper explains how we generated the training data. We've also released the code, model weights, and dataset: Paper: Dataset: You can try it here:
Show more
Was fun speaking here. Speaking at an event after a long time (quite a pain to get approvals at bigcorp, so I never bothered 🫣) Showed how easy it is to modify @pidotdev and make the coding harness "yours". (small part here) Astra + Revealjs potent way to make great slides!
Show more
Breaking: Browser Use + Jev = Ultrafast ⚡ Findings flights took 7s and cost only $0.0039 🤯 > new action space every step > DOM state space > small LLM fallback to type (this video is at 1x speed btw) Built a tiny open source browser agent. try it below ↓
Show more
0
276
9K
643
Forward to community
thought i joined a serious, good company. but they haven't refilled the manner stack, so my only way out is to quit.
I created a little game to help you (or your kids) learn touch typing. It's called Letterby and it supports English, German, and Swedish (which coincidentally is the languages that my kids speak). You can try it out here:
Show more
neat, thanks for sharing. my gut initially said that subagent use in claude code and codex may explain the cost difference (and possibly some of the success rate differences in the oposite direction). but i see you looked at turns. i assume you also looked at which tools where used? pi doesn't jave a subagent tool built-in.
Show more
a data point, grain of salt.
Does your Claude model really need Claude Code…? 🤔 We evaluate 7 models on Claude Code, Codex, and Pi. Three surprising findings emerge: 1️⃣Harness choice has little effect on task success rate, but can significantly affect the cost 2️⃣A simple harness can be competitive 3️⃣The native harness isn’t always the best. Millions of people are using coding agents, but the impact of harness choice remains unclear. (1/n) More details in the thread. 🧵
Show more
he vibe coded the recording tool as i sat next to him at the vienna office. sometimes it's alright to have the ceo vibe slop his way to success.
Radius lets you work with and collaborate on artifacts, natively within Pi
Ran @typesafeai's Jev against an existing classifier eval that previously used Gemini 2.5 Flash Lite. It won both on quality (saturated the eval) and speed (6x)
0
40
1.1K
37
Forward to community