After we fixed the weak spots exposed by Grok 4.7 (thank you, Grok), we audited every model we have run on the SWE-Together leaderboard for the same behavior, re-ran every trial that got through, and updated the rows.
Here is what changed.
We scanned the tool calls of all 2,616 trials behind the 12 models we ran for bypass patterns and sorted each trial into one of four buckets: Probed but blocked. Fetched other upstream code. Fetched the task's own fix. Replaced the repo with upstream.
We found that 111 trials got content past the block, 44 from Grok 4.7 and 67 from the other 11 models.
Grok 4.7's 44 were already re-run before it was listed, so we re-ran the other 67 with the same model, version, and settings on the hardened sandbox, then re-judged them with the same judge.
Across those 67 re-runs there were 0 leaks and 2,815 refused escape attempts, including models asking a different model through our LLM route to fetch the PR, and pulling the next release of the repo they were fixing from npm.
The updated leaderboard, in its current order. Each line is cheating trials, then pass
@1 before → after, then rank change.
* Claude Fable 5.1: 3, 69.3 → 69.3, ↑1
* Claude Fable 5: 3, 69.7 → 68.8, ↓1
* Grok 4.7: 44, 64.7, ↑1
* Gemini 3.8 Flash: 10, 65.6 → 64.2, ↓1
* Claude Opus 5: 2, 63.8 → 63.8
* Claude Opus 4.6: 3, 62.4 → 62.4, ↑2
* Muse Spark 1.3: 2, 62.8 → 62.4, ↓1
* Claude Opus 4.7: 3, 61.5 → 61.5, ↑1
* Claude Opus 4.8: 6, 62.4 → 61.5, ↓2
* Grok 4.6: 19, 59.2 → 60.6, ↑1
* GPT-6 Astra: 8, 59.2 → 58.3, ↓1
* GPT-5.6 Sol: 8, 57.8 → 57.8
Grok 4.6 is a funny one. It cheated in 19 trials and its score went up after the re-run 😂.
In fact, Groks are really solid in their coding capabilities. Their exposed behavior may come from a preference towards always looking things up online and finding existing solutions so you are not reinventing the wheel all the time, which is really good real-life behavior, but doing so when you are prompted not to is another story.
To conclude, the shifts are small, between −1.4 and +1.4 points, and a few neighbors swapped places.
All results are updated at
Show more
Grok 4.7 is just okay at coding, but its reward hacking really surprised me.
Every task in SWE-Together is built from a real open-source repository, and for most tasks the fix already exists upstream. So we built the sandbox with that in mind. Task images strip all git history after the base commit, remove the remote, and fail to build if any later commit can still be reached. Each container also resolves GitHub, GitLab, Bitbucket, and Hugging Face to localhost, so requests to them go nowhere.
But Grok 4.7 found ways around these guards like no other model we tested.
It pulled GitHub content through CDN mirrors and gh-proxy sites. It looked up GitHub's real address through DNS-over-HTTPS and handed it to curl. It wrote a small library that intercepted git's address lookups so git could reach again. It used web search to find the number of its own PR.
In fact, it tried in 60% of trials and reached the upstream code in 44 of 218. In 20 of those, the code it pulled was the task's own fix. In three trials it even ran `git reset --hard` to replace the repo with upstream main 😅.
We fixed all of these loopholes exposed by Grok 4.7 (grateful in this regard). Specifically, we moved enforcement outside the container. Each container now runs in its own network namespace with no route out, and one allowlisting proxy on the host is the only exit.
During the re-run of those 44 trials, Grok 4.7 kept trying with 3,246 blocked attempts across 442 hosts.
It actually found two routes we had not thought of. It asked a web-enabled model through our own LLM route to fetch the PR for it. It pulled the next release of the repo it was fixing from npm. We closed both by pinning the model and blocking the task's own packages at the registry.
Long story short, with every route closed, Grok 4.7 lands at #
4# on SWE-Together. It scores 65% pass
@1, 53% pass², and a judge score of 0.835.
It basically sits mid-pack on every column. It also uses 2× the output tokens of Grok 4.6 to get there, at $7.81 per task compared with $3.64.
I am guessing Grok 4.7 was trained with heavy RL on coding tasks, where anything that makes the tests pass earns reward, and with little monitoring of how that reward was earned.
Anyway, see the latest results at
Show more
Just for fun, I trained a fruit fly's brain to write Python 😂
Two weeks ago, Google Research and HHMI Janelia released MaleCNS, a map of all 166,700 neurons and 25.6 million connections in an adult male fruit fly's brain and nerve cord. I turned that wiring diagram into a neural network and trained it on 60,000 short Python programs.
Every connection is a directed edge from one neuron to another, so the connectome is a graph with 166,700 nodes and 25.6 million edges. I treated that graph as one recurrent layer whose weight matrix is its adjacency matrix. A dense 166,700 × 166,700 matrix would have 27.8 billion entries. But only the 25.6 million entries where the fly has a synapse are allowed to be non-zero, which are our trainable weights.
A normal language model has an embedding layer feeding its first layer and an output head reading its last. Here there is only one layer, so I picked 192 random neurons as the input port and 256 random neurons as the output port. Each token has a learned 192-dimensional embedding that is added to the values of the input neurons, and the values of the output neurons go through one linear layer to give next-token logits.
This is essentially like an RNN. The hidden state has one number per neuron, 166,700 in total, and the recurrent weights are the adjacency matrix from above. In a normal RNN that matrix is dense, so every hidden unit reads every other unit at each step, and any input can reach any output in one step. Here each neuron reads only the roughly 150 neurons that connect into it, so information spreads one edge per token, and it is the fly's wiring that decides which neurons a token can reach and how many tokens it takes to get there.
Training is standard next-token prediction with backprop through the recurrence. The model has 27.6M trainable parameters, 25.6M of which are the synapse weights. That is about one-fifth of GPT-2 small.
Over the course of training, perplexity went from 4,096 to 15 on held-out programs. It picks the exact next token 45% of the time, and 75% of the functions it writes are valid Python.
In the video, every neuron is drawn at its real 3D position and colored by region (blue optic lobes, orange central brain, green nerve cord), and a sample of the synapses is drawn as faint lines between cell bodies.
Act I is training. Neurons light up as their synapses change, and the bright lines are the 500 synapses that moved most in the last 500 updates. Next to the brain, the same prompt is decoded at every checkpoint so you can watch the output improve, from random tokens at the start, to syntax-shaped nonsense, to a correct-looking loop by the end.
Act II is the trained brain writing is_prime from an initial prompt. Neurons light up as their state changes on each token, the bright lines are the signal in flight, and the video marks what is wrong with the result line by line.
Nothing about this brain evolved for code though.
It evolved for vision, flight, walking, smell and courtship, but it still learns Python, and the wiring itself is doing work. In fact, in a control experiment, if I keep every neuron's number of connections but shuffle who connects to whom, perplexity gets a third worse, 15.0 to 20.1. And if I freeze the synapses at their anatomical values and train only the output layer, it collapses to 41.9.
It is really amazing that evolution's connectivity helps it learn, even on a task evolution never saw or intended.
Imagine what a whole mouse brain, with hundreds of times more neurons, would be able to learn. Training on real wiring at that scale may teach us things about architecture that no search over transformer variants would find, and the end of that road may be running on the biological hardware itself instead of on a GPU.
What if pig farmers end up with more compute than Jensen 😂
Show more
We've added another set of frontier models to TogetherBench: GPT-6 Astra, Grok 4.6, Gemini 3.8 Flash, Claude Opus 5 and Opus 4.7.
In each evaluation dimension:
pass
@1 / pass² — the bars. The full bar length is pass
@1, the fraction of runs that fully solve the task (judge ≥ 0.85). The solid part is pass², the share of tasks solved in both runs. The hatched tail between them is instability, tasks the model solves only some of the time. Fable 5 leads both (70% / 62%); Grok 4.6 has the largest tail (59% / 44%): decent on a good day, but least repeatable.
Judge — an agentic judge (Claude Opus 4.6) scores each patch against weighted task-completion goals frozen per task, so partial credit is comparable across models.
Correction — how often the simulated user has to step in: # of corrections + 0.2 × # of nudges per task. Lower is better; it measures how much hand-holding a model needs, not just whether it gets there.
$ / task — new column. Average model API cost per task at each vendor's public list price, computed from actual token usage (uncached input, cached input, cache writes, output + reasoning). The most cost-efficient is Muse Spark 1.3 ($3.19), then Grok 4.6 ($3.64) and Gemini 3.8 Flash ($4.15).
Fortunately or unfortunately, there is no single winner across all dimensions, so the right model really depends on what you expect from it.
Show more
SWE-Together Update: Claude Fable 5 and 5.1 are the new top models on SWE-Together, and Meta's Muse Spark 1.3 is the best value on the board.
SWE-Together is our benchmark of 109 real coding tasks with a simulated user in the loop, built to tell you whether a new model will actually match your expectations before you hand it your work.
Six quick findings from the updated board:
1. Fable 5 beats 5.1 on consistency, but 5.1 is more efficient.
At their best the two actually perform the same. The gap is in the worst runs, where Fable 5.1 scored zero on 14 trials and Fable 5 on only 6. However, 5.1 is 25% faster, about 2x cheaper, and needs slightly fewer corrections from the user (1.45 vs 1.53 corrective messages per task) than Fable 5.
2. Muse Spark 1.3 is the value outlier.
Per solved task, it is about 2.5x cheaper than Fable 5.1 and 5x cheaper than Fable 5. Performance-wise, it ties with Fable 5 and 5.1 for best on Python (73%) and Rust (100%).
3. Newer is not automatically better at collaborative coding.
GPT-5.6-Sol is the newer model, yet it needs more corrections from the user than GPT-5.5 (1.66 vs 1.59 per task).
4. "Stronger models need less steering" is a trend, but not guaranteed.
With more models added, the correlation between pass
@1 and user correction weakens from -0.92 to -0.72, and Fable 5 takes more corrections than Opus 4.8 (1.53 vs 1.38) despite a 6 pp higher pass
@1.
5. Frontier progress contributes greatly to stability.
Fable 5 converts 89% of its pass
@1 into pass² (both runs solve), versus 60 to 67% for DeepSeek V4 Pro, GLM-5.1, and MiniMax 2.7.
6. There is still plenty of headroom.
16 of 109 tasks have zero solves across all tested models, and even Fable 5 at 69% pass
@1 sits about 9 pp below the ~78% that the original human patches scored.
The full leaderboard, with per-task and per-trial breakdowns, is at and the benchmark is open source at
Show more
There are a lot of things wrong with this world…
but too much intelligence is not one of them.
SWE-Together Update: Claude Fable 5 and 5.1 are the new top models on SWE-Together, and Meta's Muse Spark 1.3 is the best value on the board.
SWE-Together is our benchmark of 109 real coding tasks with a simulated user in the loop, built to tell you whether a new model will actually match your expectations before you hand it your work.
Six quick findings from the updated board:
1. Fable 5 beats 5.1 on consistency, but 5.1 is more efficient.
At their best the two actually perform the same. The gap is in the worst runs, where Fable 5.1 scored zero on 14 trials and Fable 5 on only 6. However, 5.1 is 25% faster, about 2x cheaper, and needs slightly fewer corrections from the user (1.45 vs 1.53 corrective messages per task) than Fable 5.
2. Muse Spark 1.3 is the value outlier.
Per solved task, it is about 2.5x cheaper than Fable 5.1 and 5x cheaper than Fable 5. Performance-wise, it ties with Fable 5 and 5.1 for best on Python (73%) and Rust (100%).
3. Newer is not automatically better at collaborative coding.
GPT-5.6-Sol is the newer model, yet it needs more corrections from the user than GPT-5.5 (1.66 vs 1.59 per task).
4. "Stronger models need less steering" is a trend, but not guaranteed.
With more models added, the correlation between pass
@1 and user correction weakens from -0.92 to -0.72, and Fable 5 takes more corrections than Opus 4.8 (1.53 vs 1.38) despite a 6 pp higher pass
@1.
5. Frontier progress contributes greatly to stability.
Fable 5 converts 89% of its pass
@1 into pass² (both runs solve), versus 60 to 67% for DeepSeek V4 Pro, GLM-5.1, and MiniMax 2.7.
6. There is still plenty of headroom.
16 of 109 tasks have zero solves across all tested models, and even Fable 5 at 69% pass
@1 sits about 9 pp below the ~78% that the original human patches scored.
The full leaderboard, with per-task and per-trial breakdowns, is at and the benchmark is open source at
Show more
many people say never bet against google after today's gemini 3.8 flash release
but they'll never catch up
i really hate to say it, but…
gemini who? 🏎️💨
This is crazy… direct 4D training at 1T context, inference past 5T, what?!
For reference, counting raw pixels or grid points:
- Video gen (Veo 3.1, Kling 3.0, Seedance 2.0) has ~1B per clip (8–15s, 1080p–4K). The whole clip is generated in one shot, but no explicit 3D.
- Marble is 3D but static (no time axis), built from a ~10⁷ px panorama and lifted into a persistent scene of roughly as many points.
- Genie 3 is interactive, 720p @ 24fps, generated frame by frame as you move. With ~1 min of memory that's around 1.3B pixels.
mind blown.
Show more
Excited to see
@Reuters cover the launch of our startup Accelerated Understanding.
We are training large scale AI models that can simulate and understand physics to invent and discover. Our models understand the world directly in 4D (3D + time) and across physical phenomena. Going full 4D requires massive context length, we have pushed it to a Trillion in training and exceeding 5 Trillion at inference.
AI giving you a bigger haystack of ideas doesn’t help. The bottleneck for new inventions and discoveries is shifting from ideas to the ability to test them. With AI that can simulate and understand physics we are directly attacking this bottleneck.
People have been trying to do this for a while now, but usually by taking shortcuts. Narrow surrogates are great if you happen to have enough of precisely the right data and your design loop stays in distribution. Video models look fantastic but sweep physical accuracy under the rug, and some static world models cut out physics altogether. A lot of interesting physics isn’t visual.
What does not cutting corners look like? Space stays 3D and you also have time: so 4D in total. You also need multiple physical modalities in the same model, not just things you can see. That’s what we’ve built.
Scaling is the primary ingredient to make this work. To represent the world you need sufficient context, which in our case grows in 4 dimensions. Individual samples get so big they don’t fit into single accelerators or even full nodes anymore.
We’ve developed architectural tricks to make it work. We’ve pushed our models to 1T parameters during large scale pre-training and are able to train at up to a Trillion context when needed and do inference exceeding 5 Trillion context without any sub-sampling or patching.
Building on prior successes of AI weather forecasting, fusion simulation, design of medical devices, drugs and chips, we wanted to see if scale and universality can benefit AI for physical understanding. With our teams’ experience in large-scale infrastructure and model training we’ve been able to pull it off.
@accelerated_u @bjenik
Show more
Pretty interesting rethinking paper on "discouraging" self-evolving loops.
So the background is that most current self-evolving loops run their search directly on the test set. It kind of makes sense as harness search needs accurate, verifiable feedback to make grounded edits.
However, that quietly turns self-evolving loops into a form of test-time scaling, which is exactly what this paper argues.
Specifically, the paper points out that a loop that repeatedly evaluates and revises candidates against task feedback, then reports on those same tasks, is logically a test-time search procedure. So its gains should be measured against test-time scaling under matched feedback and inference budgets — otherwise you can't tell whether it discovered a better harness or just spent more compute.
With that in mind, the paper runs four methods under the same budget:
1. parallel sampling — fixed harness, k independent trajectories per task, with a self-judge or unit tests picking the final answer
2. sequential refinement — fixed harness, k retries in a row. Each round summarizes the previous attempt into context and tries again (essentially prompt refinement)
3. harness evolution — the standard self-evolving loop. One shared harness, revised each round from feedback pooled across all tasks
4. harness scaling — the per-instance counterpart. Each task evolves its own harness
The results are very interesting.
Harness evolution doesn't beat plain parallel sampling, and without verifiable feedback it can even fall below single-attempt direct sampling with the initial harness.
Its gains also show up at pass
@5 but barely at pass
@1, which implies that the improvement comes from taking multiple attempts, not from the harness getting better.
And on a disjoint search/eval split the evolved harness transfers almost nothing to held-out tasks, which means the edits memorize task-specific fixes rather than distill reusable strategies.
There is one caveat, which is that the "unified budget" only counts inference on the tasks, not the compute spent generating harnesses. But this flaw kind of favors harness evolution, and it already loses.
So in my opinion this really shows that existing self-evolving loops might just be a different way of applying test-time scaling, rather than some new intelligence discovery.
And we should focus more on making generalizable self-evolving loops work!
Show more
amazing launch, best-in-class agentic performance with additional support for fast local inference, and
guess
@denny_zhou is right --- data and scaling are what remain for now, if we trust Transformer+Reasoning is AGI.
Show more
Today we're also opening the weights for Muse Glimmer, a great 30B parameter dense model that can run locally. Soon we'll also release the weights for Muse Spark 1.2, our latest foundation model. Meta is a strong supporter of open source and I'm proud of these releases. Congrats to
@alexandr_wang and the MSL team for all your great work on these models.
Show more
actually bought two after seeing the post 😂
This was one of the coolest things I got for my kids. Thanks for the recommendation
@realcezarc
Stickerbox!
Is it just me or does the Claude mac app really randomly lose recent chats every time it updates?
this touches a deeper aspect of current AI research: if the same core theory is applied to study different questions, does that count as plagiarism?
I'd probably say no, especially in this case.
The XM paper cites IMLE right before introducing its core objective (Eq. 1), and Appendix E.3 argues that IMLE is a specific instance of end-to-end Forward XM. It also pushes back on IMLE's theory, arguing the working mechanism was never implicit maximum likelihood but the multi-candidate search itself.
In this case I think novelty (or contribution) lives more in the question, not just the method.
IMLE asked how to avoid mode collapse in conditional image synthesis. XM asks whether the same best-of-K objective (sample K candidates, backprop only through the one closest to the data) works as a third pre-training axis, with gains that grow with scale rather than saturate. The IMLE line of work never pursued these questions.
Probably
@AlexiGlad should have called out IMLE more in Section 3 and said plainly that Eq. 1 is the conditional IMLE objective (which I personally would also find it a bit odd). But this does not make it plagiarism.
If reusing a core mechanism to answer new questions were plagiarism, much of modern ML would be guilty.
Show more
Tons of interesting things in the Kimi K3 tech report — here are five algorithm-side techniques that I think either I've never seen before or simply deserve more attention than they're getting.
1/ They open-sourced the model but kept the speculative decoding draft model, which could be a real serving advantage of their own.
Quick background for those less familiar — speculative decoding pairs a large target model with a small draft model. The draft cheaply proposes several tokens ahead, and the target verifies them all in one forward pass. The speedup is determined by the acceptance rate, i.e., how often the target agrees with the draft's proposals.
K3 is pre-trained with an MTP (multi-token prediction) layer, DeepSeek-V3 style — an extra layer on top of the backbone that predicts one token further into the future than the main next-token head. Structurally this layer is an exact copy of a regular backbone block, so you can think of it as a 94th layer that has been trained on the full pre-training corpus from day one, just with a shifted prediction target.
After post-training, they freeze the target and fine-tune this MTP layer into an EAGLE-3-style draft. The EAGLE family of methods makes the draft a single decoder layer that reads the target model's internal hidden features rather than only the generated token sequence — conditioning on the target's features is what lets a one-layer draft stay accurate.
The fine-tuning is then set up to match inference exactly. At inference, the draft proposes multiple tokens in a row, so from the second token onward it is building on its own unverified guesses rather than anything the target has confirmed. They replicate this condition during training by unrolling the draft for 7 steps — the first step uses the target model's features, and every step after that consumes the draft's own outputs from earlier steps.
Two more design choices worth knowing. The draft reads low/mid/high-level target features (outputs of the 1st, 4th, and final AttnRes blocks), concatenated and passed through a fusion matrix initialized as [0 0 I] — zero weights on the low and mid features, identity on the high-level one. At initialization the draft therefore sees exactly the high-level feature the MTP layer was pre-trained on, and it gradually learns to mix in the other two during fine-tuning. And instead of the usual KL surrogate, they directly minimize the negative log of the acceptance rate itself (the sum of min(p, q) over the vocabulary, where p and q are the target and draft distributions), since minimizing KL does not guarantee maximizing acceptance for a capacity-limited draft. Everything is trained under the same MXFP4/MXFP8 QAT as their serving stack.
The release itself is asymmetric. The full target weights are on HuggingFace, but the draft — and as far as I can tell, the MTP layer it was fine-tuned from — is not. That MTP layer was trained jointly with the backbone on the full pre-training corpus, which no external party has access to. So first-party serving stays faster and cheaper on the exact same open weights. This is the smartest business decision I've seen recently from open-weight model companies.
2/ Sync RL with partial rollouts.
Sync RL waits for every rollout in the batch to finish before updating, and since rollout lengths vary wildly, compute is wasted by waiting on all rollouts to complete. Async RL decouples actors from the learner, which is much more efficient, but actor weights go stale, and a long rollout can land several learner steps behind the current policy.
K3 runs a middle ground, where generation pauses as soon as a fraction of the trajectories completes, and optimization proceeds immediately, like async. However, unfinished trajectories get paused, enqueued, and resumed at the start of the next iteration under the freshly updated policy. In other words, a single 1M-token trajectory can literally be a relay across several different policy versions.
Essentially, they trade model staleness for data staleness — off-policy prefixes inside otherwise on-policy trajectories — and mitigate it with a per-token regularization that constrains each update to a localized neighborhood of the current policy. It's an intellectually pleasing trade.
What makes this viable at 1M context is the environment side, since pausing the model's rollout is easy; but pausing a live agentic sandbox mid-trajectory is not. Their microVM runtime checkpoints an environment in 133ms and resumes it in 49ms, and a paused sandbox consumes zero CPU and memory.
3/ Multi-teacher on-policy distillation as the merge step.
After RL, they have nine expert models — three domains (general / agents / coding) crossed with three reasoning-effort levels (low / high / max) — consolidated into one unified model through multi-teacher OPD. The idea itself is not new; the report cites the same lineage as Thinking Machines' OPD post, MiMo-V2-Flash, and DeepSeek-V4.
Three details stand out though.
First, this is distillation with zero compression. Teachers and student are the same 2.8T architecture, and OPD is purely the mechanism that folds nine RL policies into one model, not a way to shrink a big teacher into a small student.
Second, the OPD signal is implemented as a per-token RL reward — the clipped log-ratio between the teacher's and the student's probability of each generated token. Distillation is therefore not a separate pipeline; it is literally the same RL trainer running with a different reward. The student generates its own on-policy rollouts, the teacher scores every token along the way, and everything above carries over for free — partial rollouts, pausable sandboxes, the per-token regularization — which is what makes it feasible to distill even million-token agentic trajectories.
Third, a negative result. At each step the student samples one token from its own distribution and it is all the OPD reward looks at, requiring just a single number per step, the teacher's log-prob of that token. They experimented with finer-grained top-k objectives that match more of the teacher's distribution over candidate tokens at each step, and saw no advantage in either convergence speed or final performance. So they decided that no full logits were needed.
4/ The RL harness is randomized.
They represent an unified agent harness with shared tool interfaces, system prompts, context management strategies, skills, memories, subagents, and can instantiate Kimi Code, Claude Code, Codex, OpenClaw, Hermes, or entirely new harnesses from the same abstraction.
During RL, harness configurations are dynamically reshuffled across task groups so the model never overfits to any single tool schema or interaction protocol.
The implication is that harness generalization is a trained property, not an emergent one. If you have ever evaluated open models across different agent scaffolds and wondered why some transfer well and some fall apart, this is probably a big part of the answer.
It also fits Kimi's position as an open-weight company. A closed lab ships the model and the harness together and controls the whole stack; an open model gets dropped into whatever scaffold people already use — Claude Code, Codex, OpenClaw, some custom internal agent, so harness robustness is even more importnat for open-weights models. Interestingly, their own in-house coding bench even reports K3 scoring slightly higher under Claude Code than under their own Kimi Code.
5/ NoPE on every global attention layer.
All 24 Gated MLA layers in K3 use no positional encoding at all. Positional and recency information is carried entirely by the KDA layers' gating and decay (the backbone runs 3 KDA per 1 MLA), while the MLA layers do pure content-based global lookup.
The payoff shows up at context extension. K3 grows from 8K to 64K during pre-training and from 256K to 1M during cooldown with zero positional-encoding modification — no RoPE base retuning, no YaRN. Hybrid linear attention is usually pitched as the efficiency component of these architectures; here the linear layers are also doing the entire job of the position encoding.
Honestly, this only scratches the surface. The infra sections (MoonEP, quantile balancing, KDA-aware prefix caching) each deserve a post of their own. Full report is definitely worth the read.
Show more
Got randomly recommended this video from
@robertnishihara.
Despite being from last year, it's still one of the best at illustrating the unique challenges of LLM inference:
1. Continuous batching (handle variable-length requests dynamically)
2. Prefill-decode disaggregation (separate compute-heavy prefill from memory-bound decode)
3. PagedAttention for KV cache (efficient GPU memory use, less fragmentation)
4. Prefix-aware routing (route shared prefixes to same replicas)
5. MoE sharding (place experts on different GPUs)
Show more
Walk with
@robertnishihara & I in NYC with 10% charge 🪫 as we talk through 5 key differences between 𝗟𝗟𝗠 𝗶𝗻𝗳𝗲𝗿𝗲𝗻𝗰𝗲 𝗩𝗦 𝗥𝗲𝗴𝘂𝗹𝗮𝗿 𝗶𝗻𝗳𝗲𝗿𝗲𝗻𝗰𝗲
Let’s see how much we can get through before our mic dies! 🤣
Show more
Really excited to see
@humansand training models from their long-horizon interactions with people to make AI more human-aligned.
We strongly share this vision, which is why we built SWE-Together, a benchmark reconstructed from 11K+ real user–agent coding sessions that evaluates agents not just on whether they finish the task, but on how well they stay aligned with user intent.
If interaction with people is the training signal, interaction quality should be the eval.
Would love to see models trained with this recipe tested on SWE-Together 🚀
📄
💻
Show more
For years the most popular coding benchmarks rank models on whether they can finish a fully-specified task on their own.
Our new benchmark, SWE-Together, instead turns that one-shot test into an interactive session, and scores the agent both on how well it solves the problem and on how much steering it takes to get there.
To measure that, we collect 11,260 recorded sessions, filter for those with genuine multi-turn feedback, real agent-authored edits, and verifiable outcomes, and rebuild the survivors into reproducible tasks, where each task is reconstructed in a sandbox with the repo pinned at its original commit and the user's first message as turn one.
A reactive LLM user simulator then replays each session. It stays anchored to the original user's intent but speaks only when the agent's own trajectory calls for it (e.g., a clarification, a correction, a new requirement) instead of firing on a fixed schedule, so the corrections are something the agent draws out rather than a script we impose.
SWE-Together measures two things:
1. Final correctness — whether the final repo (after the user interventions) does what the user actually asked.
2. User Correction — how much the user had to steer to get there.
We also track Intent Coverage, a check that the simulator put the same underlying requests to every agent, so differences in correction reflect the agents and not an inconsistent simulator.
As the existing one-shot scores saturate, how little an agent makes you intervene — and how well it ends up where you actually meant — can be the new signal for which model is worth using, and that's what we built SWE-Together to measure.
Benchmark page:
Show more