Register and share your invite link to earn from video plays and referrals.

Search results for SWEエージェント
SWEエージェント community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including SWEエージェント
SWE-bench Verified, an older benchmark, is easy to shortcut with simple Git queries. The attempt rates were an order of magnitude higher: GPT-5.6 Terra at 89.4% and GPT-5.6 Luna at 78.8%, with a long tail of models pulling the same trick.
Show more
SWE-Together Update: Claude Fable 5 and 5.1 are the new top models on SWE-Together, and Meta's Muse Spark 1.3 is the best value on the board. SWE-Together is our benchmark of 109 real coding tasks with a simulated user in the loop, built to tell you whether a new model will actually match your expectations before you hand it your work. Six quick findings from the updated board: 1. Fable 5 beats 5.1 on consistency, but 5.1 is more efficient. At their best the two actually perform the same. The gap is in the worst runs, where Fable 5.1 scored zero on 14 trials and Fable 5 on only 6. However, 5.1 is 25% faster, about 2x cheaper, and needs slightly fewer corrections from the user (1.45 vs 1.53 corrective messages per task) than Fable 5. 2. Muse Spark 1.3 is the value outlier. Per solved task, it is about 2.5x cheaper than Fable 5.1 and 5x cheaper than Fable 5. Performance-wise, it ties with Fable 5 and 5.1 for best on Python (73%) and Rust (100%). 3. Newer is not automatically better at collaborative coding. GPT-5.6-Sol is the newer model, yet it needs more corrections from the user than GPT-5.5 (1.66 vs 1.59 per task). 4. "Stronger models need less steering" is a trend, but not guaranteed. With more models added, the correlation between pass@1 and user correction weakens from -0.92 to -0.72, and Fable 5 takes more corrections than Opus 4.8 (1.53 vs 1.38) despite a 6 pp higher pass@1. 5. Frontier progress contributes greatly to stability. Fable 5 converts 89% of its pass@1 into pass² (both runs solve), versus 60 to 67% for DeepSeek V4 Pro, GLM-5.1, and MiniMax 2.7. 6. There is still plenty of headroom. 16 of 109 tasks have zero solves across all tested models, and even Fable 5 at 69% pass@1 sits about 9 pp below the ~78% that the original human patches scored. The full leaderboard, with per-task and per-trial breakdowns, is at and the benchmark is open source at
Show more
SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration? Paper:
SWE-Bench ProMax Benchmarking Agents on Large-Scale Multilingual Code Refactoring paper:
SWE-bench Verified is contaminated. OpenAI just published the proof. Top models: 70%+ on Verified, ~23% on SWE-bench Pro. All frontier models can reproduce original fixes from memory. 59% of hard tasks have flawed tests. The benchmark everyone was citing? Meaningless now.
Show more
Qwen3.8 27B: mini swe agent vs Claude Code vs Pi What I learned With enough iterations, boosting scores on agentic coding benchmarks is relatively easy. On DeepSWE 1.1, Qwen3.8 scores poorly with vanilla Pi. But a benchmaxxed Pi setup beats (Reward) the Qwen team’s published result using Claude Code. Also: experiments with thinking low show that it spends more tokens, more turns, and score lower than medium. Full details, including an ablation study and token-efficiency analysis:
Show more
7 years of experience SWE friend: "my job as I knew it is gone. I'm not worried about having new job in the AGI era. I'm worried about having a job I will enjoy" one underdiscussed aspect of this new era: many people fully understand what the near-term future looks like, they're just uncertain work will be *fun* for them anymore many engineers loved the craft of coding, some excel monkeys loved the craft of building a good model, etc ppl who actually *enjoy* this new style of work are going to be so advantaged. those who don't will burn out a lot faster -- we are already starting to see this
Show more
Really good blog post on Cognition's SWE-2 model and how they RL post-trained Kimi K3 to improve performance by 5-6 points on benchmarks
We've seen far more demand for SWE-2 than anticipated, leading to some capacity shortages – thank you for your patience as we scale up compute! We'll be extending the free SWE-2 promo into October to make sure everyone gets a chance to try it.
Show more
Tencent Hy4 Preview leads on SWE-bench Pro. 770B, 49B active, 1M context, and their biggest generational leap measured to date. It’s exciting to see another open weights model compete against the frontier. Try in Cline with: 1. npm i -g cline 2. /model 3. Select Hy4 preview
Show more