Register and share your invite link to earn from video plays and referrals.

Search results for DeepSWE
DeepSWE community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including DeepSWE
Claude Fable 5 has debuted on DeepSWE Bench with a 66% Pass@1, claiming the #1# spot and edging out GPT-5.5. The result reinforces a broader trend across recent coding benchmarks: strong raw performance combined with consistent reliability and efficiency in real-world software engineering tasks.
Show more
Claude Opus 4.8 has landed on DeepSWE Bench, posting a 58% Pass@1 and taking #2# overall behind GPT-5.5. It continues a broader trend: slightly behind on raw score, but among the most reliable and efficient coding models across recent benchmarks.
Show more
GLM-5.2 is looking surprisingly strong. According to this evaluation across 8 tough benchmarks (SWE-bench Pro, Terminal-Bench, DeepSWE, ProgramBench, Tool-Decathlon, etc.), Zhipu’s latest model is beating or matching the top closed models in several key areas, especially coding and agentic tasks. It leads in: SWE-bench Pro Terminal-Bench 2.1 DeepSWE ProgramBench MCP-Atlas Tool-Decathlon Claude Opus 4.8 still looks very competitive in some areas, but GLM-5.2 is clearly in the conversation now. Interesting to see Chinese labs continuing to close the gap this aggressively on practical coding/agent benchmarks.
Show more
Cost per performance in AI seems to be falling more than ~200x annualized (and higher at higher levels of benchmark performance.) And exceeding 500x annualized cost declines moving along the efficient frontier of the performance curve. Probably more relevant practically for users, at $1 per task whereas today you might only have 50/50 odds that the agent succeeds, by the end of year, you should have an 89% chance of success and by a year from now a 97% chance (at least on deepSWE 1.1 type tasks.) The moving pareto frontier makes it difficult to cleanly report a performance cost improvement due to the shape of the cost performance curve. At the highest asymptote of performance you go from literally not being able to achieve such a low error rate *at any cost* to being able to get that performance for 10s of $s per task (an infinite cost decline). We forecast that top-end rate separately. In the belly of the curve you need only cross a performance cost threshold, but as the curve steepens you can also less expensively buy more performance (though not 100% clear that the curve really is steepening that much.) We also model that performance curve, though given uncertainty effectively zero out continued improvement for the conservative case. And at lower levels of benchmark performance you get a cleaner understanding of the underlying cost-per-performance improvement though at thresholds that are less interesting practically. If anything I suspect that people are wildly under indexing on the rate of improvement of these models. Even last month's AI spend will fall to a fraction of a fraction within the year (if we weren't going to so predictably deploy the new capabilities against new workloads, tasks and challenges.) Spending a lot of time trying to painstakingly carve out specific current workflows into more efficient models, at this point in the improvement curve, seems, if anything, a fool's errand designed to line the pockets of a consultant. Let the tokens flow.
Show more
DEEPSEEK RESUMES FUNDRAISING, SEEKING $8B AT $74B VALUATION
Deepseek transform!!
deepseek flash is currently having capacity issues from the unprecedented volume you may see errors - we're working on a fix
0
253
4.9K
130
Forward to community
DeepSeek Flash did 8T tokens on August 1st 5T of free usage + 3T on OpenCode Go
0
196
5.9K
220
Forward to community
DeepSeek-V4-Flash-0731 (all 284B parameters) running on a single DGX Spark. ⚡ ~17 tok/s generation, 40 tok/s prefill 📦 One 80GB GGUF file, mixed IQ2_XXS/Q8 imatrix quant 🧠 256-expert MoE, 6 routed per token 🆓 MIT licensed frontier model (kinda) on your desk 😎 🤗
Show more