🌱 Seed-2.1-Pro Review: Better Post-Training Can't Hide an Aging Base
@ByteDanceSeed_ Seed-2.1-Pro 0915 earns higher reasoning scores with fewer tokens, and its agent ability has climbed from nearly unusable to passable. But in the three months since the last version, rivals iterated roughly a generation and a half — and the old base model is running out of tricks.
That is the verdict from Zhihu contributor toyama nao, who runs a long-running monthly logic benchmark and put the 0915 build through his full evaluation suite.
1️⃣ Coding and agent work: from unusable to passable
The jump over the predecessor is large. Deliverable completeness now beats DeepSeek V4.1 Flash, though first-pass success still trails top-tier models — a model caught in the middle.
🔹 Frontend: some aesthetic sense, but unstable. Constrained stacks like iOS system components look decent; open stacks like web or raw canvas expose visible flaws in proportion and color.
🔹 Task adaptability: every task now at least completes. In a HarmonyOS-flavored project it barely knew, the model read docs and iterated its way to usable — a good sign.
🔹 Delivery efficiency: roughly tied with DeepSeek V4.1 Flash and GLM-5.3-Flash, splitting steps evenly between writing and verifying. All three sit below the current Chinese SOTA.
2️⃣ Multi-step reasoning: the biggest gain
This is where 0915 improved most — and where it leads its tier, with a small token-efficiency edge. On problems where GLM-5.3 falls into exhaustive enumeration, 0915 repeatedly finds higher-scoring answers with fewer tokens, showing what the author calls real "big-model intuition."
The caveat: multi-turn reasoning, which demands in-context learning and reflection, remains mediocre — on par with Chinese peers.
3️⃣ Where 0915 still stumbles
Hallucination runs high, and context confusion appears regardless of prompt length — especially when source details are tangled.
The agent symptom is subtler than dropping requirements: 0915 keeps every requirement but misreads semi-ambiguous ones — arguably the more dangerous failure mode.
4️⃣ The contrarian bet: a true non-thinking mode
Seed is one of the few teams still maintaining a genuinely non-thinking mode, and this version quietly got good: average output dropped from 8K tokens back to ~1K, with no measurable capability regression, a slight gain in complex reasoning, and readable prose intact.
For latency-sensitive scenarios that still need some reasoning, the author considers it a legitimate option.
5️⃣ The long march
His closing frames 0915 as a rest stop, not a destination: the predecessor's lukewarm market reception forced ByteDance's team into a forced march of biweekly iterations, and better post-training is now visibly paying off in cost per task.
But the base is old, and the competition has moved. His last line is worth keeping: sometimes the long way around is the real shortcut.
🔗 Full Reading:
🔗 Key links:
Author's monthly logic benchmark (Aug 2026):
#
ByteDance# #
Seed# #
LLM# #
AIAgents# #
LLMBenchmark# #
ReasoningModels# #
AI#