GLM-5.3 is showing up across top cyber agents.
On CyberGym’s leading-systems board, the two highest-scoring agents both run it. Four of the top six do.
🚨 DeepSeek V4.1-Flash is kind of insane.
It’s already beating GPT-5.6 Sol on several coding + agent benchmarks:
• DeepSWE: 74.2 vs 73.0
• AutomationBench: 54.8 vs 45.8
• Agents’ Last Exam: 31.8 vs 26.7
• CyberGym: 88.1 vs 84.5
And this is only the Flash model.
V4.1-Pro is still coming.
Show more
We’re also introducing Gemini 3.8 Flash Cyber, our most capable cybersecurity model. It shows frontier-level performance in discovering vulnerabilities and patching them at scale, with Flash-level speed & pricing.
That includes achieving 86.2% on the important CyberGym industry benchmark, plus 47.2% on CWE-Bench for patching. We saw a 70%+ success rate in discovering vulnerabilities across 20 programming languages on our internal benchmark.
Show more
DeepSeek V4 Pro 0813 is live on OpenRouter.
@deepseek_ai reports large agent gains over V4 Pro Preview: DeepSWE 62.7 (+49.9), CyberGym 83.3 (+30.6), NL2Repo 61.5 (+23.0), and Terminal Bench 2.1 87.9 (+15.8)
More providers coming online soon
Use it now:
Show more
We want to help all companies be secure, working with the USG and the security ecosystem.
*The full version of GPT-5.5-Cyber is here; state of the art performance on CyberGym.
*Patch The Planet and Codex Security will help solve security problems instead of just finding them.
Show more
Our new multi-model agentic security system brings together more than 100 specialized agents across frontier and custom models to find exploitable bugs, delivering top performance on the CyberGym benchmark.
We used it ahead of Patch Tuesday to help find and fix 16 vulnerabilities. Today we’re announcing that customers can sign up to test it in private preview.
Show more
👑 Atria Dawn Preview is here, built to complete real research and engineering work. 📜 MIT.
🤖
⚙️ Built on a 744B MoE foundation with a 256K context window. Standard and FP8 weights are available.
🔬 Discovery workflows cover evidence gathering, deep research, experiment design, execution, analysis, and recovery from failure.
🏆 Leads the reported comparison on AutomationBench, BFCL v4, CyberGym, DeepSearchQA, and BrowseComp. Scores include 53.8, 77.0, 86.5, 96.0, and 92.5 respectively.
🧩 Creation and delivery capabilities span software, interactive apps, ML systems, visualizations, reports, and presentations.
🛡 Cybersecurity support covers analysis, vulnerability validation, remediation, and retesting in authorized environments.
Show more
Deepseek-V4.1-Flash is available now on Fireworks!
It is a 552B Parameter MoE built for coding, cybersecurity, and agents. It is the ideal workhorse model that outperforms Opus 5 and GPT 5.6 Sol at 1/40th the cost on DeepSWE,CyberGym, and Automation Bench!
Interested in higher quality? Reach out, as we’re bringing this model to Fireworks Training soon.
Try Deepseek-V4.1-Flash now on Fireworks:
Show more
DeepSeek-V4.1-Flash just landed on ModelScope! 🐋
A 552B multimodal MoE with 1M context that activates only 8B params on prefill and 16B on decode. MIT license. 🤖
🏆 New SOTA on DeepSWE v1.1 (74.2), Terminal-Bench 2.1 (90.6), CyberGym (88.1), and Agent's Last Exam (31.8), beating Claude Opus-5.0 and GPT-5.6 Sol.
⚙️ A Causal Encoder-Decoder design built for input-heavy agent work: read huge codebases with 8B active, decode with 16B.
📉 Global KV cache compressed to 890 bytes per token, 4x smaller than DeepSeek-V4-Flash and 437x smaller than V1, making 1M context cheap to serve.
🎛️ Reasoning effort is a continuous dial from 1 to 100, not three presets. Trade cost for accuracy as finely as you need.
Show more
DeepSeek V4 Pro now on
@Bitdeer_AI Model Studio. ⚡
💫💫💫
𝗜𝗻𝗽𝘂𝘁: $0.413 / 1M tokens
𝗖𝗮𝗰𝗵𝗲𝗱: 0.0036 / 1M tokens
𝗢𝘂𝘁𝗽𝘂𝘁: $0.8265 / 1M tokens
✈️ Faster decode, zero extra infra: DSpark speculative decoding and no secondary draft model.
🧠 The Agentic Leap is Real: Massive gains vs. Pro Preview: up to 49.9 points increase on DeepSWE,
@terminalbench and CyberGym
📉 1M Context & Fraction of the Cost: Hybrid CSA+HCA uses. Production ready for full-codebase agent loops.
🤑 Compute on Your Terms: 3 reasoning_effort dials (Non-Thinking/High/Max). Pay for what you need.
The lowest rates are available today on
@Bitdeer_AI Model Studio:
@deepseek_ai 🔥
Show more