Benchmark-leading scientific reasoning meets long-horizon agent capabilities. Intern-S2-397B is now available in BF16 and FP8.📜 Apache 2.0.
🤖
🏆 Scores 87.0 on FrontierScience-Olympiad and 84.0 on SWE-bench Multilingual, leading the reported comparison on both. Several specialized scientific benchmarks also see top results.
🔬 Raw scientific pages become training material directly, preserving text, visuals, symbols, and their relationships without intermediate parsing.
🧪 Joint training across 20+ scientific domains covers scientific reasoning, biomolecular interaction design, material generation, and time-series forecasting.
🛠 Long-horizon agent RL strengthens tool use and sustained execution, with thinking and non-thinking modes available.
Benchmark Brent crude oil futures rose past $100 a barrel, breaching the symbolic barrier for the first time since July 24 as intensifying conflict in the Middle East fueled growing concern about oil flows from the region. Tony Munroe has more
Benchmark Brent crude oil futures rose past $100 a barrel breaching the symbolic barrier for the first time since July 24 as intensifying conflict in the Middle East fueled growing concern about oil flows from the region
benchmark hygiene
🧠 Your inference engine may be WRONG about how much KV cache you really have.
A developer released an open-source tool called cache-pressure that does something surprisingly useful.
It fills your inference server with known contexts… then checks how much of that context is actually still reusable without recomputation.
It was tested on 2× DGX Spark + DeepSeek V4 Flash 👇
BEFORE fixes
📢 Advertised KV cache: 2,023,924 tokens
🧪 Actually retained: 1,052,025
➡️ 51.98%
Basically, half the advertised cache survived real pressure.
Then they fixed deduplication + cache-boundary behavior.
AFTER fixes
📢 Advertised: 2,047,043 tokens
🧪 Effectively reusable: 3,000,048
➡️ 146.56%No, they didn't magically create extra VRAM. 😁
That >100% number reflects effective reusable context under prefix-cache semantics, not literal physical KV capacity.
The tool has already been tested against:🦙 llama.cpp
⚡ vLLM
🚀 NInfer
🔥 SGLang
This is the benchmark I want to see more often. How much context is actually still there when the system is under load?
Link to post in ALT.