đ° A 1.56GB local model just took the #
1# spot for open models under 4B on Artificial Analysis. Let's run it on ThumbLLM.
OpenBMB released MiniCPM5-2B TODAY.
Stats đ
đ§ 2.52B dense parameters
đž Q4_K_M GGUF: just 1.56GB
đ 131K native context
âī¸ Apache 2.0
đĻ llama.cpp
đĸ Ollama
đĨī¸ LM Studio
đ MLX
đ¤ Native tool calling + agent training
Artificial Analysis Intelligence Index v4.2 ...
đĨ MiniCPM5-2B â 15
Qwen3.5-4B Reasoning â 14*
Qwen3.5-9B Reasoning â 15*
Granite 4.2 3B â 11
*Qwen scores are estimated by Artificial Analysis.
So a 1.56GB Q4 GGUF is landing in the same AA Intelligence Index tier as Qwen3.5-9B.
đ¤ GDPval-AA v2 Elo â 831
đĻ ĪÂŗ-Banking â 21%
⥠19K output tokens/task vs 56K for Ling 3.0 Tiny
OpenBMB even released the training data and an official DSpark speculative decoding model.
No trustworthy local tok/s numbers yet.
That is the benchmark I want next. đĨ Let's run it on ThumbLLM!
đ HF /openbmb/MiniCPM5-2B