Chinese researchers did it again!
OpenBMB just open-sourced MiniCPM5-2B, a dense 2B-parameter model built for reasoning, coding, and tool use on resource-constrained hardware.
Artificial Analysis ranked it highest among models under 4B in its Agentic Index comparison. It scored 20, while Granite 4.2 8B scored 9.
The model is particularly strong at coding and tool calling, so I tested both capabilities locally.
I pulled it onto my machine, connected it to a constrained CI repair agent, and gave it one issue:
> A customer reports that retrying checkout with the same idempotency key returns a larger total. The first request returns $109, while the retry returns $118. Find the root cause, fix it without changing the public API contract, and verify the complete test suite.
The Python checkout service had 18 tests. Sixteen passed, while two failed on the retry path.
The agent could list files, search code, read selected ranges, run approved tests, apply a patch, and inspect its diff.
It reproduced the failure, then followed the checkout and idempotency paths through the repository.
The model found that shipping was added to mutable order state before the cached result was checked. On retry, the same order already contained shipping, so the calculation added it again.
It generated a narrow patch that moved the idempotency check ahead of the mutation without changing the public API.
The agent ran the targeted tests and the complete suite. All 18 tests passed.
The model was never told which file contained the issue or what change to make. Each test result, search result, and code inspection determined its next action.
The video below shows the full trajectory, including the investigation, tool calls, generated patch, diff, and final verification.
Everything ran 100% locally on my machine throughout the run.
MiniCPM5-2B supports llama.cpp, Ollama, vLLM, SGLang, iOS, Android, and HarmonyOS for local deployment.
The model weights, training recipes, reasoning datasets, and UltraX data-refinement system are open-source.
GitHub Repo:
A 2B model can now inspect a repository, reason across multiple files, modify code, and verify its patch while remaining small enough to target local hardware.
Show more