LLMs now beat humans on math accuracy. But does that mean they're actually building understanding on top of the right foundations?
Title: Do LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective from Knowledge Space Theory
URL:
❓ Do LLMs actually build on prerequisite knowledge to get answers right?
💡 Accuracy favors LLMs (92.5% for the best model, Qwen3-80B, vs. 79.6% for humans), but "perfect prerequisite satisfaction" tells a different story: 72.7% for humans vs. only 48.16% for the best LLM. Lots of correct answers rest on shaky foundations.
❓ Does giving them prerequisite hints help?
💡 Surprisingly, prerequisite-grounded context barely outperformed unrelated examples. That points to surface-level pattern matching rather than genuine structured reasoning over prerequisites.
❓ Do strong models at least share a consistent knowledge structure with each other?
💡 Human learner groups overlap at 0.9+ in their knowledge structure, but LLM pairs only overlap 0.38-0.6 — and stronger models diverge even further from humans.
❓ So what's the takeaway?
💡 Accuracy alone hides how differently LLMs "understand" math. Knowledge Space Theory offers a lens that exposes the fragmented structure lurking behind impressive scores.
#
LLMEval# #
MathReasoning#