Register and share your invite link to earn from video plays and referrals.

Search results for MathReasoning
MathReasoning community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including MathReasoning
LLMs now beat humans on math accuracy. But does that mean they're actually building understanding on top of the right foundations? Title: Do LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective from Knowledge Space Theory URL: ❓ Do LLMs actually build on prerequisite knowledge to get answers right? 💡 Accuracy favors LLMs (92.5% for the best model, Qwen3-80B, vs. 79.6% for humans), but "perfect prerequisite satisfaction" tells a different story: 72.7% for humans vs. only 48.16% for the best LLM. Lots of correct answers rest on shaky foundations. ❓ Does giving them prerequisite hints help? 💡 Surprisingly, prerequisite-grounded context barely outperformed unrelated examples. That points to surface-level pattern matching rather than genuine structured reasoning over prerequisites. ❓ Do strong models at least share a consistent knowledge structure with each other? 💡 Human learner groups overlap at 0.9+ in their knowledge structure, but LLM pairs only overlap 0.38-0.6 — and stronger models diverge even further from humans. ❓ So what's the takeaway? 💡 Accuracy alone hides how differently LLMs "understand" math. Knowledge Space Theory offers a lens that exposes the fragmented structure lurking behind impressive scores. #LLMEval# #MathReasoning#
Show more
We just shipped DFlash speculator checkpoints for Qwen3.5-397B-A17B. On math_reasoning: ~5 out of 7 draft tokens accepted on average. On code (HumanEval): ~4.5 out of 7. Both checkpoints trained with the open source Speculators library from @vllm_project. Apache 2.0. Validated on NVIDIA H200. One flag to enable in vLLM: --speculative-config '{ "model": "RedHatAI/Qwen3.5-397B-A17B-speculator.dflash", "num_speculative_tokens": 7, "method": "dflash" }' Check it out:
Show more
Qwen3-8B now has a DFlash speculator! 82.2% first-token acceptance on math reasoning. 3.74 avg tokens accepted per step. Built with the Speculators library. Training compute sponsored by @modal. 🙏
Show more