Shanghai AI Lab and SJTU’s LUMIA Lab release NCP-ArchPreview, an open-weight 8.9B language modell. 📜 Apache 2.0.
🤖
📄
⚡ Trained on 5.73T Dolma 3 tokens, it reaches OLMo-3-7B’s final Stage 1 loss with only 51.3% of the tokens, a 1.95× convergence gain.
🏆 Its Stage 1 macro-average rises from 46.59 to 49.04, with +5.99 on GSM8K and +4.28 on HumanEval.
🧠 NCP jointly predicts tokens and a concept sequence at one-quarter the length, then feeds those concepts back to guide generation.
🛠 Domain adaptation updates only the 17M-parameter concept module while keeping the token backbone frozen.
🚀 Concept-conditioned drafting improves mean accepted length by 4.17% with negligible overhead.