๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

ModelScope
@ModelScope2022
Driving innovations with open communities. ๐Ÿ’ฌ Join our Discord:
๊ฐ€์ž… April 2024
183 ํŒ”๋กœ์ž‰ ์ค‘    16K ํŒฌ
Shanghai AI Lab and SJTUโ€™s LUMIA Lab release NCP-ArchPreview, an open-weight 8.9B language modell. ๐Ÿ“œ Apache 2.0. ๐Ÿค– ๐Ÿ“„ โšก Trained on 5.73T Dolma 3 tokens, it reaches OLMo-3-7Bโ€™s final Stage 1 loss with only 51.3% of the tokens, a 1.95ร— convergence gain. ๐Ÿ† Its Stage 1 macro-average rises from 46.59 to 49.04, with +5.99 on GSM8K and +4.28 on HumanEval. ๐Ÿง  NCP jointly predicts tokens and a concept sequence at one-quarter the length, then feeds those concepts back to guide generation. ๐Ÿ›  Domain adaptation updates only the 17M-parameter concept module while keeping the token backbone frozen. ๐Ÿš€ Concept-conditioned drafting improves mean accepted length by 4.17% with negligible overhead.
๋” ๋ณด๊ธฐ