๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Ant Ling
@AntLingAGI
MoE model series with foundation (Ling), reasoning (Ring) and any-to-any (Ming) from Ant Groupโ€™s AGI initiative, @TheInclusionAI.
๊ฐ€์ž… April 2025
5 ํŒ”๋กœ์ž‰ ์ค‘    13.1K ํŒฌ
๐Ÿงต Weโ€™ve open-sourced 6 Base Model checkpoints for Ling-3.0-tiny & Ling-3.0-flash, covering pre-trained, mid-trained, and WSM-merged stages. None has undergone post-training, giving researchers flexible starting points for continued pre-training, fine-tuning, and further research. Two key highlights: - We use WSM to replace LR decay with weighted checkpoint merging, making the training process better suited for continual pre-training while enabling offline exploration of different LR decay strategies. - With one shared training recipe, the community can validate strategies on tiny-base, then scale them to flash-base.
๋” ๋ณด๊ธฐ