๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Teortaxesโ–ถ๏ธ (DeepSeek ๆŽจ็‰น๐Ÿ‹้“็ฒ‰ 2023 โ€“ โˆž)
@teortaxesTex
We're in a race. It's not USA vs China but humans and AGIs vs ape power centralization. @deepseek_ai stan #1#, 2023โ€“Deep Time ยซCโ€™est la guerre.ยป ยฎ1
๊ฐ€์ž… September 2010
3.3K ํŒ”๋กœ์ž‰ ์ค‘    76.9K ํŒฌ
> One near-term release candidate is the hybrid GDN-2-3B latent MoE, which follows the Nemotron-3 Nano architecture but replaces the Mamba-2 layers with GDN-2. cool!
Weโ€™ve already trained larger variants of this model, and they outperform competing approaches, including Mamba-2, GDP, and KDA, by a substantial margin. One near-term release candidate is the hybrid GDN-2-3B latent MoE, which follows the Nemotron-3 Nano architecture but replaces the Mamba-2 layers with GDN-2. A public release requires several approvals, and weโ€™re working hard to secure them!
๋” ๋ณด๊ธฐ