๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

SemiAnalysis
@SemiAnalysis_
๊ฐ€์ž… January 2024
30 ํŒ”๋กœ์ž‰ ์ค‘    160.9K ํŒฌ
Congrats to @Alibaba_Qwen on the release of Qwen3.8-Flash-Next, using the same architecture innovations as their upcoming Qwen4 model! Such innovations include: ๐ŸŸ  51-billion-param N-gram Embedding to look up a table with very little extra computation, which means the embedding table can be offloaded to slower & less expensive tiers of DRAM ๐ŸŸ  Gated Residual (GR): it seems like a lot of Chinese labs are now innovating on the res connections, like Kimi's AttentionRes and DeepSeek's mHC ๐ŸŸ  Qwen Sparse Attention (QSA): lightning indexer to select context at micro-block granularity Glad to see great Chinese open innovations along with end-to-end model weights to show these innovations can compose well together!
๋” ๋ณด๊ธฐ