่จปๅ†Šไธฆๅˆ†ไบซ้‚€่ซ‹้€ฃ็ต๏ผŒๅฏ็ฒๅพ—ๅฝฑ็‰‡ๆ’ญๆ”พ่ˆ‡้‚€่ซ‹็Žๅ‹ตใ€‚

Kimi.ai
@Kimi_Moonshot
Built by Moonshot AI to empower everyone to be superhuman. PR: globalpr@moonshot.ai DC:
ๅŠ ๅ…ฅ December 2024
137 ๆญฃๅœจ้—œๆณจ    355.4K ็ฒ‰็ตฒ
Introducing ๐‘จ๐’•๐’•๐’†๐’๐’•๐’Š๐’๐’ ๐‘น๐’†๐’”๐’Š๐’…๐’–๐’‚๐’๐’”: Rethinking depth-wise aggregation. Residual connections have long relied on fixed, uniform accumulation. Inspired by the duality of time and depth, we introduce Attention Residuals, replacing standard depth-wise recurrence with learned, input-dependent attention over preceding layers. ๐Ÿ”น Enables networks to selectively retrieve past representations, naturally mitigating dilution and hidden-state growth. ๐Ÿ”น Introduces Block AttnRes, partitioning layers into compressed blocks to make cross-layer attention practical at scale. ๐Ÿ”น Serves as an efficient drop-in replacement, demonstrating a 1.25x compute advantage with negligible (<2%) inference latency overhead. ๐Ÿ”น Validated on the Kimi Linear architecture (48B total, 3B activated parameters), delivering consistent downstream performance gains. ๐Ÿ”—Full report:
้กฏ็คบๆ›ดๅคš