๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Badr Youbi Idrissi
@byoubii
๊ฐ€์ž… August 2015
133 ํŒ”๋กœ์ž‰ ์ค‘    308 ํŒฌ
What happens if we make language models predict several tokens ahead instead of only the next one? In this paper, we show that multi-token prediction boosts language model training efficiency. ๐Ÿงต 1/11 Paper: Joint work with @FabianGloeckle
๋” ๋ณด๊ธฐ