๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Will Held
@WilliamBarrHeld
Open LLM Training @ Formerly ML PhD w/ @Diyi_Yang, ๐Ÿฆ™ @AIatMeta, Assistant @GoogleAI, ุงู„ู„ุบุฉ ุงู„ุนุฑุจูŠุฉ @NYUAbuDhabi Burqueรฑo
๊ฐ€์ž… October 2012
1.1K ํŒ”๋กœ์ž‰ ์ค‘    3.2K ํŒฌ
To train better open models, we need predictable scaling. Delphi is Marinโ€™s first step: we pretrained many small models with one recipe, then extrapolated 300ร— to predict a 25B-param / 600B-token run with just 0.2% error. Getting there took some work ๐Ÿงต
๋” ๋ณด๊ธฐ