๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Philip Kiely
@philipkiely
Author of Inference Engineering | Early @baseten | Not an LLM (yet)
๊ฐ€์ž… December 2018
888 ํŒ”๋กœ์ž‰ ์ค‘    10.3K ํŒฌ
1M context at these speeds ๐Ÿ‘€
@baseten model performance team is absolutely cracked. @Zai_org GLM 5.2 is now 4x faster running at full 1M context! Already available to use in your favorite coding harnesses, here it is COOKING in @FactoryAI Droid and @opencode Docs for how to get it in comments
๋” ๋ณด๊ธฐ