๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Chutes
@chutes_ai
The most secure AI inference on earth. Powering your next AI App.
๊ฐ€์ž… November 2024
10 ํŒ”๋กœ์ž‰ ์ค‘    11.7K ํŒฌ
Week 16. The rundown: โ†’ One year of Chutes production traffic is now a research paper: 6.12B requests, 35.8T input tokens, 314,970 users. Co-authored with Harvard and UChicago. Request-level metadata only, nobody's prompts are in it: โ†’ The paper's finding: the users are becoming machines. Output lengths fell from hundreds of tokens to under 100 over the year. Agents call often and read short โ†’ Our TEE stack and E2E encryption layer are open source: Security you can't read is a promise. Ours is an audit away โ†’ NVIDIA researchers made the small-model case for agents. Route, extract, format, call a tool. Small-model work at small-model prices Full article below ๐Ÿ‘‡
๋” ๋ณด๊ธฐ