登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Red Hat AI
@RedHat_AI
Accelerating AI innovation with open platforms and community. The future of AI is open.
参加 May 2018
2.1K フォロー中    12.6K ファン
Is an agent that gets it right every time better than one that's nearly as good in a quarter of the time? For agentic work, that's the real question. A 20-call task at ~2s per call is about 45 seconds of wall-clock time. Get each call under 1s and the same task lands around 13. Every step waits on the last, so small per-call latency compounds fast. Which is why "which model is smartest" is giving way to "which is smart enough and fast enough for this step." And you don't always have to trade off. Speculative decoding is lossless and cuts latency 1.5-3x. FP8 quantization keeps 99%+ accuracy. You can make a capable model fast on your own hardware. @_soyr_ on choosing models for agentic work:
もっと見る