What a week with Kimi K3 and Opus 5 launch!
We compared these two great models across SWE (480), Algorithmic(100), Terminal (83), the task level quality is very close with Opus 5 being on par or better, and K3's per-task cost is 2x to 4.6x cheaper, using serverless pricing.
Key measurement is per-task cost, not per-token cost. In general, open models tend to be more verbose than close models. Hence the quick study.
We aim to continue to increase per-task serving efficiency (via Fireworks Inference) and quality (via Fireworks Training).
Share what you find out for the tasks you care about. 👇
More details of our study --