Today, the
@OrnnExchange team published a research paper on The Economics of Open-Weight Inference.
We measure cost per completed task across eleven open-weight and eight closed models, the cost of a million output tokens on rented
@NVIDIA A100 through B200, and rental term prices from one month to five years.
On the same model, the A100 can produce a million output tokens for less than half the cost of an H100. A new GPU generation does not make its predecessor economically obsolete.
GPU financing is changing.