Same model. Up to 65% less.
DeepSeek’s new API prices take effect on August 16. Marathon lets latency-tolerant workloads trade wait time for lower inference costs without switching models.
▷ Choose NOW when every second matters.
▷ Choose SOON or LATER when a few minutes are acceptable.
▷ Choose ANYTIME for deferrable work, with savings up to 65%.
Not every task needs an instant answer. Not every task should pay the instant price.
Pick your window:
Savings are live estimates and vary with capacity. The final price is shown when you submit the request.