We're releasing Inference AutoTune
Distill any frontier model into a 1-30B parameter task-specific SLM with only 25 lines of code
automatically route requests to reduce cost and latency by >90%
~2 hours and <$250 to train. You own the weights
Available in private beta today