Some workloads demand either the highest throughput or the lowest latency. Embedding workloads need both.
We built Baseten Embeddings Inference (BEI) to meet that need, and we're thrilled to partner with
@turbopuffer to power BEI-optimized models in tpuf!