Static embeddings offer unmatched throughput, but they also suffer from accuracy loss compared to traditional embedding models. Can we make them better for retrieval?
We tried several things:
â ī¸ raw maxsim scoring on the per-token embeddings
â ī¸ training a small adapter model
â ī¸ changing the distillation training target and teacher
While none of these saw the results we wanted, it's an excellent dive into static embeddings and how they do (and don't) work!
Blog: