Retrieve more, rerank, get better results. That's what most of us expect from search, but it turns out to be wrong! 🤯
I'm SUPER EXCITED to publish the 141st episode of the Weaviate Podcast with Mathew Jacob (
@mat_jacob1002)! Mathew led the work behind "Drowning in Documents" during his time at Databricks and is now a Ph.D. student at the University of Washington working on ML systems!
This episode dives deep into "Drowning in Documents". This has been one of the most influential papers for us
@weaviate_io as we are exploring scaling reranked retrieval. I think it is a must read for those working in Search and Information Retrieval. 🙌
We begin with an overview of the paper, and then dive into full scoring with cross encoders and phantom hits. We then cover Listwise Rerankers, what next generation cross encoders might look like, and ranking cascades.
On the topic of ranking cascades, I loved learning about Mathew’s work with Melissa Pan (
@melissapan), Negar Arabzadeh (
@NegarEmpr) and collaborators on “Natural Language Query to Configuration for Retrieval Agents”. Per-query “effort” prediction is certainly going to be a huge component on the future of these search systems! (Congratulations to OpenRouter 😆)
We then dove into Mathew’s work on TraceLab, a super exciting effort to understand coding agents such as Claude Code and Codex. ⌨️
This was a super fun conversation, and I really hope you find it useful!
YouTube:
Spotify: