I thought this part was particularly interesting:
How can we know that two engines are dependent? It is easy to measure. Engines converging on relevant results proves nothing. What they shouldn't share is mistakes. When two engines keep returning the same irrelevant documents for the same queries, the simplest explanation is a shared upstream index. Two students with the same right answer studied. Two students with the same wrong answer sat next to each other. We combine raw overlap and shared-mistake rate into a single chart.