🔥We release the first open-source 1.4T-token RAG datastore and present a scaling study for RAG on perplexity and downstream tasks!
We show LM+RAG scales better than LM alone, with better performance for the same training compute (pretraining+indexing)
🧵