Did you know you can now put the text of almost all of arXiv on one cheap SSD.
Someone just packaged an enormous arXiv snapshot on Hugging Face
📚 3,148,796 papers
📄 PDFs for 99.47%
💽 Full archive: 16.08 TB
Sounds ridiculous for Local AI…
Until you see that …
🔥 paper_text = 70GB
That contains the resolved TeX text for …
🧠 2,856,227 papers
📊 90.7% of all arXiv
And you can download just that dataset by itself.
No 16TB NAS is necessary.
So theoretically you could build a completely local research system using ThumbLLM.exe from a thumb drive!!!
💾 70GB arXiv text corpus
🔎 local search / embeddings
🧠 local LLM
🚫 no API required
Ask your local model to search nearly 3 million scientific papers sitting on your own SSD. 👀
Caveats
⚠️ It’s a snapshot, not live arXiv
⚠️ latest submission is Aug. 27
⚠️ the text is resolved TeX, so it still contains LaTeX syntax/macros and needs some cleaning for ideal RAG
But 2.8M papers in ~70GB is a Localmaxxer dataset.
Link in ALT