Did you know you can now put the text of almost all of arXiv on one cheap SSD.
Someone just packaged an enormous arXiv snapshot on Hugging Face
๐ 3,148,796 papers
๐ PDFs for 99.47%
๐ฝ Full archive: 16.08 TB
Sounds ridiculous for Local AIโฆ
Until you see that โฆ
๐ฅ paper_text = 70GB
That contains the resolved TeX text for โฆ
๐ง 2,856,227 papers
๐ 90.7% of all arXiv
And you can download just that dataset by itself.
No 16TB NAS is necessary.
So theoretically you could build a completely local research system using ThumbLLM.exe from a thumb drive!!!
๐พ 70GB arXiv text corpus
๐ local search / embeddings
๐ง local LLM
๐ซ no API required
Ask your local model to search nearly 3 million scientific papers sitting on your own SSD. ๐
Caveats
โ ๏ธ Itโs a snapshot, not live arXiv
โ ๏ธ latest submission is Aug. 27
โ ๏ธ the text is resolved TeX, so it still contains LaTeX syntax/macros and needs some cleaning for ideal RAG
But 2.8M papers in ~70GB is a Localmaxxer dataset.
Link in ALT