๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | ๐Ÿ”” Follow for AI & Vibe Coding Tips ๐Ÿ‘‡
๊ฐ€์ž… July 2023
549 ํŒ”๋กœ์ž‰ ์ค‘    11.2K ํŒฌ
Did you know you can now put the text of almost all of arXiv on one cheap SSD. Someone just packaged an enormous arXiv snapshot on Hugging Face ๐Ÿ“š 3,148,796 papers ๐Ÿ“„ PDFs for 99.47% ๐Ÿ’ฝ Full archive: 16.08 TB Sounds ridiculous for Local AIโ€ฆ Until you see that โ€ฆ ๐Ÿ”ฅ paper_text = 70GB That contains the resolved TeX text for โ€ฆ ๐Ÿง  2,856,227 papers ๐Ÿ“Š 90.7% of all arXiv And you can download just that dataset by itself. No 16TB NAS is necessary. So theoretically you could build a completely local research system using ThumbLLM.exe from a thumb drive!!! ๐Ÿ’พ 70GB arXiv text corpus ๐Ÿ”Ž local search / embeddings ๐Ÿง  local LLM ๐Ÿšซ no API required Ask your local model to search nearly 3 million scientific papers sitting on your own SSD. ๐Ÿ‘€ Caveats โš ๏ธ Itโ€™s a snapshot, not live arXiv โš ๏ธ latest submission is Aug. 27 โš ๏ธ the text is resolved TeX, so it still contains LaTeX syntax/macros and needs some cleaning for ideal RAG But 2.8M papers in ~70GB is a Localmaxxer dataset. Link in ALT
๋” ๋ณด๊ธฐ