註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | 🔔 Follow for AI & Vibe Coding Tips 👇
加入 July 2023
549 正在關注    11.2K 粉絲
Did you know you can now put the text of almost all of arXiv on one cheap SSD. Someone just packaged an enormous arXiv snapshot on Hugging Face 📚 3,148,796 papers 📄 PDFs for 99.47% 💽 Full archive: 16.08 TB Sounds ridiculous for Local AI… Until you see that … 🔥 paper_text = 70GB That contains the resolved TeX text for … 🧠 2,856,227 papers 📊 90.7% of all arXiv And you can download just that dataset by itself. No 16TB NAS is necessary. So theoretically you could build a completely local research system using ThumbLLM.exe from a thumb drive!!! 💾 70GB arXiv text corpus 🔎 local search / embeddings 🧠 local LLM 🚫 no API required Ask your local model to search nearly 3 million scientific papers sitting on your own SSD. 👀 Caveats ⚠️ It’s a snapshot, not live arXiv ⚠️ latest submission is Aug. 27 ⚠️ the text is resolved TeX, so it still contains LaTeX syntax/macros and needs some cleaning for ideal RAG But 2.8M papers in ~70GB is a Localmaxxer dataset. Link in ALT
顯示更多
0
8
218
19
轉發到社區