登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | 🔔 Follow for AI & Vibe Coding Tips 👇
参加 July 2023
549 フォロー中    11.2K ファン
Did you know you can now put the text of almost all of arXiv on one cheap SSD. Someone just packaged an enormous arXiv snapshot on Hugging Face 📚 3,148,796 papers 📄 PDFs for 99.47% 💽 Full archive: 16.08 TB Sounds ridiculous for Local AI… Until you see that … 🔥 paper_text = 70GB That contains the resolved TeX text for … 🧠 2,856,227 papers 📊 90.7% of all arXiv And you can download just that dataset by itself. No 16TB NAS is necessary. So theoretically you could build a completely local research system using ThumbLLM.exe from a thumb drive!!! 💾 70GB arXiv text corpus 🔎 local search / embeddings 🧠 local LLM 🚫 no API required Ask your local model to search nearly 3 million scientific papers sitting on your own SSD. 👀 Caveats ⚠️ It’s a snapshot, not live arXiv ⚠️ latest submission is Aug. 27 ⚠️ the text is resolved TeX, so it still contains LaTeX syntax/macros and needs some cleaning for ideal RAG But 2.8M papers in ~70GB is a Localmaxxer dataset. Link in ALT
もっと見る