가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Daniel van Strien
@vanstriendaniel
Machine Learning Librarian @huggingface 🤗 I like big datasets and small models.
가입 September 2014
1.5K 팔로잉 중    6.7K 팬
The real tokenmaxxing move: use frontier agents to build small classifiers for large-scale data curation. Small classifiers already help curate the data used to train LLMs. I wanted to see how far an agent could take me in building one. I used Astra, SetFit and @huggingface Jobs to turn 200 agent-labelled examples into a reusable document-purpose classifier, reviewing the categories and tricky cases along the way. It classified 191,724 FinePDFs-Edu documents for ~$0.70 in inference compute, versus an estimated $13–26 to label the same document excerpts with low-cost batch LLMs. Training experiments added ~$2.90 in compute. Workflow, mistakes, reusable model and a prompt to try on your own data:
더 보기