註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Daniel van Strien
@vanstriendaniel
Machine Learning Librarian @huggingface 🤗 I like big datasets and small models.
加入 September 2014
1.5K 正在關注    6.7K 粉絲
The real tokenmaxxing move: use frontier agents to build small classifiers for large-scale data curation. Small classifiers already help curate the data used to train LLMs. I wanted to see how far an agent could take me in building one. I used Astra, SetFit and @huggingface Jobs to turn 200 agent-labelled examples into a reusable document-purpose classifier, reviewing the categories and tricky cases along the way. It classified 191,724 FinePDFs-Edu documents for ~$0.70 in inference compute, versus an estimated $13–26 to label the same document excerpts with low-cost batch LLMs. Training experiments added ~$2.90 in compute. Workflow, mistakes, reusable model and a prompt to try on your own data:
顯示更多
0
6
164
13
轉發到社區