登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Daniel van Strien
@vanstriendaniel
Machine Learning Librarian @huggingface 🤗 I like big datasets and small models.
参加 September 2014
1.5K フォロー中    6.7K ファン
The real tokenmaxxing move: use frontier agents to build small classifiers for large-scale data curation. Small classifiers already help curate the data used to train LLMs. I wanted to see how far an agent could take me in building one. I used Astra, SetFit and @huggingface Jobs to turn 200 agent-labelled examples into a reusable document-purpose classifier, reviewing the categories and tricky cases along the way. It classified 191,724 FinePDFs-Edu documents for ~$0.70 in inference compute, versus an estimated $13–26 to label the same document excerpts with low-cost batch LLMs. Training experiments added ~$2.90 in compute. Workflow, mistakes, reusable model and a prompt to try on your own data:
もっと見る