登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Ari Morcos
@arimorcos
CEO and Co-founder @datologyai working to make it easy for anyone to make the most of their data. Former: RS @AIatMeta (FAIR), RS @DeepMind, PhD @PiN_Harvard.
参加 April 2009
1.7K フォロー中    7.1K ファン
This is more an example of a bad filter rather than why you shouldn't filter. That said, I absolutely agree with @giffmana that overfiltering is commonn and that lots of highly useful data is filtered out in most publicly curated datasets. There other public examples of this: Fineweb and nemotron v1 organic have similarish performance and are derived from largely the same cc dumps, but select very different documents. A bug in the stack v2 pipeline over-deduplicated significantly relative to stack v3. While there is still a role for more precise filters, there's more alpha in thinking about importance weighting and in particular selecting which source data to rephrase. We have found that the choice of source data is absolutely critical and can have a massive impact on the quality of the resultant synthetic data.
もっと見る