This is more an example of a bad filter rather than why you shouldn't filter. That said, I absolutely agree with
@giffmana that overfiltering is commonn and that lots of highly useful data is filtered out in most publicly curated datasets.
There other public examples of this: Fineweb and nemotron v1 organic have similarish performance and are derived from largely the same cc dumps, but select very different documents. A bug in the stack v2 pipeline over-deduplicated significantly relative to stack v3.
While there is still a role for more precise filters, there's more alpha in thinking about importance weighting and in particular selecting which source data to rephrase. We have found that the choice of source data is absolutely critical and can have a massive impact on the quality of the resultant synthetic data.