Random note: I’ve been squinting at the statistical validity of published AI benchmarks in the last few days, and while I vaguely believe the mean, I pretty much believe none of the confidence intervals.
Everybody is playing very fast and loose with stats.
As a human, I can't publish a book and prohibit labs from training on it under the idea that training on copyrighted material is fair use. But some labs seem to think no such fair use applies to their model outputs?
If you want to ban training on model outputs, abolish fair...