There’s a correct way to do this using hierarchical bootstrapping to estimate CIs but this leads to CIs that visually appear way too high so nobody does this.
Random note: I’ve been squinting at the statistical validity of published AI benchmarks in the last few days, and while I vaguely believe the mean, I pretty much believe none of the confidence intervals.
Everybody is playing very fast and loose with stats.