We found that no model is close to client-ready deliverables. Claude Opus 4.8 leads with 69.4% accuracy, ahead of Claude Sonnet 5 (66.3%) and GPT 5.5 (64.5%). When creating models from scratch, numerical correctness is the primary bottleneck: Opus 4.8 passes 87% of formula checks and 74% of presentation checks, but only 61% of numerical checks. A formula can point at the right cell and still compute the wrong value when an upstream input is off. The writing looks correct while the numbers do not.