EMB grades each task in two modes. In Template mode, the agent fills in a provided workbook skeleton and is scored on exact cell matches against a gold model. In Scratch mode, it builds the workbook from nothing, with an agent-as-judge scoring the workbook on numerical accuracy, formula wiring, and presentation. One mode tests whether an agent can follow the structure while the other tests whether it can create one.