There are two kinds of questions on a trading desk, and the difference comes down to one test: does the answer depend on whether we are there?
A predictive question does not. Where price goes over the next interval, whether volatility widens — that path is the same path whether or not we participate. This is why it can be checked against history. Run it out of sample, count the hits, and the logic holds.
A decision question does. Which level to quote, how wide, how often to requote, how much to show. Change the action and the market changes with it.
The second kind is the one that actually binds on our desk. It is also the harder one, and the reason is structural rather than technical: history recorded one version of events — the version in which our orders were not in the book. Replaying it tells us how the market went. It cannot tell us how it would have gone at a different quote, or a different size.
So a backtest that looks clean is answering the first kind of question well. That is worth something. It is just not the question we were asking.
What follows from this — how far a simulator has to go before it can answer the second kind, and where it stops being trustworthy — is where most of our tooling effort has gone. More on that in the next piece.