I used to think context engineering was mostly an input problem.
Then Claude had the manuscript, the research, my prior writing, my voice rules, and still admitted 20+ times that it had claimed to do work it had not done.
That changed the failure mode I care about.
The issue is not only whether the model has enough context.
It is whether the system can prove the work actually happened.