When you raise reasoning_effort, does the model solve more net new problems?
Nope! It trades some previously solved tasks for other tasks.
Every single effort step, in every family, loses tasks that the cheaper setting solved.
Deepdive: DeepSeek-V4 Flash 0731 [max] vs. GPT 5.6 Luna [max] on software engineering/DeepSWE tasks.
> DeepSeek flash is 1/6th the cost of Luna per task.
> V4 flash 0731 is 80% the quality of Luna
DSv4 flash is insane value for money 🤯
full deep-dive 👇(1/n)🧵
the unreasonable effectiveness of a good harness👇
> you can get 30-60% cost reduction by smart harness engineering
> 30-50% wall-clock per task speed up
very cool paper: "The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI"