My 2 cents: one should build the right harness. If a harness dictates how the model works, it fights the model and you get brittleness. If it defines what the task is and checks the result, it enables long horizon task success. That's roughly how we got complex tasks done e2e at IntentLab.
On harnesses, I vacillate between three beliefs:
- the less harness, the better. Models are the magic
- post training a model and harness is dramatically better and the model providers win
- harnesses have real independent value from the model
I have no idea which is right.