Excellent paper on how to actually improve a harness
Think of a junior dev joining a team, their brain doesn't change, their notebook does... two months in they just make fewer dumb mistakes
Same here, the model stays identical, what changes is a folder of notes it loads before each task. Same idea as SKILL.md
Then they changed who decides the task failed
If the model grades its own work, the notes leave it worse than having no notes at all. If a unit test grades it, it goes up
So the reflection was never the valuable part, the verifier is. And that's the loop half of us are shipping right now
Link:
顯示更多