Evals & Harnesses beyond Code Mode & Pass Rate:
in a Math degree, assignments & exams are largely closed book + no fancy calculator
this is for good reason
tooling can help you get answers, but putting constraints on your tools helps you sharpen innate skills
developing reasoning skills on the underlying structure of problem solving is often more valuable to learn than how to use a tool
this imo totally applies to LLM RL training & the current code mode paradigm
I LOVE code mode and find it amazing how well models can use code to solve any problem,
but using code as a tool is a decision on giving a model a very powerful harness to solve problems, Python is Turing Complete ofc
this has tradeoffs such as not being able to solve other problems
Im seeing this pretty distinctly in visual reasoning tasks that don’t have code mode, models just aren’t great at persistence or visual perception without code
Maybe that’s a fine tradeoff, but AGI level should be able to do tasks without code and I bet training in harnesses without code mode will make them generally more intelligent