getting >30 pp improvements by changing to native tool calling, preserving reasoning + setting correct sampling params
so many evals have those subtle mistakes which are easy to catch if you know where to look. or just use verifiers + the native harness : )