Further analysis (and this is more surprising): When thinking is disabled on both, not seeing a clear edge so far from Qwen3.8-27B vs Qwen3.6-27B. Weirder yet, in at least several tests 3.6 no think is successfully out performing while 3.8 gets stuck in cascading loops / aborts.