I think there is a general tendency to define the strength of mechinterp as ‘ability to produce cot-esque text about model reasoning’ (e.x. NLAs)
I think this incorrectly ignores mechinterp research as a way to build fundamental understanding and intuition about model behavior
I agree that we probably won’t get to tools that make us happy by hillclimbing the first in <1 year, but I disagree that we won’t be able to make useful leaps in our understanding per the second