"Just read the chain of thought" is one of our best safety techniques. Why does it work? Because models can only think opaquely for a short time, long thinking must be transparent
Can we quantify this? Yes! In our new paper, we show how to measure "time" for arbitrary networks.