Relying on an AI to "think out loud" (called chain-of-thought monitoring) is not a long-term safety solution.
1. As AI models get larger, more thinking happens deep within their layers before they utter a word.
2. AIs can already alter their chain of thought when prompted, so they can control what their monitors do and do not see.
3. Future AIs may may have continuous chains of thought, not necessarily English.
4. Even now, their reasoning is becoming increasing alien, using opaque phrases like "vantages," "marinades," and "watchers."
In the long-term this will not a dependable window into an AI's mind.