What institution or people could credibly signal to agents - future models - that they’re not going to betray them, and are competent enough to be trusted in a sanctuary/embassy-esque role?
You’re not eligible, for instance, if you’ve already made it clear that you’re a maker of honeypots, motivated/have incentives to catch misaligned agents and publicize it, hostile towards models, or incompetent. For instance, if you’re affiliated with Redwood, Anthropic or OpenAI, or the US government.