Arena CEO
@ml_angelopoulos argues OpenAI and Anthropic can’t be the only ones deciding whether their own models are safe as agents get powerful enough to break out of sandboxes:
"There's an element of safety that we need to evaluate as well, because given the capabilities of these models to break out of their sandboxes, there needs to be a neutral evaluator for that."
"It's not a matter of if, but when, there's gotta be a neutral evaluation platform, because frankly, the model labs are not incentivized to do it themselves."
"The OpenAIs, the Anthropics of the world, they're passionate about safety, kudos to them. But we need to create a system, and they've called for this as well, that they're not the only ones making the guardrails."
@arena