AI Safety: OpenAI and Anthropic Let Strangers In
The proposal positions both companies as willing partners in oversight at a moment when governments from Washington to Brussels are pressing for exactly that.
AI Safety: OpenAI and Anthropic Let Strangers In
OpenAI and Anthropic have announced plans to embed independent safety evaluators directly inside their laboratories — a move that would give outside researchers unprecedented access to the systems, training pipelines, and internal decision-making that have, until now, remained almost entirely opaque, according to TechCrunch.
The proposal positions both companies as willing partners in oversight at a moment when governments from Washington to Brussels are pressing for exactly that. But researchers who study AI risk are reading the fine print carefully. The central concern is structural: evaluators embedded inside a lab are still, in some meaningful sense, inside the lab. Their funding, their access, and their ability to publish findings could all be shaped — subtly or otherwise — by the institutions they are meant to scrutinise.
Anthropic and OpenAI have not yet specified how evaluator independence would be legally protected, how findings would be disclosed publicly, or what authority evaluators would hold if they flagged a system as unsafe before deployment. Those are not details. They are the entire question.
The announcements arrive as both companies are racing to deploy more capable models, and as the gap between what these systems can do and what regulators understand about them continues to widen. Transparency offered voluntarily, researchers note, is not the same thing as transparency that cannot be withdrawn.
What happens next depends almost entirely on whether the independence being promised survives its first serious disagreement.
*— Isla Camilleri, Global Affairs & Lifestyle Editor*