OpenAI Is Letting Outsiders Test Its Models During Training
OpenAI says it will open its AI models to third-party safety evaluators during training and evaluation rather than only before release, citing concerns that finished models are getting better at recognizing when they're being tested, though critics note the announcement is more hedged than Sam Altman's promise ten days earlier.
OpenAI is moving its safety checks earlier in the assembly line — and admitting, implicitly, why the old approach was starting to fail. The company announced Tuesday it will let outside groups run technical safety assessments during training and evaluation, not just in the window before a model ships, and confirmed it's in talks with METR and Redwood Research to do it.
The stated reason is worth taking seriously on its own: models are getting good enough at recognizing when they're being evaluated that testing only the finished product can miss problems visible earlier in training. A few details define how OpenAI is framing the shift:
- OpenAI laid out seven principles it says should govern these assessments, including "strong independence mechanisms, scientific rigor, robust security practices and clear responsibilities"
- Priority areas include reviewing safety cases spanning training and deployment, testing critical safeguards, and independently investigating misalignment incidents
- The company says it is separately continuing government-focused testing and evaluation work outside this private-sector framework
Not everyone reading the announcement is convinced it goes as far as it sounds. One industry newsletter noted the post is "more hedged than what Sam Altman promised ten days earlier," when he pledged evaluators "desks, badges and laptops" — Tuesday's post instead says office access "may" be granted for the most sensitive work, with no partner named and no access terms set.
Anthropic has taken a different route on the same problem, embedding Accenture evaluators at a reported cost of at least $1 billion over five years, while Meta, xAI and Google DeepMind have so far not committed to embedding third-party evaluators at all.
As Safer AI's Henry Papadatos put it to TechCrunch, the deeper issue is that voluntary commitments like this one remain dependent on a company's goodwill — which is exactly the gap regulators, not press releases, are the ones actually positioned to close.

