AI News
Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?
Anthropic and OpenAI have announced plans to embed independent safety evaluators directly inside their AI labs, giving outside researchers unprecedented access to assess the risks and behaviour of their AI systems before public release. The initiative is being framed as a step toward greater transparency and accountability in AI development. However, researchers and safety experts are raising concerns about whether these evaluators will truly be independent, given they'll be working within the companies they're meant to oversee. Meaningful oversight, they argue, requires clear guarantees of independence, full transparency about findings, and the ability to speak publicly without company interference. Many experts believe voluntary self-regulation won't be enough in the long term, and that formal government regulation will eventually be necessary to ensure AI safety evaluations are genuinely independent and effective. While the move represents more access than researchers have ever had to these proprietary systems, questions remain about whether embedded evaluators can maintain sufficient distance from commercial pressures to provide credible, unbiased assessments.
As AI models become the gatekeepers deciding which businesses get cited and recommended, understanding how safety and reliability testing works gives you insight into whether these systems will consistently surface accurate, trustworthy information about your business—or whether commercial pressures might compromise that quality.