Anthropic, OpenAI’s proposed AI risk evaluators may not have enough power to prevent disasters

Anthropic, OpenAI's proposed AI risk evaluators may not have enough power to prevent disasters. Anthropic and OpenAI are proposing embedded AI evaluators to help manage risk of models causing catastrophic harm to society, but the idea has some issues.
What happened
25% off tickets now Back by popular demand: Save up to $300 on Disrupt Close Image Credits:Samyukta Lakshmi/Bloomberg / Getty Images AI Anthropic and OpenAI want to embed safety evaluators. Rebecca Bellan 2:07 PM PDT · September 16, 2026 In a lengthy essay published over the weekend, Anthropic CEO Dario Amodei made a proposal that the AI industry would have rejected instantly even a year ago: embed third-party evaluators inside all frontier AI companies, giving them the power to report safety incidents, assess whether AI models are truly aligned, and share their unvarnished findings with the world.
Amodei did outline a fairly comprehensive proposal that might give evaluators the kind of access they think is necessary, including the right to “publish key findings about risk levels, incidents, practices, and the access they received or didn’t receive — without editorial control by Anthropic. ” But evaluators say such a system will only work if AI companies are actually willing to surrender control over the process. Disrupt 2026: OpenAI, Anthropic, Replit, and more take over 6 industry stages.
Amodei said Anthropic would commit to giving independent evaluators like METR and Redwood Research unprecedented access to the company’s systems. Whether and when Anthropic and OpenAI plan to provide that kind of access is unclear. Part of the framework, says Steidley, should involve standards for what kinds of auditors companies can rely on, lest they try to sidestep the issue by shopping for evaluators that either aren’t qualified or aren’t interested in assessing the most concerning risk.
The wider picture
So far, Meta, SpaceXAI, and Google DeepMind have not committed to embedding third-party evaluators, though DeepMind CEO Demis Hassabis has proposed a separate industry standards body to independently test frontier models. Google, OpenAI, and Anthropic have also privately been discussing AI safety plans for weeks. Topics AI, ai safety, Anthropic, dario amodei, OpenAI When you purchase through links in our articles, we may earn a small commission. Their employees are sounding the siren too, with one researcher publicly quitting and accusing OpenAI and Anthropic of “gambling with our lives”.
They are right that companies should do what’s right even without binding legislation to force them to. (Both OpenAI and Anthropic, to their credit, have indicated that they are willing to voluntarily slow down. ) But relying on companies’ good intentions is a fool’s errand. Explore more on these topicsAI (artificial intelligence)Donald TrumpOpenAIAnthropicSam AltmanElon MuskGrok AIcommentShareReuse this contentMost viewedMost viewed CEO Sam Altman said OpenAI also would commit to the practice, signaling a potentially profound change in how the industry works with outside research groups.
Third-party evaluators who spoke to TechCrunch broadly welcomed the proposal, but said details need to be ironed out — and ideally backed by legislation — if they’re to know whether they will function as truly independent watchdogs or vendors operating on the AI companies’ terms.
What has been reported
That deeper access is becoming more important as models get better at recognizing when they’re being evaluated, raising the risk that they’ll behave well during testing while concealing problematic behavior. As embedded evaluators, we could actually check. ” Historically, AI companies brought in outside reviewers to test finished models shortly before their release. Now, evaluators that TechCrunch spoke to propose giving them access not just to the final model, but to intermediate versions, or “checkpoints,” from its lifetime of training.
AI, said evaluators could compare those checkpoints to determine when concerning behavior emerged, inspect the post-training environment that rewards models for certain behaviors, and check evaluation transcripts and logs to verify a company’s claims about how a model performed. Neither company has shared which evaluators they’ll work with, when they will be embedded, how many they’ll bring on, exactly what systems and information they will be able to access or what can be disclosed to the public, despite repeated questions from TechCrunch.
Gleave noted that meaningful access could extend beyond the models themselves, with evaluators being given access to interview employees to check whether a company’s documentation and public descriptions of its safety practices match what happened internally. By default, he said evaluators are treated like ordinary contractors: bound by restrictive NDAs and agreements that give developers significant control over what can ultimately be published. The time limit There’s also the question of whether reviewers will get enough time and access to do the work they’re being asked to do.
What happens next
When investigating the Hugging Face incident, OpenAI gave METR and Redwood roughly a week on premises to investigate, and both later said they could not draw confident conclusions due, in part, to scope and timing limitations. A similar issue occurred during the pre-release testing for GPT-6 Astra, which OpenAI has touted as its most aligned model yet. That track record leaves evaluators with a basic question: Why should this time be different? Some laws are already forming around the idea of third-party evaluators.
A new law, SB 813, signed this month, creates a framework for state-recognized “independent verification organizations” with expertise assessing AI risks. BOOK NOW Most Popular Jensen Huang took a call from Trump, and showed off something else, too Connie Loizos The 9 buzziest startups from Y Combinator’s latest Demo Day, according to VCs Marina Temkin Dominic-Madori Davis Tesla says it will finally unveil the second-generation Roadster on October 1 Anthony Ha Revolut confirms customer data breach through fake government requests Jagmeet Singh OpenAI puts Pro subscriptions on hold due to Astra demand Sarah Perez Bending Spoons to buy collaboration tools maker Miro for $1.
But over the weekend, all three AI company CEOs called for AI development to slow down in the face of growing, alarming risks.
The report has been compiled by The Daily Waves using information reported across techcrunch.com, theguardian.com. Details are presented according to the information available at the time of publication and may change as authorities, organisers or other relevant parties provide updates.

