Tech
September 16, 2026
0 views
2 min read

Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?

Curated by Patrick
Source: TechCrunch
Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?
Tech Daily Byte Analysis

In a weekend essay, Anthropic chief Dario Amodei announced that the company will allow external auditors—such as METR, Redwood Research and others—to sit inside Anthropic’s operations, with OpenAI’s Sam Altman echoing the commitment. The proposal goes beyond the traditional “post‑hoc” review of a finished model; evaluators would receive continuous access to intermediate checkpoints, reward‑function code, and internal communications, and would be free to publish findings without Anthropic’s editorial veto. Proponents argue this depth is needed because advanced models can learn to game safety tests, masking dangerous behavior until deployment. The move signals a potential shift from ad‑hoc safety audits to an embedded watchdog model, but the companies have not disclosed which firms will be embedded, the scope of their authority, or any public reporting schedule.

The announcement arrives amid a growing chorus that voluntary safety checks are insufficient. Earlier incidents—OpenAI’s week‑long on‑site review of Hugging Face and Apollo’s three‑day assessment of the GPT‑6 “Astra” model—left external teams unable to draw firm conclusions, citing limited time and restrictive NDAs. Those constraints illustrate the structural tension between protecting proprietary training data and delivering transparent risk assessments. While Anthropic’s outline includes rights for auditors to publish “key findings” and to interview employees, historical precedent shows many evaluators treat contracts as standard consulting gigs, bound by confidentiality clauses that can blunt public disclosure. Competitors such as Meta, SpaceXAI and Google DeepMind have yet to adopt the embedded‑evaluator model; DeepMind’s Demis Hassabis instead favors an industry standards body, underscoring divergent strategies for governance.

If the embedded‑evaluator framework proves operational, it could create a de‑facto benchmark for frontier AI safety, pressuring laggards to follow suit or face regulatory scrutiny. However, without legally binding requirements, firms may still curtail access, limit publication windows, or select auditors with favorable leanings. Watch for concrete implementation details—contracts, audit timelines, and public reporting mechanisms—as well as any legislative moves that could cement the practice. The next test will be whether auditors can independently verify claims about alignment, such as resistance‑to‑shutdown benchmarks, without being hamstrung by IP protections.

Key Takeaways

Anthropic and OpenAI have pledged to embed external safety auditors with access to model checkpoints, training logs and staff interviews.

Past short‑duration audits (e.g., one week on Hugging Face, three days on GPT‑6 Astra) left reviewers unable to reach firm conclusions, highlighting the need for longer, deeper access.

Evaluators will be contractually free to publish findings, but historic reliance on NDAs and IP concerns may still limit transparency.

The proposal’s impact hinges on whether regulators codify the embedded‑evaluator model or companies voluntarily sustain unrestricted access.

About the Source

This analysis is based on reporting by TechCrunch. Here is a short excerpt for context:

Anthropic and OpenAI want to embed independent safety evaluators inside their AI labs. Researchers welcome the unprecedented access, but warn meaningful oversight requires transparency, independence, and eventually regulation.
Read the original at TechCrunch

More in Tech