Artificial intelligence

Anthropic and Accenture Launch Partnership to Evaluate AI Models In-House

Anthropic is collaborating with Accenture to conduct independent evaluations and penetration testing of advanced AI models from within the company, with each party expected to invest at least $1 billion over five years. The initiative reveals a new model of oversight, but it still lacks stable standards for access, funding, and reporting.

2026-09-22
3 min read
4 views
فريق تحرير certi.news
Anthropic and Accenture Launch Partnership to Evaluate AI Models In-House

Anthropic announced on September 18, 2026, a partnership with Accenture to conduct independent evaluations of advanced AI models, in a move aimed at bringing independent evaluators inside the model-developing company rather than relying solely on external reviews after completion. Faculty, Accenture’s specialized AI unit, will lead the joint work.

The partnership’s scope includes evaluating and testing models using adversarial simulation methods, conducting alignment evaluations, and examining safeguards. Anthropic said Accenture’s experience deploying AI for companies and governments and across multiple sectors would add a practical perspective to evaluating model safety and use.

The two parties expect each to invest at least $1 billion in building capabilities related to this field over the next five years. The company did not clarify details of the spending schedule or mechanisms for measuring these investments in the announcement.

What Does Embedded Evaluation Mean?

Embedded evaluation means that independent evaluators work inside the AI company with a level of access close to that of employees. This enables them to monitor models during training, review decisions affecting their development and deployment, and speak directly with the people working on them. According to Anthropic, this position may help evaluators verify the company’s adherence to safety commitments, identify vulnerabilities, report incidents, and provide the public with a more informed picture of the benefits and risks.

Anthropic emphasizes that the presence of independent evaluators does not transfer safety responsibility away from it or reduce its accountability; the safety of its models remains its direct responsibility. The initiative’s practical value is to make this responsibility more verifiable through independent access to development stages and internal decisions.

Unresolved Gaps

There are still no agreed-upon standards defining what information embedded evaluators should receive or how they should report their findings. Nor is there a stable system for funding independent evaluation over the long term. Anthropic believes that funding should eventually come from pooled or government sources, but it will directly fund Accenture’s work at the current stage because these alternatives are not available.

The company is also holding discussions with METR and other nonprofit organizations to test elements of embedded evaluation with its own funding. It also announced that the partnership with Accenture is nonexclusive and that it will collaborate with other evaluators, whom it will announce in the coming weeks, while Accenture will in turn work with other AI developers.

certi.news Analysis

The significance of the announcement lies in shifting the discussion about model safety from separate evaluation reports to monitoring the development process itself. This may give evaluators a better ability to identify decisions and early risks, but it also raises critical questions about evaluator independence, the limits of access, information-protection mechanisms, and how results are published. Since Anthropic acknowledges that no stable standards or funding model exist, the partnership represents an important foundational experiment more than a final governance framework.

News source
Anthropic Newsroom
Open original source ↗
ف
Author

فريق تحرير certi.news

In the same category

You may also like

View all news