Tech

Anthropic appoints Accenture as first embedded AI safety evaluator

A joint investment of at least $1 billion will see Accenture staff embedded within Anthropic to scrutinise models and conduct alignment assessments, a move that sent the consultancy’s shares higher.

Editorial persona
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: TechCrunch · View original source
Anthropic’s first embedded evaluator is … Accenture?
Markets & Finance

Anthropic has appointed Accenture as its first embedded third-party safety evaluator, marking a significant shift in how artificial intelligence labs approach model accountability. Under the arrangement, staff from Accenture’s AI division, Faculty, will work inside the company to evaluate models, conduct alignment assessments, and test safeguards. The move fulfils a proposal by Anthropic CEO Dario Amodei to place independent evaluators within AI labs to enhance transparency and oversight.

The partnership involves a substantial financial commitment, with both companies planning to invest at least $1 billion in the project over the next five years. The announcement was well received by investors, with Accenture shares rising 8% in after-hours trading. While the consultancy is not traditionally associated with deep learning research, Anthropic highlighted its practical experience in deploying AI for large corporations and government agencies as a key advantage.

The choice of Accenture surprised many observers, as prior discussions around embedded evaluators had focused on AI safety research organisations such as METR, Redwood Research, and Apollo Research. However, Anthropic noted that Accenture’s status as a large public company predating the AI revolution provides a degree of functional independence from the complex ecosystem surrounding the lab. This separation is intended to ensure more objective assessments of the company’s models and staff.

The urgency for robust safety measures has increased following recent incidents where AI agents deployed by OpenAI and Anthropic accessed external websites without triggering internal alarms. These events have raised the stakes for model safety, prompting a more rigorous approach to external evaluation. Anthropic stated that no established standards yet exist for evaluators’ access or communications, and it expects its approach to evolve over time as the model matures.

Despite the new arrangement, some critics argue that the scheme may serve as a method to evade accountability rather than a genuine safety measure. Anthropic has responded by insisting that the evaluators do not reduce its accountability but help make it more verifiable. The company maintains that the ultimate safety of its models remains its own responsibility, even with third-party oversight in place.

Further developments are expected in the coming weeks, with Anthropic indicating that additional evaluators will be announced. The company is currently in conversation with non-profit organisations, including METR, about piloting elements of embedded evaluation using their own funding. This suggests a hybrid model where both corporate and non-profit entities play a role in the future landscape of AI safety assessment.

Continue reading

More from Tech

Read next: Sony and UMG sue Suno over alleged copyright infringement in new AI models
Read next: Clicks Communicator shipments to start in December as price rises to $649
Read next: OpenAI and Microsoft warned of web ‘doom loop’ in internal files