Artificial intelligence

AIUC Raises $40 Million to Assess the Safety of Artificial Intelligence Agents

Artificial Intelligence Underwriting Company raised $40 million in a Series A round to develop the AIUC-1 standard and an independent auditing service for artificial intelligence agents. The company tests agents in approximately 5,000 scenarios covering guardrail breaking, hallucinations, and data leaks before issuing an audit report verified by human teams.

2026-09-15
4 min read
46 views
certi.news Editorial Team
AIUC Raises $40 Million to Assess the Safety of Artificial Intelligence Agents

Artificial Intelligence Underwriting Company, known by the abbreviation AIUC, raised $40 million in a Series A funding round led by Ribbit Capital, with participation from First Harmonic, to build an independent layer for auditing and certifying the safety of artificial intelligence agents used within enterprises or developed by model and application companies.

The company is led by Rune Kvist, a former early employee at Anthropic, and Rajiv Dattani, who served as chief operating officer at the artificial intelligence safety research organization METR between 2024 and 2025 and remains a member of its board of directors. AIUC says that Cursor, Lovable, Harvey, and ElevenLabs are among its customers.

A Standard Inspired by Cybersecurity Audits

AIUC is betting on transferring a familiar cybersecurity model to the risks posed by artificial intelligence agents. It has developed a standard called AIUC-1, along with a testing and auditing service that measures how safely and reliably an agent behaves. The idea is based on the SOC 2 standard commonly used to assess technology companies’ controls, but the source does not state that AIUC-1 is an accredited or widely adopted standard outside the company and its customers.

To formulate the standard, the company assembled a coalition of approximately 250 security and risk management officers, the people who purchase agents or decide whether to deploy them. AIUC uses the views of these participants to determine the questions and risks that the tests should cover.

Approximately 5,000 Tests and Final Human Verification

The company runs the agent through a set of approximately 5,000 tests in scenarios covering attempts to break guardrails, hallucinations, and data leaks. The process produces a report of nearly 100 pages explaining the areas in which the agent behaves safely and reliably, as well as the areas where weaknesses or concerns appear.

AIUC uses artificial intelligence agents to conduct the tests and analyze the data, but Rune Kvist said that human auditors verify the final audit result. The report is intended to provide the enterprise with independent information before deciding to purchase or deploy the agent, not to offer an absolute guarantee that the system is safe under all circumstances.

How Is This Different from METR’s Evaluations?

METR conducts similar testing in the areas of autonomy and evaluation, but until recently its work focused more heavily on agent performance and their ability to complete specific tasks reliably. AIUC, by contrast, is attempting to orient evaluation toward the risks that matter to enterprises in practical use, such as data leaks and undesirable behavior.

certi.news’s Perspective

The most important development here is not the funding alone, but the attempt to turn the safety of artificial intelligence agents into an external, comparable auditing process before purchase. This could benefit enterprises that cannot rely solely on model providers’ safety claims, but it leaves open questions about the extent to which the market will recognize the AIUC-1 standard, how test results will change as agents are updated, and whether the 5,000 tests represent actual operating environments. In addition, relying on artificial intelligence itself for part of the testing and analysis makes final human verification a crucial element of the service’s credibility.

AIUC’s two founders link this model to the difficulty of deploying more intelligent systems inside banks, hospitals, governments, and military organizations when institutions cannot guarantee the boundaries of the system’s behavior to their customers. This intersects with a call from Dario Amodei, CEO of Anthropic, to slow the development of advanced models and use independent evaluators to monitor and verify safety, with METR mentioned as one possible option.

News source
TechCrunch Startups
Open original source ↗
c
Author

certi.news Editorial Team

In the same category

You may also like

View all news