Artificial intelligence

Anthropic Proposes Public Metrics to Measure the Pace of AI Progress Inside Advanced Laboratories

On September 17, 2026, Anthropic published an experimental framework for measuring how dependent AI research is on Claude, mechanisms for monitoring agents, and the share of computing allocated to safety research. The figures reveal that Claude leads 26% of the company’s research and development work, while it did not operate with full autonomy in any measured area through August 2026. The company emphasized that the measurements require standardized methodologies and independent verification.

2026-09-18
6 min read
12 views
فريق تحرير certi.news
Anthropic Proposes Public Metrics to Measure the Pace of AI Progress Inside Advanced Laboratories

On September 17, 2026, Anthropic published a proposal to measure the pace of AI development inside advanced laboratories, starting from a gap that the company says the public and governments currently cannot observe: what is actually happening inside these laboratories, and to what extent models have begun helping build subsequent generations of models? The report does not announce a new model or product; rather, it presents three sets of experimental metrics along with a snapshot of Anthropic’s internal operations in 2026.

The metrics focus on the extent to which research and development depends on AI, the company’s ability to oversee agents carrying out semiautonomous tasks, and how computing resources are distributed between capability development and safety research. Anthropic believes these indicators complement capability evaluations and risk reports published under the Responsible Scaling Policy, and could later help establish transparency commitments that can be compared across laboratories.

How Much Does AI Development Depend on Claude?

Anthropic created an initial indicator called the Anthropic R&D Automation Index by cataloging model-related research and development work, then assessing the degree of automation for each task and calculating its weight according to the human time allocated to it. The assessment uses a scale developed by Epoch AI, ranging from AL0, where AI does not participate, to AL5, where the system operates with full autonomy and without human intervention.

As of August 2026, Claude was not operating with full autonomy in any measured group of research and development work. However, under the AL4 definition, it was leading approximately 26% of this work; that is, it could complete most of the task end to end from general guidance, with human supervision. The share of work in which AI participated at AL3 or higher—where the system handles substantial portions of the task under close human direction—also exceeded 90%.

To produce the index, the company built a map of approximately 15,000 detailed tasks based on work logs and internal documents from July 2026, then organized them into a tree containing 542 nodes, including 378 leaf nodes. According to the methodological appendix, the evaluation model’s judgments exactly matched human judgments in 59% of cases, compared with a human-to-human agreement rate of 35%; the judgments were within one level of one another in 97% of cases. However, Anthropic acknowledges that the boundaries between “collaborates” and “leads” remain open to disagreement.

Monitoring Approximately 30,000 Agents

The company says that approximately 30,000 agents were carrying out research and engineering work at any given time on its most widely used internal platform during August 2026. The system includes real-time monitoring before an action is executed and retrospective monitoring to analyze behavior after it occurs. One hundred percent of these agents’ actions passed through the real-time monitor before execution, usually within seconds, and 100% of the actions were also fed into retrospective monitoring.

Anthropic analyzed more than one billion decisions made by research and engineering agents during August and found that the real-time monitor blocked 0.002% of decisions, or approximately one decision in every 47,000. Retrospective monitoring flagged approximately one to two out of every thousand conversations or activities for additional review. A system of classifiers handles most of the flags, while approximately 50 cases per week were escalated to human reviewers. The company says that human review of blocked actions takes place within a week.

These figures do not mean that risks have disappeared; they measure a specific internal platform over a limited monitoring period, and detecting rare behavior becomes more important as the number of agents increases. Anthropic acknowledges that the monitoring experiment has run for months rather than years and that it does not guarantee a complete picture of unknown behaviors.

Computing Allocated to Safety

Anthropic examined computing use between July 13 and 20, 2026, classified workloads into categories, and then estimated the share of safety research in the computing used for research and development. During that period, approximately 6% of research and development computing went to safety work, rising to approximately 12% when considering computing allocated to AI-led research and development.

The company describes these estimates as conservative because it did not count work that served capabilities and safety to an equal degree, and it did not include the computing used for safeguard classifiers. It emphasizes that the share is not a complete measure of the effort devoted to safety, because research in this area may depend more on researchers’ time than on running computationally intensive experiments. The measurement also reflects a single week and does not establish a time trend.

Why Does This Framework Matter?

The proposal’s primary value lies not only in the current figures, but in its attempt to transform AI progress from a subject visible only to laboratories into a process that can be tracked through comparable indicators. If laboratories publish rates of research automation, monitoring coverage, review times, escalation rates, and the share of computing directed toward safety, it may become possible to compare change over time and perhaps across companies.

Anthropic itself identifies fundamental limitations: there is not yet a shared methodology, some assessments rely on the company’s own models, and the “judge” may repeat the errors of the system it is reviewing. In addition, classifying safety work depends on definitions whose boundaries may be drawn differently by different laboratories or regulatory bodies, while some computing labels are based on automated data or user entries that are not fully documented.

The company says it plans to provide comparative access to multiple independent evaluators so they can verify safety practices and monitor incidents and indicators. Therefore, what has actually changed is the availability of a prototype for transparency and internal data open to discussion, not the establishment of a binding global standard. The idea’s success will depend on standardizing definitions, enabling external verification, and clarifying what can be published without exposing sensitive competitive data.

News source
Anthropic Newsroom
Open original source ↗
ف
Author

فريق تحرير certi.news

In the same category

You may also like

View all news