Programming and Software Development

Coding Agents Have Made CI a Bottleneck; Speeding Up Build Pipelines Is Not the Complete Solution

The article argues that the rising productivity of coding agents is putting pressure on continuous integration, but the problem is not limited to slow CI pipelines. Testing the repository alone does not reveal failures in interactions between distributed services, making it necessary to move system-level validation into the agent’s work loop.

2026-10-04
5 min read
5 views
certi.news Editorial Team
Coding Agents Have Made CI a Bottleneck; Speeding Up Build Pipelines Is Not the Complete Solution

Continuous integration (CI) has become a new bottleneck as the use of coding agents expands. This conclusion is based on experiences presented by engineering teams at Anthropic, Linear, and Depot during September, not on the announcement of a single tool. Anthropic’s CI job volume increased 25-fold over six months, while its engineers began shipping roughly eight times as much code per quarter as they had shipped between 2021 and 2025. At Linear, the size of the test suite approached four times its level since January, while agents began writing most of the tests.

Anthropic responded by using test impact analysis to run only the tests likely to be affected by a change, while Linear redesigned almost its entire pipeline. These measures reduce waiting time, but they address only one layer of the problem: the speed of repository checks.

Why Is CI Speed No Longer Enough?

CI was historically designed around a human workflow: a developer opens a limited number of merge requests each week, then waits for the pipeline result while moving on to another task. An agent, however, can create code, tests, and merge requests in parallel, multiplying the number of CI operations. The article notes that Blacksmith, a company that sells CI runners, observed weekly growth of between 5% and 10% in the number of jobs it runs.

The second problem is where validation occurs. The agent writes the change and then waits for the result after creating the merge request. When the result arrives 20 minutes later, it may have lost the task context, while every problem means another waiting cycle.

The Repository Is Not the System

In standalone applications, repository testing may provide a close picture of system behavior. But in a cloud-native environment, the repository represents one service within an ecosystem that may contain dozens of services, while the rest of the system is often represented by mocks or test data.

For this reason, a change can pass unit tests, pass CI, and work inside an isolated environment, then fail on the first real request that crosses service boundaries. Examples include changing the name of a field that another service depends on, reducing a timeout in a way that causes a chain of retries, modifying a schema that causes a table lock in the test environment, or an endpoint behaving differently when it is actually called by the consuming service.

The article cites data from DevOps Research and Assessment (DORA), which links increased AI adoption with both higher software delivery frequency and increased delivery instability. In other words, producing code faster does not guarantee better validation.

What Changes in Practice?

The proposed solution is not to eliminate CI or slow agents down, but to move part of validation earlier into the agent’s work loop, so that the change is tested against the actual system rather than the repository copy alone. Tools such as Cursor use isolated cloud environments; more than 30% of the merge requests merged by Cursor come from agents operating this way. Other tools, including GitHub Copilot cloud agent, Codex, Devin, and Greptile, also provide various forms of code execution inside temporary environments.

However, these environments usually contain the branch, its configuration, and whatever the setup script can install—not the other services, the real message queue, or a database resembling production data. The article therefore argues that the loop is closed, but it may be closed around the wrong thing.

Shared Environments and Governed Validation

The article proposes running a stable shared version of the services inside a Kubernetes cluster, while creating lightweight test environments that deploy only the modified service. Tagged requests are routed to that service, while the rest of the paths connect to the shared stable versions. This allows many agents to share the environment instead of copying the entire system for each agent; the article estimates that the environment’s cost may approach the cost of a single container and that starting it takes seconds, but the source does not provide independent measurements proving these estimates across all environments.

The environment alone is not enough. Platform teams need approved procedures defining which requests are sent, which logs are collected, and which contracts must be verified. The services touched by the test and its results should also be recorded, so that the record can be read by review tools and merge gates. The author emphasizes that governance is necessary to prevent agents from carrying out unsafe operations inside a shared cluster.

From certi.news’s perspective, the real change is not merely speeding up CI, but redefining what validation “success” should mean. Repository testing remains important, but it is not sufficient on its own for distributed systems produced by agents at higher speed. Questions concerning cost, isolation, data security, and the measurement of validation accuracy remain open, and the material presents an analytical thesis rather than a proven standard or a specific product.

News source
The New Stack - Software Development
Open original source ↗
c
Author

certi.news Editorial Team

What you need to know

أدى توسع استخدام وكلاء البرمجة إلى زيادة ضغط وظائف التكامل المستمر، لكن تسريع خطوط CI لا يحل مشكلة التحقق من تفاعلات الخدمات الموزعة. يقترح المقال نقل جزء من التحقق إلى حلقة عمل الوكيل، باستخدام بيئات مشتركة ومحكومة تختبر التغيير مقابل النظام الفعلي.

  • ارتفع حجم وظائف CI لدى Anthropic 25 مرة خلال ستة أشهر، كما زاد حجم شحن الشيفرة لديها بنحو ثمانية أضعاف مقارنة بالفترة بين 2021 و2025.
  • اقترب حجم مجموعة اختبارات Linear من أربعة أمثال مستواه منذ يناير، بينما أصبح الوكلاء يكتبون معظم الاختبارات.
  • تحليل تأثير الاختبار وإعادة تصميم خطوط CI يقللان زمن الانتظار، لكنهما يعالجان سرعة فحص المستودع ولا يختبران النظام الموزع كاملاً.
  • قد تمر تغييرات أسماء الحقول والمهلات والمخططات واختبارات نقاط النهاية داخل المستودع، ثم تفشل عند تفاعل الخدمات فعلياً.
  • يقترح المقال بيئة Kubernetes مشتركة تضم نسخاً مستقرة من الخدمات، مع نشر الخدمة المعدلة فقط وتوجيه الطلبات الموسومة إليها.
  • تحتاج هذه البيئات إلى حوكمة تحدد الطلبات والسجلات والعقود المسموح بها، لمنع الوكلاء من تنفيذ إجراءات غير آمنة.

FAQ

لماذا لا يكفي تسريع CI؟

لأن اختبارات المستودع قد لا تكشف أعطال التفاعل بين الخدمات الموزعة، حتى إذا نجحت اختبارات الوحدة ومر التغيير عبر CI.

ما الحل العملي المقترح؟

نقل جزء من التحقق إلى وقت أبكر داخل حلقة عمل الوكيل، واختبار التغيير داخل بيئة تتصل بالخدمات الفعلية أو بنسخ مستقرة مشتركة منها.

ما دور بيئات Kubernetes المشتركة؟

تسمح لعدد كبير من الوكلاء بمشاركة نسخة مستقرة من الخدمات، مع نشر الخدمة المعدلة فقط وتوجيه الطلبات التجريبية إليها.

هل تثبت المقالة أن تقديرات تكلفة هذه البيئات دقيقة؟

لا؛ يذكر المقال تقديرات عن اقتراب التكلفة من تكلفة حاوية واحدة وسرعة التشغيل، لكنه يوضح أن المصدر لا يقدم قياسات مستقلة تثبتها في جميع البيئات.

In the same category

You may also like

View all news