Cloud Computing and Data Centers

How Atlassian Reduced Incident Detection Time from More Than 40 Seconds to Less Than 10

Atlassian describes rebuilding its incident-detection platform using OpenTelemetry, Apache Kafka, and Apache Flink on Kubernetes, reducing the time required to turn events into metrics to less than 10 seconds and lowering operating costs by approximately 97%. However, the improvement did not eliminate coverage and false-alarm problems, nor the ingestion pipeline’s dependence on a single region.

2026-09-30
6 min read
15 views
certi.news Editorial Team
How Atlassian Reduced Incident Detection Time from More Than 40 Seconds to Less Than 10

Atlassian rebuilt the incident-detection platform that feeds its automated incident-creation system, relying on Apache Kafka, Apache Flink on Kubernetes, and OpenTelemetry. According to measurements published after 18 months of work, the time for an event to reach a metric fell from more than 40 seconds to less than 10 seconds, while sustained capacity increased from approximately 500 million events per day to more than 1 billion events per day when using 50% sampling.

The project does not present itself as a complete success story. The recall rate for incidents within the monitored scope rose from approximately 60% to a peak of 86% in June 2026, before falling to 64% in August. Precision also remained below the target level, prompting the team to distinguish between the quality of the detector itself and the extent of coverage of the products and experiments instrumented with measurements.

A Platform Built Around Event Streams

Atlassian operates more than ten cloud products for millions of tenants, and these products generate billions of events daily from user interactions. The events include the start and result of a task, along with data about the tenant, user, experiment, and HTTP code. The platform uses this data to answer three questions quickly: Is there a problem? How large is its impact? And which team should be notified, and at what severity?

In the new design, events are filtered at the Apache Kafka bus instead of consuming the entire stream inside the application. The subscription filter, consisting of approximately 770 lines of YAML, is maintained as part of the software configuration. A single Apache Flink 1.20 application then processes the events on Kubernetes, enriching them with tenant data, sending metrics through OpenTelemetry, and aggregating incident impact in 60-second windows.

The team used Apache Parquet to store aggregates and a multi-region key-value store, while the Impact API provides answers about the number of affected users and tenants. The AutoHOT engine converts alerts into incidents, applies a severity matrix, suppresses transient alerts, and reevaluates impact every minute.

Performance and Cost Gains

  • The number of virtual machines fell from approximately 90 machines, in addition to queueing and caching layers, to four Kubernetes containers.
  • Monthly operating costs declined from approximately $20,000 to about $650, a reduction of nearly 97%.
  • Recovery from a processing outage became possible by replaying Kafka events within approximately 20 minutes without data loss.
  • Impact-dashboard query time fell from approximately 10 seconds to about one second.
  • Agreement with the old path reached 99.9% during a two-week parallel run.

The design used HyperLogLog to count unique affected users instead of storing user IDs, with an error of approximately 1.5%. Idempotent storage keys were also used to make Kafka replay recoverable without duplicating rows. The team retained the old metric names through a StatsD path, allowing existing SLO dashboards and detectors to continue operating without changes during the transition.

Why Performance Numbers Are Not Enough

Incident results show that accelerating the data pipeline does not necessarily mean better coverage. Over nine months, 263 major incidents occurred, but only 117 incidents, or 44.5%, affected instrumented experiments. Eighty of those incidents were detected, meaning that the system captured only 30.4% of all major incidents during that period.

Precision also varied by alert type. The precision of Sev2 incidents created automatically by the system was approximately 85% during the fiscal year, while the precision of lower-severity early warnings ranged between 70% and 79%. Suppression of fluctuations helped reduce rejected transient-alert tickets by approximately 80%. Conversely, delays in the metrics system caused missing data to be treated as zero, generating a wave of false alerts; this was addressed by adding a minimum delay of 120 seconds to volume-drop detectors.

Remaining Open Limitations

The largest logical problem is that silence can look like health. A completely stopped database may prevent a page from loading and therefore produce no failure events at all. A failure of the event pipeline itself can also blind the detector. The team therefore added data-volume-drop detectors and measurement-freshness checks, and plans to incorporate independent signals such as edge 5xx errors and synthetic probes.

The team also found that detecting an incident and measuring its impact are separate problems. In one incident, the ticket was created early, but the system estimated the impact at approximately 2,000 users, while the actual number exceeded 80,000. The Impact API also collapsed under the load of approximately 100 concurrent users because it merged HyperLogLog sketches at read time without pre-aggregation or caching.

The most important weakness remains that the event gateway, Kafka subscription, and Flink job operate in a single region, even though the decision engine runs in an active-active mode across two regions. This means that a regional failure could leave computation healthy while disabling the source of detection itself.

Editorial View from certi.news

The core value of Atlassian’s experience lies not in choosing Flink or Kafka individually, but in connecting architectural decisions to operational metrics that can be reviewed: filtering at the bus, idempotent keys, monitoring the platform with OpenTelemetry tools themselves, and distinguishing recall from coverage. The results show that improving efficiency can clearly reduce cost and latency, but it does not automatically address measurement gaps or incidents that produce no client signals.

Atlassian plans to move detectors to a time-series store fed through OpenTelemetry, expand the Flink pipeline to an active-active model across regions, and add daily aggregation and caching in front of the Impact API. It also targets recall and precision of more than 90% each within the instrumented scope, with a P90 time to detect Sev3 incidents of less than 90 minutes. These remain future goals rather than achieved results according to the published material.

News source
c
Author

certi.news Editorial Team

In the same category

You may also like

View all news