Cloud Computing and Data Centers

Practical Tests Reveal Why a Kubernetes Backup Does Not Mean Successful Recovery

The CNCF presents three reproducible scenarios that illustrate the difference between having backups and actually being able to recover stateful Kubernetes applications. The guidance emphasizes that testing data, separating declared state from stored state, and coordinating multi-volume snapshots are essential elements of any recovery plan.

2026-09-10
6 min read
108 views
certi.news Editorial Team
Practical Tests Reveal Why a Kubernetes Backup Does Not Mean Successful Recovery

On September 10, 2026, the CNCF published guidance based on three reproducible failure scenarios to test disaster recovery for stateful Kubernetes applications, rather than merely checking that a backup ended with a Completed status. The experiments, which can be run on a laptop from the lab repository, used a PostgreSQL application containing known data consisting of four rows, allowing the actual recovery result to be verified instead of relying on general status indicators.

The material was prepared by Saiyam Pathak and Saloni Narang, both CNCF Ambassadors. Its scope is limited to recovering application state within Kubernetes; it does not address compliance frameworks, product comparisons, or recovery of the underlying cloud or data-center infrastructure. Tools such as Velero and CSI Snapshot APIs appeared as reference implementations for the scenarios, while the failure patterns apply to tools performing the same roles.

A Completed Backup Does Not Prove Recoverability

In the first scenario, the Kubernetes backup components were separated into YAML resource definitions and persistent-volume data. The lab used Velero with a data-movement mechanism to an S3-compatible object store outside both clusters. Inspecting the DataUpload objects showed that 47,989,888 bytes of volume data had actually been transferred to the external store.

After the namespace, including the PVC, was deleted, the same four rows were restored in about two minutes. However, the CNCF warns that protecting volume data does not automatically make a database backup application-consistent; applications may require flush or quiesce procedures. Recovery on different infrastructure may also require StorageClass alignment and other transformations that the team must design and test.

Most importantly, backup tools restore resources to an existing cluster; they do not create nodes, networking, load balancers, or DNS. Therefore, the recovery plan must clearly identify the environment that will receive the backup, with Kubernetes recovery assigned to infrastructure as code or to Cluster API when necessary.

Git Restores Intent, Not Stored State

In the second scenario, the production cluster was shut down, while the recovery cluster already existed and contained a GitOps controller connected to a Git repository and a backup tool connected to the shared store. Application synchronization succeeded, the StatefulSet was running, and the dashboard appeared healthy, but querying the database returned the error: relation "attendees" does not exist.

The cause was not a Kubernetes or GitOps failure. The repository contained only the definitions, so the controller recreated the StatefulSet, Service, and a new empty volume. In practice, Git stores declared state, or the team’s intent, while backups store the actual data; neither can restore the complete application on its own.

The lab’s recovery method involved removing the empty application created by synchronization, then restoring the application with its volumes from the backup store, and finally comparing the data with the expected content. The journey from shutting down production to the appearance of verified data took four minutes in the live experiment and just under two minutes in the rerun. The CNCF notes that these figures cover only the scripted portion and exclude incident detection, decision-making, traffic switching, and returning to the original environment.

Individual Snapshots May Produce a Recovery Point That Never Existed

The third scenario tested an application using two interdependent volumes: one for orders and another for payments. The lab wrote matching pairs at a rate of five times per second, with the condition that every payment correspond to an order. When two separate snapshots were taken five seconds apart, each snapshot appeared ready and healthy on its own, but recovery revealed that the latest recorded order was 108352, compared with 108377 payments—25 payments without matching orders.

The result shows that the success of each individual storage operation does not guarantee application consistency across multiple volumes; together, the two snapshots may describe a point in time that never actually existed. The material notes that the gap may grow in production when the backup tool processes a large number of PVCs one after another.

VolumeGroupSnapshot, which reached GA status in Kubernetes 1.36, provides a mechanism for identifying volumes with a single label and requesting a consistent recovery point through CSI. In the coordinated experiment, the latest order and payment numbers matched at 109169, and validation of the relationship between them succeeded. However, support depends on the driver; support for ordinary VolumeSnapshots does not prove support for group snapshots, and most major cloud drivers examined for the lab were not implementing them as of mid-2026. CRDs, the snapshot feature, and the related add-ons must also be enabled explicitly.

What Should a Recovery Test Measure?

  • Restore a complete application to a clean target on which it has never run before, rather than merely deleting a Pod and watching it be recreated.
  • Verify the data and the user-access path using expected content, rather than relying only on resource status or dashboard colors.
  • Measure the entire process with a clock, recognizing that actual recovery time includes detection, decision-making, traffic switching, and possibly returning to the original environment.
  • Use two independent failure domains, such as a production cluster and a recovery cluster, with the backup store located outside both.

Editorial Reading: A Gap Between Tools and Recovery Sequencing

The practical change highlighted by these experiments is shifting the success criterion from “the backup completed” to “the correct data returned and is accessible.” The broader limitation is that core Kubernetes does not define a shared contract for coordinating data, applications, clusters, traffic, and identity, nor does it provide a standard durable resource describing the complete recovery unit for an application and its external dependencies. Therefore, the boundaries between GitOps, backup, infrastructure, and the user path remain the team’s responsibility even when mature tools are available for each layer. The CNCF states that the Cloud Native Business Continuity initiative of the CNCF TAG Operational Resilience seeks to collect contributions on ecosystem gap analysis, recovery guidance, and reference architectures.

News source
c
Author

certi.news Editorial Team

In the same category

You may also like

View all news