backenddrills

Cloud and delivery

Cloud rollouts: readiness, draining and compatible rollback

By the BackendDrills editorial team · Published and checked October 6, 2026 · 7-minute read

Before you start: Service requests, background tasks and deployment replicas. Find this reading in a study path →

A healthy process, a service ready to accept traffic and a worker safe to terminate are different conditions. Define all three before relying on a rolling deployment. This guide uses Kubernetes lifecycle concepts; check the version and traffic infrastructure you actually deploy.

Four contracts to draw

ContractQuestion
StartupHas initialization progressed far enough for regular checks?
LivenessIs restarting this container an appropriate response?
ReadinessShould this Pod receive service traffic now?
TerminationCan admitted work finish or be recovered before the deadline?

A failed readiness check makes a Pod unavailable as a ready service endpoint; it does not itself restart the container. Liveness failures can cause a restart. A startup probe can protect initialization from premature liveness and readiness evaluation. Choose checks that correspond to those responsibilities rather than making every endpoint run the same expensive dependency query.

A deployment with in-flight work

In a synthetic upload service, a request has already stored a file and is writing its metadata when termination begins. Immediately killing the process can leave an uncertain outcome. The application needs to stop admitting new work, bound the completion of admitted work and expose enough identity to reconcile unfinished effects. Kubernetes allows a termination grace period and lifecycle hooks, but the application's actual shutdown behavior still matters.

Write a timeline for endpoint withdrawal, signal handling, connection draining and the final process exit. These changes are not a universally atomic sequence. Account for your ingress or load balancer, long-lived connections and background consumers. Avoid a fixed sleep that consumes most of the available grace period while leaving the application unable to drain.

Exercise the timeline

  1. Start a long synthetic request and record its operation identity.
  2. Begin a controlled rollout while a second client sends requests.
  3. Record readiness changes, admitted work, completion and termination time.
  4. Assert one correct effect for each accepted operation, including retried ones.
  5. Repeat with a delayed dependency and a bounded drain deadline.

Rollback also has a data contract

A previous binary may not understand a new schema or event shape. For a practice migration, keep old and new readers running against the transitional state and demonstrate the compatibility boundary. Decide when removing the old representation becomes irreversible. Finish with a rollout checklist that names traffic admission, admitted-work recovery, schema compatibility and the evidence required to expand the deployment.

Check the underlying behavior

Original illustrative examples, prepared with AI assistance and checked against the linked primary documentation. No customer incident or vendor endorsement is claimed. Our editorial approach.

Practice the next decision

Try a complete free backend drill: inspect evidence, make three decisions and review the reasoning. No account or card required.

Try a free incident drill →

Explore 1008 scenarios · $28 one-time

Selected practice from the study paths

Free readings need no account. Full edition drills require verified access; opening a paid link does not expose its answers.

Recommended next readings