Clusterward
← Back to the blog
OperationsFlorian Apel

Kubernetes in production: the checklist most architecture diagrams leave out

Popular Kubernetes production architecture diagrams usually stop at the registry. This checklist starts there: rollback, secrets, TLS, resources, backups with tested restores, security, alerts, upgrades and the question of who owns the cluster.

Cover image: Kubernetes in production: the checklist most architecture diagrams leave out

Friday, 5:40 pm. The lunchtime release breaks the checkout, and the rollback fails on one question: which image ran before? The tag latest already points at the broken build. The team had built its cluster from one of those diagrams shared as the “complete Kubernetes production architecture” – Git, CI/CD, registry, cluster, ingress, pods, a Grafana logo in the corner. Everything on it was in place. What was missing never appears on it.

A Kubernetes production checklist is a list of operational questions that must be answered – and tried once – before real customers land on the cluster: how you roll back, where secrets live, who gets the alert at night. It does not list components; it lists states you can check.

The typical diagram shows the path of the code: commit, build, push, deployment, ingress – the part that works on day one. What hurts in production comes later: on day 60, when a certificate is not renewed; in month ten, when the Kubernetes version leaves support; on the evening someone drops a table. The list below is grouped by area, each item with its reason. Only tick what you have tried.

Deploys, configuration, networking: what must be in place before the first customer?

Deploys and rollback

  • Roll out images by digest (image@sha256:…), not by tag – a tag can be moved, a digest cannot.
  • The previous version comes back in one step, without a new build – rehearsed at least once.
  • Database migrations stay backwards compatible: old code must cope with the new schema for a while, or the fastest rollback is useless.
  • A readiness check and a timeout per service: a slow start must not count as a failed deploy.

Configuration and secrets

  • No secrets as Base64 in the Git repository – Base64 is an encoding, not encryption.
  • By default Kubernetes stores Secrets unencrypted in etcd; encryption at rest and RBAC on Secrets are part of the basics.
  • Changed configuration visibly reads “saved, not rolled out yet” – otherwise a change counts as done that never reached the pods.

Networking and TLS

  • By default cert-manager renews two thirds of the way through a certificate’s lifetime. Monitor whether renewal works, not just whether it is configured.
  • Pass the real client IP through with the proxy protocol – otherwise rate limits, allowlists and logs only see the load balancer’s address.
  • Set the body size limit deliberately: NGINX allows 1 MB per request by default (client_max_body_size); larger uploads get a 413.

Resources and data: how much headroom, which backups?

Resources and scaling

  • Requests for every container: the scheduler places pods by requests, not actual usage.
  • Memory limits with room above normal usage: exceeding one risks an OOM kill; a CPU limit only throttles.
  • Autoscaling needs a CPU request – the HPA works in percent of it. Anything that cannot afford an outage runs at least two replicas.
  • Headroom in the node pool: when a node fails, its pods must fit on the rest.

Data

  • Database backups and regularly tested restores – a never-restored backup is a hope.
  • Restore into a new database, not over the running one, so the pre-restore state survives.
  • Volume snapshots on a schedule and before every volume is deleted.
  • Bucket backups in another project and region; versioning in the same bucket is not a backup.

Retention, restore tests and evidence in practice: Backups & recovery.

Security: who can reach cluster and data?

  • RBAC with proper roles instead of one admin kubeconfig on three laptops.
  • Nodes without a public IP on a private network, outbound traffic through a gateway.
  • The API server is reachable only from allowed addresses.
  • Network policies: without them every pod accepts connections from every other – and they only work with a network plugin that enforces them.
  • Least privilege for machines too: CI tokens and operators with exactly the rights they need and an expiry date.

Roles, tokens and network access in detail: Security & access.

Observability and upgrades: who notices first?

Observability

  • Logs beyond a pod’s life: once it is replaced, its log is gone. A central store such as Loki or OpenSearch keeps it.
  • External uptime checks on the real hostnames – “pod running” does not mean “site reachable”.
  • Alerts that reach a person: in the team chat or by email, once per event. A dashboard nobody has open alerts nobody.

Upgrades

  • Kubernetes maintains the three most recent minor versions, each for roughly a year. Version 1.34 reaches end of life on 27 October 2026 (as of October 2026).
  • Plan minor upgrades calmly, one version at a time, watching for removed APIs – before the provider pushes.
  • Add-ons have lifecycles of their own: Ingress NGINX has had no releases and no security fixes since March 2026.

How a Kapsule cluster stays current: Cluster updates.

Who owns the cluster at 2 a.m.?

The items above are technology. The last area is organisation – and the most common reason a well-built cluster goes feral within a year.

  • At least two people can deploy, roll back and start a restore.
  • Runbooks for the five most common incidents exist before they happen.
  • An audit log shows who changed what and when – CI tokens included.
  • Upgrades and certificates have a named owner, holidays included.

Where Clusterward takes items off the list

On Scaleway, Clusterward ticks off several items for you. New clusters get nodes without a public IP and an API server that accepts only allowed addresses; the proxy protocol is on by default. Images from the cluster’s registry roll out by digest, and any earlier healthy version can be restored (Deployments). Secrets are stored encrypted or in Scaleway Secret Manager. Database backups restore into a new database, volumes are snapshotted on a schedule, buckets are backed up nightly and restore-tested monthly. Uptime checks and wildcard certificate warnings reach a person by webhook or email, and the next Kubernetes minor upgrade runs once you confirm it.

Not covered: logs beyond a pod’s life, canary deployments and a Prometheus stack. And no software writes your runbooks for you.

Conclusion

The architecture diagram shows how code gets into the cluster. The checklist asks what happens when something breaks afterwards. Go through it once with the team and mark every item as “tested”, “set up, never tested” or “missing”. The middle column is usually the longest – and the most dangerous.

Gaps in your checklist? Tell us how your cluster is run today. We’ll tell you which items to tackle first. Ask a question →

Sources and further reading

Frequently asked questions

  • Nine areas: deploys and rollback, configuration and secrets, networking and TLS, resources and scaling, data with tested restores, security, observability, upgrades and ownership. Every item should not only be set up but tried once – a rollback, a restore, an alert that really reaches a person.