Rollbacks on Kubernetes: why a tag is not a version
A rollback that pulls the same tag again brings back the broken version. If you want to roll back reliably, you need the digest of every deployment, and a plan for the database.

Friday afternoon, a release goes out, the error rate climbs. Reaching for a rollback is a reflex: kubectl rollout undo. Kubernetes reports success, the pods restart, and the error rate stays put. The rollback worked; it just didn’t bring anything back. The reason is in one line of the manifest: image: registry/shop:main.
A tag is a pointer
An image tag such as main, latest or v2 is a mutable pointer. Every build that pushes the same tag moves it. kubectl rollout undo restores the previous pod template, and it contains the same tag. The nodes pull the image again, reliably with imagePullPolicy: Always or on a fresh node, and get the new, broken version. Kubernetes did exactly what the manifest said.
Only the digest is immutable, the checksum of the image manifest: registry/shop@sha256:4f1c…. A digest always points to the same image, no matter where a tag moves later.
Three rules for rollbacks that actually roll back
- Every deployment records the digest. Not the tag that was in the build, but the digest it pointed to at that moment. If you roll that out in the pod template, an undo really does roll out the old version.
- Every build also gets a fixed tag. For example
sha-<commit>. Themaintag can keep moving, but every version stays findable under a name nobody overwrites. - A restart doesn’t silently pull a new version. If you change environment variables and restart, you expect the same version with new configuration. If the tag now points somewhere else, a new build comes along unintentionally.
Helm does it differently, but not completely
helm rollback restores the chart version and values of an earlier revision, and Helm keeps the history in Secrets in the namespace. That’s a real step back for everything in the chart. But if the chart says tag: main, the same applies as above: the values come back, the image doesn’t necessarily.
What no rollback brings back
The most important section, because this is what goes wrong most often.
The database. If the new release ran a migration, deleted or renamed a column, the old version runs against a schema it doesn’t know. The application rollback becomes the second outage. The answer is two-step migrations: expand first (new column, the old one stays), roll out, and only clean up in the next release. Then every version can run against the next one’s schema.
Configuration and secrets. An image rollback leaves environment variables, secrets and domains as they are. That’s usually right, because a rotated key shouldn’t be rotated back. But it means: if you changed configuration along with the release, you have to revert it yourself.
Data on volumes. Whatever the new version wrote stays written. That’s what snapshots are for, not rollbacks.
A rollback in practice
Here’s what a good rollback looks like, in this order:
- Find the last healthy deployment in the history, with its digest.
- Check whether the faulty release included a migration. If so: is the old schema still compatible?
- Roll out exactly that digest, without a new build.
- Watch the instances’ error rate and logs until they’re stable.
- Fix the bug in the code and deploy normally again, instead of staying in the rollback state for weeks.
Tags, digests and build pipelines
You can get the digest in three places, and none of them takes much effort:
- At build time.
docker buildx build --metadata-file meta.jsonwrites the digest of the pushed image to a file. The pipeline passes it on to the deployment. - From the registry. Every registry that follows the Docker standard returns the digest for a tag in the
Docker-Content-Digestheader. Tools likecrane digest registry/shop:mainquery exactly that. - From the cluster. The status of a running pod contains, under
imageID, the digest the node actually pulled. That’s the most honest proof of what is running right now.
What matters is that the digest is stored where you’ll look for it when things go wrong: in the deploy history, next to the time, the commit and the person who rolled it out.
Migrations that survive a rollback
An example makes the pattern clear. A name column is to be split into first_name and last_name.
- Release A: expand. Create the new columns; the application writes to both old and new, and keeps reading from the old one. A rollback of A is harmless; the old version ignores the new columns.
- Backfill. Populate existing rows, in the background, in small batches.
- Release B: switch over. The application reads from the new columns and keeps writing to both. A rollback to A works because the old column is still being maintained.
- Release C: clean up. Only once B has run stably for a while is writing to the old column dropped, and a later migration deletes it.
That’s three releases instead of one. In return, every single step is reversible, and no rollback runs into a schema the old version doesn’t understand.
Rollback or fix forward?
Not every bug needs a rollback. The rule of thumb: if users are affected right now and the cause isn’t clear within minutes, go back to the last healthy version. If the bug is narrowly contained, the fix small and the pipeline fast, fixing forward can be the shorter path. What never works: spending half an hour hunting for the cause in production while users see errors. Stabilize first, then understand.
How Clusterward does it
Clusterward rolls out every deployment of an image from Scaleway Container Registry with the digest the tag had at that moment, and records it in the deploy history. New builds also get a fixed sha- tag. “Restore this version…” rolls out exactly the image of an earlier healthy deployment, without a build; for Helm services, the chart version and values. A restart keeps the running image instead of pulling a tag that has moved. Configuration, secrets, domains, volumes and the database deliberately stay as they are, and the dialog tells you so beforehand. More in Deployments; how databases and volumes come back in Backups & recovery. Whether the rollback worked is shown by the logs of all instances in Logs & monitoring.
Conclusion
A rollback is only as good as the version it points to. Roll out digests instead of tags, give every build a fixed tag and plan migrations so the old version can live with the new schema. Then Friday afternoon is one click, not a lost evening.
Want to make your rollbacks safe? Describe your build and your deployment. We’ll tell you where a rollback misses its target today. Ask a question →
Sources and further reading
Frequently asked questions
- Because the previous pod template contains the same mutable tag, such as image: registry/shop:main. When the nodes pull the image again, reliably with imagePullPolicy: Always or on a fresh node, they get the new, broken build. Kubernetes does exactly what the manifest says. Only the digest lets you roll back reliably.