Clusterward
← Back to the blog
ArchitectureUpdated Florian Apel

Helm charts and their disks: what is left after helm uninstall

A Helm chart often brings its own disks, for example for the database bundled in the chart. Some disappear on uninstall, data and all, others stay around forever. You want to know which before you start.

Cover image: Helm charts and their disks: what is left after helm uninstall

A WordPress chart, a chart for Matomo, one for a search index: many Helm charts bring their own disks. The values say persistence.enabled: true plus a size, and after installation the namespace contains one or more PersistentVolumeClaims, PVCs for short. What happens to them when the chart is upgraded, uninstalled or reinstalled is surprisingly inconsistent. And it decides whether data is lost or keeps costing money forever as a forgotten disk.

Two ways a chart creates disks

As its own object in the chart. The chart contains a pvc.yaml template. Helm creates the claim, and it belongs to the release just like a Deployment or a Service. Helm marks it with the annotations meta.helm.sh/release-name and meta.helm.sh/release-namespace.

Via a StatefulSet. The chart contains a StatefulSet with volumeClaimTemplates, typical for databases such as MariaDB or PostgreSQL bundled in the chart. Here, it isn’t Helm that creates the claim but Kubernetes’ StatefulSet controller, one per replica: data-wordpress-mariadb-0, data-wordpress-mariadb-1. Helm doesn’t know these claims as its own objects.

The difference sounds academic. It isn’t.

What helm uninstall deletes

Disk

On helm uninstall

Consequence

PVC from a chart template

is deleted

With reclaim policy Delete, the disk and its data are gone

PVC with helm.sh/resource-policy: keep

stays

Data stays, nobody cleans it up

PVC from a StatefulSet’s volumeClaimTemplates

stays

Data stays, Helm no longer knows about it

The first row is the dangerous one. The usual storage classes have the reclaim policy Delete: when the claim disappears, the cloud provider deletes the disk behind it. A helm uninstall to clean up a test environment that accidentally pointed at the production data has happened exactly like this. That’s why some charts set resource-policy: keep on their claim, and some don’t. A look at the template before the first installation pays off.

The third row is the expensive one. StatefulSet claims survive every uninstall. Newer Kubernetes versions do support persistentVolumeClaimRetentionPolicy on the StatefulSet, which can remove claims on deletion, but most charts don’t set it. After a year of test installations, the cluster is full of disks nobody can attribute anymore.

An example: WordPress with a built-in database

A popular WordPress chart with MariaDB enabled creates two claims after installation: one for uploads and plugins, as its own template in the chart, and one for the database, from MariaDB’s StatefulSet. Remove the release with helm uninstall and you get both at once: the upload claim disappears, unless the chart marks it with keep, and all the images go with it. The database claim stays, along with all the posts. Reinstall afterwards and the old posts are back, but without images, and the newly generated database password no longer matches the old database. The site shows a connection error, and it takes a while to figure out why. How to set up a WordPress site without this trap is shown in WordPress from the app catalog; more traps are covered in WordPress on Kubernetes: seven traps between the chart and the first page.

Everyday pitfalls

Resizing. In a StatefulSet, the volumeClaimTemplates are immutable. Increase the database disk size in the values and the next helm upgrade fails with an error. The way forward is through the claim itself: increase the size on the PVC, provided the storage class allows expansion, and have the StatefulSet recreated with --cascade=orphan so the pods aren’t deleted along with it.

Reinstalling. When a release is reinstalled under the same name, the StatefulSet finds its old claims and mounts them again. That’s handy if you want it and surprising if you expected an empty installation. A password that happens to be regenerated in the new release then no longer matches the old database’s data.

Restoring. Restoring a claim from a snapshot creates it anew. If it lacks the labels and annotations Helm had marked it with, the next helm upgrade fails with a message saying the object doesn’t belong to the release. A restored claim therefore has to be created with exactly this metadata.

Snapshots of chart disks

Block Storage snapshots work for chart disks just as they do for any other. Two things are worth knowing:

  • A snapshot of a running disk is crash-consistent. It corresponds to the state after a power failure. Most databases cope with that, but it isn’t guaranteed. For important data, add a snapshot taken with the application stopped, or a logical dump.
  • Before restoring, everything that uses the disk has to be stopped. A disk is attached to exactly one node. As long as a pod of the StatefulSet has it mounted, it can’t be replaced.

The better question: does the database belong in the chart?

For a small application, the bundled database is convenient. For anything customers use, there’s a lot to be said for a managed database outside the cluster: automatic backups, point-in-time recovery, hands-off updates, and no StatefulSet that has to move when a node pool is upgraded. Most charts let you switch off the built-in database and specify an external one. What that looks like with one database per service is described in Managed databases.

A checklist for charts with data

  1. Check in the chart which claims it creates and whether they carry resource-policy: keep.
  2. Know the reclaim policy of the storage class in use.
  3. Set up scheduled snapshots before real data exists, not after.
  4. Practice a restore once, ideally in a test environment.
  5. After every uninstall, check which claims are left and consciously delete or keep them.
  6. Consider a managed database for databases holding customer data.

How Clusterward handles chart disks

Clusterward finds a Helm service’s disks live in the cluster: claims Helm has assigned to the release, and claims a pod of the release has mounted, including those from StatefulSets. The Volumes card shows them with mount path and size, marked as “from the chart”. Snapshots run immediately or on a schedule per disk and show up as ready about a minute after completion. When restoring, Clusterward stops every Deployment and StatefulSet that uses the disk, backs up the current state, recreates the claim with Helm’s metadata and starts everything again, so the release still recognizes it as its own. If a disk is left behind after an undeploy, you can delete it in the cockpit, with a final snapshot beforehand. The size is still determined by the chart. More in Volumes & snapshots and Backups & recovery; how Helm services are rolled out in Deployments.

Conclusion

A Helm chart with data is two things: the application Helm manages, and the disks an uninstall either deletes or leaves lying around forever. Before the first installation, check which kind your chart creates, back up the disks from the start and practice a restore once. And for customer data, ask yourself whether the database belongs in the chart at all.

Running Helm charts with data? Name your charts. We’ll tell you which disks they create and how to back them up. Ask a question →

Sources and further reading

Frequently asked questions

  • That depends on how the chart creates them. A PVC from a chart template is deleted, and with reclaim policy Delete the disk and its data go with it. A PVC with helm.sh/resource-policy: keep stays, as do PVCs from a StatefulSet’s volumeClaimTemplates. Someone has to clean those up deliberately later.