Clusterward
← Back to the blog
SecurityUpdated Florian Apel

Customer offboarding: why every drop needs a dump first

Deleting is easy, restoring is impossible. Offboarding needs a fixed order: back up, verify, then tear down. And a retention rule that someone has written down.

Cover image: Customer offboarding: why every drop needs a dump first

Onboarding a customer gets attention: it’s in the sales pitch, the demo script, the roadmap. Offboarding gets a ticket someone works through on a Friday. Yet it’s the more dangerous operation. Onboarding can be repeated. Offboarding can’t.

The three typical accidents

The export afterwards. The customer cancels, the database gets deleted. Three weeks later their tax advisor gets in touch and needs the invoices from the system. They no longer exist.

The half deletion. The script deletes the database, fails on the bucket because it isn’t empty, and stops. The bucket keeps costing money, the hostname points nowhere, nobody notices.

The wrong customer. Two customers with similar names, a typo in the script, and the paying customer’s database is gone. Without a backup, no apology will help.

All three have the same cause: deleting was the first step.

In what order does offboarding run?

  1. Back up. Every database of the customer as a dump, every non-empty bucket as an archive, into storage that belongs to the vendor, not the customer.
  2. Verify. Is the object really there? A HEAD request on the dump is the difference between "the upload went through" and "the file is there".
  3. Tear down. Only now: workload, records, database, buckets, in the reverse order of creation.
  4. Retain. The dump stays downloadable for as long as the rule says, with a log of who downloaded it when.

And the most important rule: if step one or two fails, step three doesn’t happen. An offboarding that carries on without a backup isn’t an offboarding, it’s data loss waiting to happen.

Where does the dump run?

A detail that surprises people in production: a managed database with a private endpoint is not reachable from outside, not even from the control plane. The dump has to run where the network is, i.e. in the cluster. A job with pg_dump or mariadb-dump that uploads its result to the backup bucket via a signed URL is the usual pattern. Password and URL come from a Secret that dies with the job.

Which format should the backup have?

For PostgreSQL, the custom format of pg_dump: compressed, selectively restorable with pg_restore. For MySQL, an SQL dump, restorable with the client. Both are standard formats you can read without the vendor. That matters: a backup only the platform can read is half a backup.

Buckets are the bigger problem

Databases are small and structured. Buckets are large and arbitrary. A customer with 40 GB of uploads needs an archive at offboarding that is larger than any database, and a bucket can only be deleted once it’s empty. Three rules:

  • Archive before even a single object is deleted, and verify the archive.
  • Then delete exactly the archived objects, never empty the bucket blindly. Whatever was added between archiving and deleting stays.
  • Set an upper limit. An archive beyond a few gigabytes is a manual case, not automation.

Retention is a decision, not a default

How long does the dump stay? Data protection says: as briefly as necessary. The customer says: until I have the export. Commercial law says, for some data: years. There is no universally valid number. There is only the obligation to set one, configure it per storage and log deletions. And the rule that automation only deletes what it wrote itself. Other objects in the same bucket stay untouched.

An offboarding is good when it feels boring: back up, verify, tear down, log. The same every time.

What this looks like in Clusterward

Offboarding a tenant starts with a dump of every database and an archive of every non-empty bucket, both into the workspace’s backup bucket, both verified via HEAD. Only then does the onboarding pipeline run in reverse. An error during backup stops everything before anything has been changed. Dumps remain downloadable according to the instance’s retention rule, and every download is recorded in the audit log. Details under Tenant pipelines and Managed databases. Verified backups like these are also part of the evidence NIS2 requires; see Audit & NIS2.

The checklist for the process

  • Is the retention period for dumps defined, configured per storage and written down somewhere the data protection officer can find it?
  • Does the offboarding know which resources belong to the customer? A list from onboarding, not a search pattern over names.
  • Does the dump run where the database is reachable, and is the result verified before anything is deleted?
  • Are buckets archived before even a single object is deleted, and is there an upper limit on archive size?
  • Does the process halt on every error instead of skipping ahead?
  • Is every step logged with time and person, including later downloads of the dump?
  • Can the process be rerun after an error without touching the data that has already been backed up?

If you can answer yes to all seven, you have an offboarding. If you hesitate on one, you have data loss that just hasn’t happened yet.

What the customer gets

A good offboarding ends with a message to the customer: your data was backed up on date X, the export is available until date Y, after which it will be deleted. The export is the dump in a standard format, downloadable via a short-lived link. That’s not just courtesy. It’s proof that you neither lost data nor kept it longer than necessary, and it answers the tax advisor’s request before it comes in.

Testing the offboarding

The only way to trust an offboarding is a test customer that is regularly created and torn down again. Once a month, onboard a tenant, add data, offboard it, download the dump, load it into an empty database. Ten minutes that make the difference when it counts.

Running the restore test in the cockpit

The monthly test gets easier when the restore doesn’t touch anything running. In Clusterward, “Restore into a new database…” loads a backup next to the running database, with its own user and password. That way you verify every backup without touching production. How this works together with volumes and earlier versions is described under Backups & recovery.

Conclusion

Write down your offboarding before your tenth customer, with a fixed order and fixed retention. After that it becomes a process anyone on the team can run, even on a Friday.

Setting up offboarding for your app? Tell us what a customer owns in your system: databases, buckets, domains. We’ll show you how backup and teardown run in a fixed order. Discuss offboarding →

Sources and further reading

Frequently asked questions

  • Back up, verify, tear down, then retain. Every database is backed up as a dump and every non-empty bucket as an archive, and the backup is verified via a HEAD request. Only then are workload, records, database and buckets removed, in the reverse order of creation. If backing up or verifying fails, nothing is deleted.