Clusterward
← Back to the blog
OperationsUpdated Florian Apel

Bus factor 1: when the cluster belongs to a single colleague

In many small teams, the cluster belongs to one colleague. That works fine until they go on vacation. Six steps to make sure operations no longer hinge on one person.

Cover image: Bus factor 1: when the cluster belongs to a single colleague

The bus factor is the number of colleagues who would have to drop out for a project to grind to a halt. In many small teams, it is exactly one for Kubernetes operations. One colleague set up the cluster, knows the Terraform files, has the admin token on their laptop and knows the order in which the onboarding script runs for a new customer. That works for years. Until they take two weeks off and a certificate expires.

How do you recognize bus factor 1?

  • Deployments only happen when one particular person has time.
  • There is a script for new customers, but only one person knows which parameters it needs.
  • The kubeconfig with admin rights lives on exactly one machine – or on five, and nobody remembers which.
  • Alerts go to a personal email address.
  • The docs in the wiki describe the state of things two upgrades ago.
  • Ask how to restore a database and the answer is: ask him.

If you recognize three of these, you have bus factor 1. That is no criticism of the colleague. It is the natural result of one person doing the work while everyone else was building features.

What does bus factor 1 cost?

The obvious costs are outages nobody can fix and releases that have to wait. The less obvious ones are more expensive:

Security. A personal admin token that is never rotated is the most valuable target in the entire company. When the colleague leaves the team, it would have to be replaced – and often nobody knows everywhere it has been stored.

Speed. Every infrastructure question lands with one person. They become the bottleneck, however fast they are.

Dependency. Whoever is the only one who understands the cluster holds a bargaining position that neither they nor the company should have. And they never really get a vacation.

How do you get to a bus factor of two or more?

1. Take inventory

Write down what exists: clusters, databases, buckets, domains, certificates, access credentials. Not as a wiki page that goes stale, but as a list generated by the system itself. This exercise alone usually turns up two forgotten access credentials and a certificate that expires in three weeks.

2. Personal access instead of a shared token

Everyone involved in operations gets their own access with a role. Developers may deploy but not delete clusters. The CI pipeline gets its own token with an expiry date. The admin token is no longer needed day to day and is stored safely – not on laptops. Incidentally, this is exactly what every auditor asks about; see NIS2 in Kubernetes operations.

3. Processes as data instead of scripts

An onboarding script is knowledge locked in code that one person understands. A described process that the system executes is knowledge anyone can operate: create the database, create the bucket, set up the domain, roll out the application – in that order, with the customer’s values. Once you have described it this way, you no longer need anyone who remembers the parameters. Tenant pipelines shows what this looks like as a pipeline.

4. The second person does it, the first one watches

The next Kubernetes upgrade, the next add-on update, the next restore: the colleague who has never done it runs it, while the experienced one sits alongside. After two rounds, there are two people who can do it.

5. Alerts to a channel, not to a person

An alert sent to a personal address disappears during vacation. An alert in a team channel is seen by everyone, and whoever takes it says so.

6. Docs where the work happens

The best docs are the ones you don’t have to look for: the hint next to the button, the log that shows how it was done last time, the list of changes with names. A wiki complements this, but it doesn’t replace it.

The vacation test

A simple test shows whether it has worked: the colleague who built the cluster takes two weeks off and is unreachable. During that time, the team deploys, onboards a new customer and handles an alert. If that works, the bus factor is greater than one. If not, you know exactly where it got stuck.

A typical course of events

Take a SaaS vendor with six developers, one cluster and around forty customers – a setup we see often. The cluster was set up three years ago by a colleague who has since handled every upgrade, every new customer onboarding and every overnight incident. He announces extended parental leave, which leaves six weeks.

The first two weeks typically go into the inventory, and it brings things to light such as access credentials of former employees, a wildcard certificate that is renewed by hand, or an onboarding script with eleven parameters. After that, every developer gets their own access, the CI gets its own token, and the admin token moves into a vault. Onboarding is described as a process, and two colleagues each onboard a customer under guidance. The next Kubernetes upgrade is done by a colleague who had never worked in the cluster before.

The goal is reached when incidents during the parental leave get handled without anyone picking up the phone.

What you should not do

Schedule a documentation sprint. Two weeks of writing wiki pages produces pages that are out of date after the next upgrade. Better: restructure the work so that it documents itself.

Make everyone learn everything. Not every developer needs to be able to do Kubernetes upgrades. Two or three people who master operations are enough for a small team, provided the processes are operable by everyone else.

Hand out a second admin token. If you raise the bus factor by giving the admin token to a second colleague, you have doubled the risk instead of spreading it. More people need more access with fewer rights, not more copies of the one with all of them.

How Clusterward helps

Clusterward takes over the parts that would otherwise live in one person’s head: clusters, databases, buckets, domains and deployments are created according to the same rules, no matter who creates them. Everyone has their own access with mandatory multi-factor authentication and a role, the CI has its own token, and the audit log shows who did what and when. Onboarding processes are pipelines, and alerts go to team channels. What this means for technical leadership is covered under For CTOs; how it compares cost-wise with running operations yourself, under Clusterward vs. DIY.

Conclusion

Bus factor 1 is not a people problem but a structural one. Hand out access instead of tokens, describe processes instead of writing scripts, and let the second person do the next maintenance. The best proof is a vacation during which nobody calls.

What is your bus factor? In a demo, we show how a second colleague deploys on day one – without private scripts and without a shared admin token. Request a demo →

Sources and further reading

Frequently asked questions

  • The bus factor is the number of colleagues who would have to drop out for a project to grind to a halt. With bus factor 1, Kubernetes operations hinge on one person: they set up the cluster, know the Terraform files, hold the admin token and alone know how the onboarding script for new customers runs.