Clusterward
← Back to the blog
ArchitectureUpdated Florian Apel

From ingress-nginx to Envoy Gateway: switching controllers without a dark window

Change the ingress class and deploy: the new controller serves, the old one no longer does, and DNS still points at the old one. The host answers 404 until the TTL expires. Here’s a better way.

Cover image: From ingress-nginx to Envoy Gateway: switching controllers without a dark window

The Kubernetes Gateway API is the successor to the Ingress object, and Envoy Gateway is one of its most mature implementations. If you maintain nginx annotations today, you’ll want to get there eventually: typed routes, policies instead of annotations, timeouts and retries as fields. The path there has a trap you only see once a host answers 404.

What happens with a naive switch?

A service has an ingress class. You change it from nginx to envoy, deploy, and the new controller renders a Gateway and an HTTPRoute. The old Ingress is cleaned up because it no longer belongs to the desired state. Tidy.

Except: every controller has its own load balancer with its own public address. DNS still points at the old one. Visitors land on nginx, which no longer knows the host, and get the default backend 404. Until the DNS TTL has expired and all resolvers have the new address, the host is dark. With a TTL of one hour, that’s an hour. We experienced this live, on a host with real users.

The right process: cutover

  1. Serve from both. The service is rendered for the new controller, but the old Ingress stays. nginx and Envoy both answer for the host, each via its own load balancer. Whichever address DNS returns, the host works.
  2. Move DNS. The record now points at the address of the new load balancer. While the TTL expires, visitors arrive sometimes here, sometimes there. Both are correct.
  3. Verify. Every host of the service resolves to the new address – checked from outside, not from within the cluster – behind the Cloudflare proxy through the DNS provider’s API, because public DNS there only shows Cloudflare’s addresses.
  4. Clean up. Only now, and with an hour’s gap so cached answers expire, are the old Ingress objects deleted. The old load balancer no longer serves the host, but nobody arrives there anymore either.

Four steps, of which the first is the unfamiliar one: deliberately keeping two renderings of the same service in the cluster. That is not drift – it is a state with a name.

What has to stay the same

Both controllers must behave identically, otherwise users notice the switch:

  • TLS. A Let’s Encrypt certificate per host or a wildcard secret. Envoy Gateway needs cert-manager with Gateway API support, which is a feature gate or a configuration option depending on the cert-manager version; it has to be running before the first Gateway.
  • Redirects. www to apex, old domains to new ones. An annotation with nginx, a RequestRedirect filter with Envoy. The status code should be the same.
  • Access lists. Allowed source addresses as an nginx annotation or as a SecurityPolicy with Envoy, with the proxy protocol on both load balancers.
  • Body size. nginx limits it to 1 MB if nobody sets anything; Envoy has no limit. If you had allowed 64 MB with nginx, you won’t notice a difference with Envoy. The other way around, you will.
  • Proxy protocol. The new load balancer needs the same setting as the old one, otherwise Envoy only sees load balancer addresses.

What does Envoy do differently from nginx?

One thing that stands out during the switch: Envoy Gateway can merge several Gateway objects into one proxy fleet. One Gateway per service, all behind one load balancer. For operators with many services, that is the real gain: policies per service, one load balancer for all. With nginx, per-Ingress settings are free-text annotations; with Envoy, they are typed objects that the API server validates. For new setups, the Envoy Gateway documentation now recommends ListenerSet from the Gateway API over merging: one central Gateway, with each team defining ports, hostnames and certificates in its own ListenerSets.

When the switch isn’t worth it

If you have five hosts, need no timeouts or retries and are happy with nginx annotations, you gain little from the switch and temporarily pay for a second load balancer. The Gateway API pays off once per-service network behavior becomes an issue: rate limits per customer, conditional retries, header manipulation, sticky sessions via headers.

Preparing the cluster

Before the first service can switch, the cluster needs the platform for the Gateway API. That takes more steps than you might think:

  1. Install Envoy Gateway as a second controller alongside nginx, with its own load balancer and – if nginx has it – the proxy protocol as well.
  2. cert-manager with Gateway API support. A feature gate or a configuration option, depending on the version. cert-manager then has to restart, because it only creates the Gateway informers at startup if the CRDs already exist.
  3. A GatewayClass and an EnvoyProxy configuration that merges the Gateways of several services into one proxy fleet.
  4. A shared HTTP Gateway on port 80 that serves cert-manager’s HTTP-01 challenges and the redirect to HTTPS.
  5. A second ClusterIssuer with an HTTP-01 solver via the Gateway, using the same account as the nginx issuer.

Five objects that all have to be right before the first certificate is issued via Envoy. On staging first.

Six mistakes we made

  • cert-manager with the feature gate, but without a restart. Gateway support was configured, yet the gateway shim still wasn’t running, because cert-manager had started before the CRDs. The certificate stayed at "pending", without an error message. A rollout of cert-manager after installing the CRDs solved it.
  • A simple class change in the settings. The first cutover wasn’t one: class changed, deployed, nginx Ingress gone, DNS still pointing at nginx. Forty minutes of 404 on a host with users. Since then, the simple switch no longer exists – only the cutover.
  • The old controller with the new class. During the cutover the old controller kept rendering its objects, but already with the new class. nginx no longer felt responsible and answered 404, although DNS still pointed at nginx. In a cutover every controller therefore renders with its own class.
  • www without a certificate on Envoy. A host’s certificate still belonged to the nginx Ingress, and cert-manager’s gateway shim does not take over a certificate an Ingress owns. During the cutover the Gateway therefore uses the secrets nginx already fills; only at the end does the certificate pass to the Gateway.
  • Error 525 behind Cloudflare. Hosts without a certificate of their own behind Cloudflare in “Full” mode got nginx’s placeholder certificate, Envoy had no HTTPS listener at all, and Cloudflare answered 525. Such hosts now get a self-signed origin certificate.
  • Finished too early. With a TTL of 3600 seconds, visitors kept hitting the old load balancer for an hour after completion, and it no longer knew the host. Since then a cutover finishes only one hour after every host points to the new address.

When should you go back to nginx?

The way back is a cutover too, in the other direction. As long as nginx stays installed, it is possible at any time. Only once the last service runs on Envoy and a quarter has passed without incident is it worth uninstalling nginx and giving up the second load balancer.

How Clusterward encapsulates this

“Switch class…” in the networking block of the service page shows for every host before it starts who moves the record – Clusterward for its own, adopted or missing records and for hosts only a wildcard answered, or you at your DNS provider – and what happens to the certificate. Starting rolls out the service with both controllers, each with its own class and valid certificates. As soon as the new controller serves, Clusterward points the records in registered Cloudflare zones at the new address and reads them back through the Cloudflare API; that is how it detects the move behind the proxy too. One hour after every host points to the new address, Clusterward completes the cutover and rolls out once more to remove the old objects. Until then the switch can be aborted; if a host already points to the new address, the way back runs as a cutover again. A switch via a simple setting is rejected, because that is exactly what creates the dark window. The same applies to the cluster’s default class: as long as services without their own class serve a host through it, Clusterward refuses the change until those services are pinned to the old class. Only a host behind the Cloudflare proxy in a zone that is not registered in Clusterward stays manual. Details under Networking & ingress.

Conclusion

A controller switch is not a deployment but a DNS move with two servers that both answer in the meantime. Plan it that way and there is no dark window. Simply change the class and you get one, as long as the TTL.

Planning a controller switch? Tell us how many hosts run through ingress-nginx today. We’ll show you how the move works with two load balancers and no downtime. Discuss the switch →

Sources and further reading

Frequently asked questions

  • Because every controller has its own load balancer with its own public address. If you only change the ingress class, the old Ingress is cleaned up while DNS still points at nginx. Visitors get the default backend 404 there until the DNS TTL has expired and all resolvers have the new address.