Clusterward
← Back to the blog
OperationsUpdated Florian Apel

Logs, usage and uptime without your own monitoring stack

During an incident, three questions matter: What is the application writing, how heavily was it used, and can it be reached from outside? Not every team needs its own monitoring stack for that.

Cover image: Logs, usage and uptime without your own monitoring stack

During an incident, the same three questions always come up. What is the application writing right now? How heavily was it used over the last few hours? And can it even be reached from outside? The usual answer to all three is a monitoring stack: Prometheus, Grafana, Loki, Alertmanager. For large platform teams, that’s the right call. For a team with a handful of services, it’s often a second project that needs maintaining before it answers the first question.

Question 1: What is the application writing?

kubectl logs reads a container’s output. That’s fine for one instance. With three instances it gets tedious, and few people know the most useful options:

  • kubectl logs -l app=web --prefix --timestamps reads all pods with the label and marks where each line comes from. By default at most five pods at a time, more with --max-log-requests.
  • --since=1h limits output to the last hour, --tail=500 to the last lines.
  • --previous reads a container’s previous run. This is the single most important option: when a container crashes and restarts, it starts a fresh log. The error message that led to the crash is only in the previous run.

The limits: lines from multiple pods don’t arrive sorted by time, searching means grep, and anyone without kubectl and cluster access sees nothing at all. And logs live only as long as the pod. Once it’s gone, so are they.

Question 2: How heavily was it used?

kubectl top pods shows CPU and memory, supplied by the metrics-server. But only right now. The question during an incident is almost always: What did it look like an hour ago, when it started? For that you need history, and the metrics-server doesn’t store it.

What history should show to help you find the cause:

  • CPU and memory per service, not just per pod
  • the request and the limit as a line, because memory at the limit is the most common reason for restarts
  • restarts as markers on the timeline
  • at least a week back, ideally a month, to spot patterns

The pattern you find most often: memory rises slowly, hits the limit, the container is killed, restarts, memory drops, rises again. In the history, that looks like a sawtooth with red markers. Without history, it looks like occasional, unexplained restarts.

Question 3: Can the application be reached?

Running pods don’t mean the application is responding. A wrong certificate, an ingress without a rule, a DNS record pointing nowhere: everything is running, and nobody gets through. You only see that if you check from outside, via the same address your users use.

A good uptime check:

  • checks the real domain over HTTPS, not an internal address
  • treats certificate errors and timeouts as downtime
  • only alerts after two consecutive failures, not on a single blip
  • also reports when things are back up, so nobody keeps searching

When your own stack is worth it

Prometheus, Grafana and Loki are excellent. They’re worth it if you want to analyze your own application metrics (orders per minute, queue length), need dashboards across many services, have to search logs going back months or want tracing across multiple services. The price: memory and CPU in the cluster, retention, upgrades, and someone to maintain it all. If you’re honest, you’ll often find that the three questions above cover ninety percent of cases.

A typical case

Monday morning, a customer reports that the export has been failing intermittently since the weekend. No alert, all pods running. Here’s how it plays out with the three answers:

  1. Look at usage. The memory history for the last seven days shows a sawtooth: memory climbs over hours up to the limit, drops abruptly, climbs again. Underneath, red markers, each one a restart.
  2. Read the previous run. The current logs show nothing unusual; after all, the container is fresh. The previous run ends in the middle of an export, and the pod status gives the reason as OOMKilled, exit code 137.
  3. Cause. The export loads all of a customer’s records into memory at once. For small customers it doesn’t matter; for the largest customer, it blows past the limit.

Short term, more memory; long term, an export in batches. Without history and without the previous run, this would have stayed a mystery for days.

Log lines that help you search

The best interface doesn’t help much if the logs say nothing. Four rules:

  • One line per event. Multi-line output only for stack traces, and those belong directly below the error line.
  • Level first. ERROR, WARN, INFO are what make an error filter possible in the first place.
  • A request ID. Write it into every line and you’ll find all lines of a request across all instances.
  • No secrets. No passwords, tokens or full connection strings, not even while debugging. Logs end up in downloads, tickets and chats.

Alerts nobody ignores

An alert that fires for no reason ten times a day gets ignored after a week, including on the day it’s right. So: only alert after two consecutive failures, one message per outage rather than per check, an all-clear when things are back up, and a team channel rather than a personal address.

How Clusterward answers the three questions

On the service in the cockpit: the logs of all instances merged by time, with time ranges up to 24 hours, up to 10,000 lines, search with jump to the next match, “Errors only”, the previous run of a crashed container and download as a file. Clusterward records CPU, memory and instance usage itself every minute and shows it from 15 minutes to 30 days, as a share of the limit, with the request and restarts; a click opens the chart large. Uptime checks test the service’s domains over HTTPS and report outages to Slack, Teams, webhook or email. No agent in the cluster, nothing to install. The overview is in Logs & monitoring, the alert channels in Notifications. And if the logs show the last release is to blame: Rollbacks on Kubernetes: why a tag is not a version.

What it deliberately isn’t: a log archive spanning months or a tool for your own application metrics. If you need that, add a stack and still keep the three answers on the service.

Conclusion

During an incident, three answers matter: the logs of all instances including the previous run, the usage history and a check from outside. Make sure your team has these three within a minute. Whether there’s a monitoring stack of your own behind them is a secondary question.

Monitoring for your services? Describe your applications and what you do first during an incident today. We’ll show you how much of it works without your own stack. Ask a question →

Sources and further reading

Frequently asked questions

  • Use kubectl logs --previous to read a container’s previous run. This matters because a crashed container starts a fresh log after it restarts. The error message that led to the crash is only in the previous run, not in the current logs.