In this post
I write services. I do not run the cluster. But there are a few things about Kubernetes that, if a developer does not understand them, will make a service drop requests on every deploy without a single wrong line of code. This post collects them.
Before I joined Vinova I deployed the familiar way: push to Vercel or Render, or run a process on a server and let it live there until I stopped it. The process was mine. If I did not touch it, it did not die.
Kubernetes is different. The first time I saw my pod sitting in Terminating when I had done nothing, I assumed something had broken and asked around. It turned out the system was working exactly as designed.
A warning first: I am not a DevOps engineer. I am writing as someone who builds services in Go and needs to know just enough that my code does not ruin a deploy. Anything about running the cluster itself, the platform people know far better than I do.
Pods die all the time
A pod can be stopped for many reasons, none of them my fault:
- A new version is rolling out and the old pod has to make room.
- A node needs maintenance or an upgrade, and every pod on it is moved.
- The autoscaler sees less load and removes pods.
- The pod used more memory than it was allowed.
So the question stops being "how do I keep the service from ever dying" and becomes "how does it die without anyone noticing". That sounds odd, but once the thinking changes, most of the rest follows.
Three probes, three different questions
Kubernetes does not know whether my service is healthy. It asks, by calling endpoints I declare.
| Probe | What it asks | When the answer is no |
|---|---|---|
startupProbe | Have you finished starting? | The container is killed and restarted. The other two probes do not run until this one passes |
readinessProbe | Can I send you requests? | The pod is taken out of the list that receives requests. It is not killed |
livenessProbe | Are you still alive? | The container is restarted |
containers:
- name: api
startupProbe:
httpGet: {path: /healthz, port: 8080}
failureThreshold: 30
periodSeconds: 2
readinessProbe:
httpGet: {path: /ready, port: 8080}
livenessProbe:
httpGet: {path: /healthz, port: 8080}
By default each probe runs every 10 seconds and has to fail 3 times in a row before it counts.
What I got wrong at first was treating readiness and liveness as the same thing, and making both check the database connection. It sounds reasonable: if the database is down, the service is useless anyway.
Think it through, though. The database is slow for a minute. Readiness fails and the pod stops receiving requests. Correct. Then liveness fails too, and Kubernetes restarts every pod, none of which did anything wrong. When the database recovers, they are all starting at once and all rushing to open connections.
So my rule now: /healthz only answers "this process is running" and touches nothing outside it. /ready is where I check the things the service needs in order to serve.
requests and limits, and two kinds of punishment
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
memory: 256Mi
requests is what the scheduler uses to place the pod: only a node with that much to spare can take it. The service may still use more than this if the node has room.
limits is the ceiling, and here is something I only learned from the documentation: going over the CPU ceiling and going over the memory ceiling are handled completely differently.
- Over on CPU, you are throttled. The kernel will not give you more. The service slows down but stays alive.
- Over on memory, you may be killed. The kernel cleans up with an OOM kill: no warning, no chance to tidy up. The pod shows
OOMKilledwith exit code 137.
In the example above I left out a CPU limit on purpose. Throttling produces a kind of slowness that is very hard to trace: nothing fails, responses are just late now and then. Policies differ between clusters, so check with your platform team before copying this.
What happens when a pod is deleted
This is the most important part of the post, and the part I misunderstood for longest.
I used to think the order was: Kubernetes stops sending requests to the pod, and then tells the pod to stop. That would be sensible, wouldn't it?
It is not what happens.
- 0sThe pod is deletedIt is marked Terminating. The 30 second grace period starts counting
- 0sTwo things start at the same timeTrack A: remove the pod from the list that receives requests. Track B: the kubelet starts shutting the container down
- 0spreStop runs, if you declared oneFor example, sleep 5 seconds
- 5sSIGTERM reaches process 1The service stops taking new work and finishes what is in flight
- 30sThe grace period endsAnything still running gets SIGKILL
Look at the second row. Track A and track B run in parallel, and neither waits for the other.
Track A sounds quick, but it has to spread to a lot of places: kube-proxy on every node has to rewrite its rules, the ingress controller has to find out, and so does the cloud load balancer. That takes a few seconds. During those seconds, requests are still being sent to a pod that may already have shut down.
The result is a handful of 502s on every deploy. Not many, not regular, and not reproducible on my machine. The kind of bug that makes you suspect everything except the actual cause.
How to fix it
The idea is simple: do not shut down straight away. Wait a few seconds for track A to spread, then stop.
The fix that needs the least code is a preStop hook:
containers:
- name: api
lifecycle:
preStop:
exec:
command: ["sleep", "5"]
terminationGracePeriodSeconds: 30
Two things to remember. First, the preStop time comes out of the 30 second grace period. Sleep for 5 and the service has 25 left to finish its work. If your requests can run longer than that, raise terminationGracePeriodSeconds.
Second, the sleep command has to exist in the image. A distroless image has no shell and no sleep. Recent Kubernetes versions add a built-in sleep action for preStop that needs nothing in the image, so check whether your cluster supports it.
The rest is in the code. On SIGTERM the service has to stop accepting new requests and finish the ones in progress:
ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGTERM, syscall.SIGINT)
defer stop()
go func() {
if err := server.ListenAndServe(); err != nil && !errors.Is(err, http.ErrServerClosed) {
log.Fatal(err)
}
}()
<-ctx.Done() // Kubernetes says: get ready to stop
ready.Store(false) // /ready starts returning 503
shutdownCtx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
defer cancel()
if err := server.Shutdown(shutdownCtx); err != nil { // wait for requests in flight, 20 seconds at most
log.Printf("shutdown: %v", err)
}
Twenty seconds because five went to preStop, and I want a few in reserve before the 30 second mark.
The trap in the Dockerfile
This one cost me the most time, because the code was correct and still did not run.
# shell form: process 1 is /bin/sh, and your service may never see SIGTERM
CMD ./api
# exec form: your service is process 1 and gets the signal directly
CMD ["./api"]
Kubernetes sends SIGTERM to process 1 in the container. Write CMD in shell form and process 1 is sh, which usually does not pass the signal on to your service. So the careful shutdown code above never runs. The service sits there for the full 30 seconds and then gets SIGKILL.
The tell: on every deploy, the old pod always takes exactly 30 seconds to disappear.
The list I keep next to my screen
/healthzchecks nothing outside the process./readyreturns 503 as soon as SIGTERM arrives.- A
preStopsleeps a few seconds, or the service waits by itself before stopping. - preStop time plus shutdown time is less than the grace period.
CMDis written in exec form.- Always set
requests, and always set a memory limit.
What changed in my head
After two months, the biggest difference is not the YAML. It is how I think about a process's memory.
I used to keep a few things in memory for convenience: a small cache, a list of work in progress. Now I treat everything in memory as temporary, liable to vanish at any moment. Anything that has to outlive a pod belongs somewhere else: Redis, the database or Kafka.
Written down it looks obvious. But I had to watch my own pod get killed a few times before I actually wrote code that way.
That is all. I wrote this from a developer's seat, so it certainly has gaps. If you work on the platform side and see something wrong, please tell me. Thanks for reading.