![card](card.jpg)

## Pod stuck in Terminating: what it is actually waiting for

`kubectl get pods` shows `STATUS Terminating` and it has been that way for an hour. The short answer: a Terminating pod is not dying, it is waiting. Either the grace period is still running, or some layer has not confirmed the containers are gone. Check four things, in this order: the grace period, finalizers, the node, and the runtime. Do it before you reach for `--force`.

## What does kubectl delete actually do?

It writes a timestamp. The API server sets `deletionTimestamp` and `deletionGracePeriodSeconds` on the pod and returns. Nothing has been killed yet. From that moment the pod shows as Terminating.

Everything else is a chain of requests, each layer asking the next one to make something stop:

1. The EndpointSlice controller marks the pod as terminating, and Services normally stop sending it new traffic. This runs in parallel with the rest, which is why a few requests still land on a pod that is already leaving.
2. kubelet on the node sees the timestamp and runs the preStop hook, if there is one.
3. kubelet asks the container runtime to send the stop signal to each container, SIGTERM unless the image sets STOPSIGNAL.
4. When the grace period runs out, the runtime sends SIGKILL.
5. kubelet confirms to the API server that the containers are gone, and the pod object is removed. If the pod has finalizers, it stays until a controller removes them.

The grace period is `terminationGracePeriodSeconds`, 30 by default, and preStop spends from the same 30. If preStop is still running when the clock ends, kubelet gives it a one-time 2 second extension, then moves on.

So once the grace period has passed, "stuck in Terminating" means one of these steps is not getting an answer.

## Where is it stuck?

Check in this order, cheapest first.

```bash
# 1. Is the grace period still running?
kubectl get pod $POD -o jsonpath='{.metadata.deletionTimestamp} {.metadata.deletionGracePeriodSeconds}{"\n"}'
# 2. Is a finalizer holding the object?
kubectl get pod $POD -o jsonpath='{.metadata.finalizers}{"\n"}'
# 3. Is the node alive?
kubectl get node $(kubectl get pod $POD -o jsonpath='{.spec.nodeName}')
# 4. Is the runtime stuck? kubelet's events first.
kubectl get events --field-selector involvedObject.name=$POD
# Then on the node from check 3: find the sandbox ID, list its containers.
crictl pods --name $POD
crictl ps -a --pod SANDBOX_ID
```

**1. The clock is just long.** Someone set `terminationGracePeriodSeconds: 3600` for a worker that drains a queue. `deletionTimestamp` is the deadline, grace period already included. If it is still in the future, the pod is draining and doing exactly what it was told. Wait, unless the number was set by mistake.

**2. A finalizer is holding the object.** The object stays until every finalizer is removed. A finalizer says nothing about the containers, so confirm with check 4 that they are really gone. A common one is `batch.kubernetes.io/job-tracking` on Job pods, which the Job controller removes once it has counted the pod. If the controller that owns the finalizer is healthy, wait. If it was uninstalled, remove the finalizer with a patch, knowing that you are skipping whatever cleanup it was there for.

**3. The node is gone or NotReady.** Only kubelet can confirm the containers stopped. If kubelet is dead or the node is unreachable, nobody confirms, and the pod stays Terminating until the node comes back, the Node object is deleted, or the node is NotReady and you add the `node.kubernetes.io/out-of-service` taint, once you know the machine is really down.

**4. The runtime below kubelet is stuck.** kubelet keeps retrying the kill and records `FailedKillPod` events, often with `DeadlineExceeded`. Now the problem is in containerd, the shim, or the sandbox under it. `crictl` on the node tells you whether the runtime still thinks the containers run. Before you change anything, save the events, `journalctl -u kubelet` from the node, `crictl inspectp SANDBOX_ID` for the pod sandbox and `crictl inspect` for each container. That is what the runtime maintainers will ask for.

## Can SIGKILL itself hang?

With a plain runc container, SIGKILL goes from the host kernel straight to the process, and it can't be ignored.

Add a sandbox and that changes. Take gVisor, a kernel written in Go that runs in user space. The runtime doesn't send a host SIGKILL to the app. It asks gVisor to deliver the kill inside the sandbox, and gVisor has to pause every task and wait for them to stop.

If the sandbox stops answering, kubelet retries the kill every 2 minutes, forever, and the pod stays in Terminating. That is what [google/gvisor#14405](https://github.com/google/gvisor/issues/14405) shows in a thread dump from a frozen sandbox. The cause was one lost signal, and a wait that printed a warning at 30 seconds and kept going. The fix, [google/gvisor#14201](https://github.com/google/gvisor/pull/14201), resends the signal instead of sending it once, and dumps stacks when the wait runs long.

> In a sandbox, SIGKILL is a request to the kernel you added.

## What about --force?

`kubectl delete pod $POD --grace-period=0 --force` removes the API object without waiting for kubelet, unless a finalizer still holds it. kubectl warns that the resource may continue to run on the cluster indefinitely, and it means it. The containers, the IP and the volume mounts can all stay on the node.

For a Deployment, that means the old process may run next to its replacement. Decide whether your app can live with that. For a StatefulSet it can put two pods with the same identity on the cluster at once, so first make sure the old process is stopped or its node is powered off. Either way, save the kubelet events and the runtime state before you force delete.

## Lessons

**Terminating is a wait, not a state of dying.** Read it as "someone has not answered yet" and go find who.

**Check the cheap layers first.** The grace period, finalizers and the node need no shell on the node, so rule them out before the runtime.

**Every layer needs its own timeout.** kubelet had one, the sandbox did not. A timeout that only prints a warning is not a timeout.

**Force delete removes the record, not the process.** It is the last step, not the first.

_Nahum Litvin is a Technical Lead - DevOps Engineering at Wix, working on the Kubernetes platform that runs untrusted user code._