Nahum LitvinKubernetes, containers and sandboxes, from the production side
containerd · part 22 min read

containerd part 2: 17 lines, 40 days, 30 reviews

Writing the fix was easy. The review was where the engineering happened.

The fix was 17 lines. Getting it into containerd took 40 days and 30 review rounds, and taught me more than writing it did.

Part 2 of the bricked-node story: the resilience patch got merged upstream this week. My biggest open source contribution to date.

Writing the code was the easy part. The review was where the real engineering happened.

The timeout key got renamed twice, because naming in a project with hundreds of config keys is an API commitment, not a variable name.

The default went from 30 minutes to 5 and back to 30. I pushed hard for 5, a maintainer pushed back with an argument I hadn't considered: the old behavior was unbounded, so on upgrade the least surprising default is the big one, and anyone running remote snapshotters can tune it down. He was thinking about thousands of clusters upgrading. I was thinking about mine.

My test got rewritten around Go's synctest, so it runs in milliseconds on a fake clock instead of mutating global state.

And when CI wedged in the merge queue, Michael Brown restarted the buckets, herded the reviewers, and queued it again until it landed. Maintainers do this invisible work daily, for free, for strangers.

A patch that's correct for your incident and a patch that's correct for everyone running containerd are two different things. The review is the distance between the two.

What's the hardest review you've been through, and what did it fix that the code didn't?

References

Comments