containerd part 3: the fix got reverted
People share their victories. A month later mine got reverted, and that was the right call.

People share their victories. A month ago I shared mine: my containerd fix got merged after 40 days and 30 review rounds.
13 days later it got reverted. It never made it into a release.
My fix put a timeout on the one call that hung on my nodes. It worked for my incident, and it felt like a win.
So I opened a second PR to cover more of the calls that could hang the same way.
That's where it came apart. Working through it with the maintainers, we saw the real problem was one layer out. The thing hanging was an external plugin, and it can hang on any of its calls. Patching them one at a time would never end.
The right fix is a single timeout where containerd talks to the plugin, covering every call at once. That made my first fix redundant, so with a release about to be cut, Michael Brown reverted it. I agreed.
In Part 2 I said the review is the distance between a patch that's correct for your incident and one that's correct for everyone. Turns out it doesn't stop at merge.
A new PR is on its way. Round 31. Not giving up.
References
- Original First published on LinkedIn
- PR · merged containerd/containerd#13799: metadata: bound snapshotter Remove during garbage collection
- PR · merged containerd/containerd#14119: Revert "metadata: bound snapshotter Remove during garbage collection"
- PR · merged containerd/containerd#14187: snapshots/proxy: add optional default_timeout for proxy snapshotter calls
Comments