<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Nahum Litvin</title><description>Production write-ups on Kubernetes, containerd, gVisor and the tools around them.</description><link>https://www.catchkill9.dev/</link><item><title>Mermaid charts in Claude Code&apos;s terminal: what a mod could do that my VS Code harness couldn&apos;t</title><link>https://www.catchkill9.dev/posts/prismantis.html</link><guid isPermaLink="true">https://www.catchkill9.dev/posts/prismantis.html</guid><description>prismantis is a Claude Code mod that redraws replies with colored tables, highlighted code and mermaid charts, and nudges Claude to actually draw them. A year ago I tried the same thing as a VS Code extension and dropped it.</description><pubDate>Fri, 02 Oct 2026 00:00:00 GMT</pubDate><content:encoded>&lt;!-- vale Voice.Vocab = NO --&gt;

&lt;p&gt;&lt;em&gt;One reply on the Dracula theme. Everything here is text the mod redrew.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Every day I watch two terminals. In one, Codex prints a table with a colored header, a rule under it, versions in green. In the other, Claude Code prints the same table as pipes and dashes. prismantis is the fix: a Claude Code mod that redraws every reply with colored tables, highlighted code and mermaid diagrams, right inside the terminal.&lt;/p&gt;
&lt;h3 id=&quot;what-did-i-try-first&quot;&gt;What did I try first?&lt;/h3&gt;
&lt;p&gt;What I wanted was small: mermaid charts and some basic HTML rendering in Claude&apos;s replies. About a year ago I built a harness around Claude Code as a VS Code extension to get them.&lt;/p&gt;
&lt;p&gt;It was bad. I dropped it, and the tables stayed pipes and dashes.&lt;/p&gt;
&lt;h3 id=&quot;what-changed-on-1-oct&quot;&gt;What changed on 1 Oct?&lt;/h3&gt;
&lt;p&gt;Claude Code 2.1.287 shipped &lt;a href=&quot;https://claude.com/blog/claude-code-mods&quot;&gt;mods&lt;/a&gt;. A mod is a plugin made of TypeScript function hooks that run inside Claude Code. One hook is &lt;code&gt;ui.render&lt;/code&gt;. Claude Code calls it for every piece of the transcript it is about to draw, hands you the props, and asks what to draw instead.&lt;/p&gt;
&lt;p&gt;For an assistant reply, the props are the reply&apos;s markdown. You return a tree of &lt;code&gt;Box&lt;/code&gt; and &lt;code&gt;Text&lt;/code&gt; elements, and that is what lands on screen. No extension. No second window.&lt;/p&gt;
&lt;p&gt;The next day prismantis went from 0.1.0 to 0.3.2. Four versions, all dated 2 Oct in the changelog. I read the announcement in the morning and had a working table by lunch.&lt;/p&gt;
&lt;p&gt;The name is prism plus mantis. The mantis shrimp has 12 to 16 kinds of color receptors. Your terminal had 8.&lt;/p&gt;
&lt;h3 id=&quot;why-did-i-delete-the-best-looking-feature&quot;&gt;Why did I delete the best-looking feature?&lt;/h3&gt;
&lt;p&gt;Version 0.2 drew real mermaid diagrams. It shelled out to mermaid-cli, rendered a PNG, and showed it in terminals that speak the kitty graphics protocol.&lt;/p&gt;
&lt;p&gt;Mods run in-process, inside Claude Code, and they aren&apos;t sandboxed. A mod that spawns programs and writes image files is a mod you have to trust a lot more than one that only moves text around.&lt;/p&gt;
&lt;p&gt;0.3 removed image mode and the mermaid-cli dependency. Diagrams draw as colored box art now. The mod runs no external programs, writes no files, reads no files and makes no network calls.&lt;/p&gt;
&lt;h3 id=&quot;what-did-a-second-reviewer-find&quot;&gt;What did a second reviewer find?&lt;/h3&gt;
&lt;p&gt;I built it with Claude Code, in Claude Code. I described what I wanted, Claude read the API declarations the engine writes next to every mod, and I looked at screenshots and said &quot;almost&quot; a lot.&lt;/p&gt;
&lt;p&gt;Then I asked a second model, Codex, to review it. It flagged 12 issues.&lt;/p&gt;
&lt;p&gt;The worst one was quiet. The copy button on a table gave back my redrawn table instead of the markdown Claude wrote, so a pasted table came back reformatted. The other bad one: a fence of four backticks closed on the first three-backtick line inside it, which broke every code example nested in a code example.&lt;/p&gt;
&lt;p&gt;I fixed most of them, documented one and skipped two as rare. The fixes came with 10 regression tests, and I ran them against the old code first. 7 failed there, which is the only way I trust a regression test.&lt;/p&gt;
&lt;h3 id=&quot;why-didnt-claude-draw-any-charts&quot;&gt;Why didn&apos;t Claude draw any charts?&lt;/h3&gt;
&lt;p&gt;Once the charts rendered, I noticed something annoying. Claude almost never wrote one. It doesn&apos;t know the terminal can draw a mermaid block, so it writes a table or a paragraph instead.&lt;/p&gt;
&lt;p&gt;The obvious fix is one line in the system prompt. Plugins can&apos;t do that. Claude Code&apos;s built-in &lt;code&gt;sec-default&lt;/code&gt; policy keeps the system prompt for the organization, not for whatever you installed.&lt;/p&gt;
&lt;p&gt;So 0.3.2 added &lt;code&gt;diagramHints&lt;/code&gt;. The mod attaches a short note to each prompt you type, read by the model and never shown, saying tables, code, mermaid diagrams and &lt;code&gt;xychart-beta&lt;/code&gt; charts render here. It costs about 80 tokens per prompt. It is off whenever mermaid is off, and skipped for headless &lt;code&gt;claude -p&lt;/code&gt; runs.&lt;/p&gt;
&lt;p&gt;After that, charts started showing up without anyone asking for them.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Drawing the chart was the easy part. Getting Claude to write one was the work.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3 id=&quot;what-does-it-draw&quot;&gt;What does it draw?&lt;/h3&gt;
&lt;p&gt;Tables get a colored header, rules, column alignment from &lt;code&gt;:---:&lt;/code&gt; and colored numbers, sized to the terminal. Code gets a language header and Prism highlighting in 24 languages. Mermaid covers flowcharts, sequence, state, class and ER diagrams, plus bar and line charts, one color per box, participant and bar.&lt;/p&gt;
&lt;p&gt;Every code block, table, list and quote has a copy button; diagrams have two, one for the mermaid source and one for the drawn art. Tool calls shrink to one line, &lt;code&gt;Ran gh pr view 12&lt;/code&gt;, and collapsed groups to a summary like &lt;code&gt;Ran 3 commands, read 2 files&lt;/code&gt;. 15 MIT themes ship, plus &lt;code&gt;mono&lt;/code&gt;, and each of the 20 color slots can be overridden in &lt;code&gt;/config&lt;/code&gt;.&lt;/p&gt;
&lt;h3 id=&quot;what-did-the-mod-api-teach-me&quot;&gt;What did the mod API teach me?&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Hooks run without Node.&lt;/strong&gt; No &lt;code&gt;require&lt;/code&gt;, no package imports, no dynamic &lt;code&gt;import()&lt;/code&gt;. Every library has to be bundled into one ESM file. I bundle two, beautiful-mermaid and Prism, 206 KB together, both MIT.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The engine handle can&apos;t cross a file import.&lt;/strong&gt; &lt;code&gt;claude plugin validate&lt;/code&gt; doesn&apos;t notice. &lt;code&gt;claude plugin test&lt;/code&gt; does, with &quot;hooks module did not load&quot;. All my code that touches &lt;code&gt;$&lt;/code&gt; lives in one file now.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The built-in Markdown element takes no colors.&lt;/strong&gt; So a mod that themes replies has to parse markdown itself. Mine is regex, not CommonMark. It covers what Claude writes.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;There is no font size, and the clipboard takes text only.&lt;/strong&gt; Headings get bold, underline or a rule instead. A diagram can&apos;t be copied as a picture.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The whole thing is about 950 lines of TypeScript and 52 tests. CI runs them on Linux, macOS and Windows against the newest Claude Code, and on Linux against 2.1.287, the oldest with mods. It also type-checks, rebuilds the bundled libraries byte for byte, fails if a bundled package isn&apos;t MIT, and installs the plugin into a clean config.&lt;/p&gt;
&lt;h3 id=&quot;what-doesnt-work-yet&quot;&gt;What doesn&apos;t work yet?&lt;/h3&gt;
&lt;p&gt;The parser is not CommonMark, so nested quotes and HTML draw as plain text. CJK and emoji count as two columns in tables, and terminals disagree on a few emoji, so those can still be off by one. Pie charts don&apos;t render. I&apos;ve hand-tested it in the macOS terminal only.&lt;/p&gt;
&lt;p&gt;All of it is reversible. Press ctrl+o on any reply to see the original, or turn the mod off in &lt;code&gt;/plugin&lt;/code&gt; and Claude Code&apos;s own renderer comes back.&lt;/p&gt;
&lt;h3 id=&quot;how-do-i-try-it&quot;&gt;How do I try it?&lt;/h3&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;background-color:#fff;--shiki-dark-bg:#24292e;color:#24292e;--shiki-dark:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;text&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;/plugin marketplace add NahumLitvin/prismantis&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;/plugin install prismantis@prismantis&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Needs Claude Code 2.1.287 or later. Then ask Claude to print &lt;code&gt;docs/demo.md&lt;/code&gt; from the repo verbatim as its whole reply. Every feature is in that one file. Pick a theme in &lt;code&gt;/config&lt;/code&gt;. MIT licensed, source at &lt;a href=&quot;https://github.com/NahumLitvin/prismantis&quot;&gt;https://github.com/NahumLitvin/prismantis&lt;/a&gt;.&lt;/p&gt;
&lt;h3 id=&quot;credits&quot;&gt;Credits&lt;/h3&gt;
&lt;p&gt;Diagrams come from &lt;a href=&quot;https://github.com/lukilabs/beautiful-mermaid&quot;&gt;beautiful-mermaid&lt;/a&gt;, highlighting from &lt;a href=&quot;https://github.com/PrismJS/prism&quot;&gt;Prism&lt;/a&gt;. The palettes belong to their authors: Catppuccin, Dracula, Nord, Tokyo Night, Gruvbox, Rosé Pine, Everforest, GitHub Primer, One Dark and Solarized, all credited in the repo&apos;s &lt;a href=&quot;https://github.com/NahumLitvin/prismantis/blob/main/docs/THIRD_PARTY_NOTICES.md&quot;&gt;third-party notices&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Nahum Litvin, Lead Platform Engineer.&lt;/em&gt;&lt;/p&gt;</content:encoded><category>notes</category><category>claude-code</category><category>typescript</category><category>open-source</category></item><item><title>A zero timeout in Go is not no timeout</title><link>https://www.catchkill9.dev/posts/go-zero-timeout-is-expired.html</link><guid isPermaLink="true">https://www.catchkill9.dev/posts/go-zero-timeout-is-expired.html</guid><description>context.WithTimeout(ctx, 0) hands you a context that is already dead.</description><pubDate>Tue, 29 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;While working on snapshotter timeouts in containerd, I assumed a timeout of 0 meant &quot;no limit&quot;. It doesn&apos;t.&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;go&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;ctx, cancel &lt;/span&gt;&lt;span style=&quot;color:#F97583&quot;&gt;:=&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt; context.&lt;/span&gt;&lt;span style=&quot;color:#B392F0&quot;&gt;WithTimeout&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;(context.&lt;/span&gt;&lt;span style=&quot;color:#B392F0&quot;&gt;Background&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;(), &lt;/span&gt;&lt;span style=&quot;color:#79B8FF&quot;&gt;0&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#F97583&quot;&gt;defer&lt;/span&gt;&lt;span style=&quot;color:#B392F0&quot;&gt; cancel&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;()&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;fmt.&lt;/span&gt;&lt;span style=&quot;color:#B392F0&quot;&gt;Println&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;(ctx.&lt;/span&gt;&lt;span style=&quot;color:#B392F0&quot;&gt;Err&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;()) &lt;/span&gt;&lt;span style=&quot;color:#6A737D&quot;&gt;// context deadline exceeded&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The deadline is now, so the context is expired before the first call. containerd&apos;s &lt;code&gt;timeout.WithContext&lt;/code&gt; passes the configured value straight to &lt;code&gt;context.WithTimeout&lt;/code&gt;, so a key set to 0 fails every call on the spot.&lt;/p&gt;
&lt;p&gt;That&apos;s why the new &lt;code&gt;default_timeout&lt;/code&gt; for proxy snapshotters is off when empty, and only wraps calls whose context has no deadline yet: &lt;a href=&quot;https://github.com/containerd/containerd/pull/14187&quot;&gt;https://github.com/containerd/containerd/pull/14187&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;A 12 second &lt;code&gt;go run&lt;/code&gt; settled it. Cheaper than arguing about it.&lt;/p&gt;</content:encoded><category>containerd</category><category>go</category></item><item><title>One lost signal, five days stuck, 45,000 frozen threads: fixing a gVisor hang upstream</title><link>https://www.catchkill9.dev/posts/gvisor-one-lost-signal.html</link><guid isPermaLink="true">https://www.catchkill9.dev/posts/gvisor-one-lost-signal.html</guid><description>One lost signal hangs pod deletion. The maintainer&apos;s smaller fix was the better one.</description><pubDate>Tue, 29 Sep 2026 00:00:00 GMT</pubDate><content:encoded>
&lt;p&gt;It looked like a deadlock in gVisor. It was a single missed wake-up, hiding behind a symptom teams have been reporting since 2021, and two strangers and a maintainer fixed it in about a week.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The hang in one picture: one signal, no retry, and a deadline check whose only action is a log line.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;On a Tuesday morning our alerting claimed that 559 pods were stuck Terminating in production. The real number was one. The rule counts kubelet failed-kill events rather than pods, and a sandbox that never dies gets retried forever, so a single wedged pod read as a fleet on fire. That pod turned out to be the interesting part.&lt;/p&gt;
&lt;p&gt;At Wix, our Velo grid runs users&apos; backend JavaScript as untrusted code inside gVisor sandboxes on Amazon EKS. gVisor puts a userspace kernel between that code and the host kernel, so untrusted code never talks to the host kernel directly. Close to two thousand sandboxes per node group, short lifecycles, heavy churn. Deleting a pod should take seconds. This one had been Terminating since 07:11. kubelet&apos;s kill requests got DeadlineExceeded every 2 minutes and would have continued forever. We intervened by hand after about 90 minutes. A week earlier, eight pods on four nodes had sat like that for days.&lt;/p&gt;
&lt;h2 id=&quot;why-nobodys-timeout-saved-us&quot;&gt;Why nobody&apos;s timeout saved us&lt;/h2&gt;
&lt;p&gt;Before the forensics, the shape of the problem. Deleting a pod is a relay: each layer asks the next one, politely, to make something die. Every layer should hold two more things beyond the polite ask: a bound on how long it waits, and an escalation that works &lt;em&gt;without the cooperation of the thing it is waiting on&lt;/em&gt;. In our case the relay runs from kubelet (the Kubernetes node agent) to containerd (the container runtime) to the gVisor shim (the per-sandbox process containerd talks to) to the sandbox itself. Here is that stack, graded on all three:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Three arrows per hop: blue attempts, amber detects, violet recovers. Blue failed once, at the bottom. Amber either did not exist, fired into a log line, or fired and could only re-send blue. Violet existed nowhere, so one lost message at the bottom of the stack became the whole stack&apos;s hang.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Look at the violet column. Every escalation asked the very component being escalated against to cooperate. containerd&apos;s dead-shim cleanup only fires when the shim disconnects. Ours kept answering Connect and State, which never touched the lock, while Kill and Stats blocked behind it. The connection stayed up, so cleanup never fired. Alive enough to prevent its own cleanup.&lt;/p&gt;
&lt;h2 id=&quot;what-a-frozen-sandbox-looks-like&quot;&gt;What a frozen sandbox looks like&lt;/h2&gt;
&lt;p&gt;That userspace kernel is called the sentry. On systrap, the platform we run, the app&apos;s threads run in stripped-down stub processes and the sentry coordinates them. Our wedged pod had a sentry that was alive but answering nothing. runsc kill, gVisor&apos;s own command for stopping a sandbox, hung. runsc debug --stacks, the tool that is supposed to tell you why something hangs, also hung. The panic log was 0 bytes.&lt;/p&gt;
&lt;p&gt;The node told the rest of the story. Load average climbing about 3 per hour while actual CPU sat near 30%. Hundreds of sentry threads parked, waiting on shared memory. Stub processes accumulating. On the worst nodes, load kept climbing for days with the machine mostly idle. The only remediation that worked was SIGKILL to the whole sandbox process tree, by hand, over AWS SSM, a remote shell into the node. That became a runbook. Runbooks like that are a debt.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Stub processes pile up behind the frozen sentry: load average keeps climbing while the CPU does almost nothing. Threads stuck waiting, not working: the rising top panel against the flat bottom one is the bug.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;So we filed &lt;a href=&quot;https://github.com/google/gvisor/issues/14408&quot;&gt;gvisor#14408&lt;/a&gt; with what we had: the outside view, two affected releases, and an honest admission that we could not get stacks because the debug tooling was wedged along with the sandbox. Our proposed fix, PR #14201, had been open since the week before.&lt;/p&gt;
&lt;h2 id=&quot;then-a-stranger-showed-up-with-the-other-half&quot;&gt;Then a stranger showed up with the other half&lt;/h2&gt;
&lt;p&gt;The same day, another engineer filed &lt;a href=&quot;https://github.com/google/gvisor/issues/14405&quot;&gt;gvisor#14405&lt;/a&gt;. Same bug at a different company, and they had the one thing we could not get: a full dump of the sentry&apos;s goroutines, Go&apos;s lightweight threads, from inside a frozen sentry. Their dump showed one of the sentry&apos;s internal worker threads sitting in the same wait loop for over five days, waiting for a stub thread that had missed a single wake-up signal.&lt;/p&gt;
&lt;p&gt;In their shim, about 45,000 waiting threads were queued behind one lock held by a kill that would not finish, and their containerd and kubelet memory grew until nodes ran out.&lt;/p&gt;
&lt;p&gt;The bug itself is painfully simple. The sentry sends the stub one interrupt signal. One. If that signal is lost, and nobody has yet pinned down why it sometimes is, the sentry keeps waiting for the answer with no retry and no escape. There is even a 30 second deadline in the code that notices the wait is too long. Here is the whole thing it did about it:&lt;/p&gt;
&lt;p&gt;if time.Now().After(deadline) {
log.Warningf(&quot;Systrap task goroutine has been waiting on &quot; +
&quot;ThreadContext.State futex too long. ...&quot;)
// no retry, no escalation
}&lt;/p&gt;
&lt;p&gt;Where did 45,000 waiting threads come from? cAdvisor kept polling every container&apos;s Stats. Each request added a goroutine that blocked behind the hung kill&apos;s lock. Its caller timed out, but a mutex wait cannot observe cancellation, so the goroutine stayed. Over five days, 45,000 accumulated: roughly one every 10 seconds, the scrape interval fossilized in a thread dump. The metrics system asking &quot;how are you?&quot; every 10 seconds had pushed the shim to about 600MB.&lt;/p&gt;
&lt;p&gt;A timeout that only logs a warning is not a timeout. It is a diary.&lt;/p&gt;
&lt;p&gt;And this exact wait sits on the teardown path. Killing a sandbox starts with a freeze-everything step that waits for every worker thread to park, so one stuck worker means the kill command can never answer, the shim call hangs, kubelet gets DeadlineExceeded, and each retry parks another goroutine behind the first one. The pod stays Terminating for days.&lt;/p&gt;
&lt;h2 id=&quot;so-how-did-we-fix-it&quot;&gt;So, how did we fix it?&lt;/h2&gt;
&lt;p&gt;The merged fix escalates in stages, gentlest first:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;1. Tap again.&lt;/strong&gt; The interrupt is resent on each 5 second checkup wakeup instead of once. When the resend lands, a lost signal costs the workload a stall of a few seconds, then it carries on.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;2. Photograph the scene.&lt;/strong&gt; If the stub is still unresponsive after the 30 second deadline, the sentry dumps every internal stack trace to the log. A frozen sandbox can not be debugged from outside unless it was started with --panic-signal; ours were not, theirs were, which is where their dump came from. So this failure now documents itself. The next team to hit anything like this gets for free what took two companies a week to assemble.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;3. Remove only the broken part.&lt;/strong&gt; Then the sentry kills just the one stuck subprocess, through the same path it already uses when a stub dies naturally. The blocked task unwinds, teardown can finish, and the healthy subprocesses in the sandbox are untouched.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Update, 1 Oct 2026: a gVisor maintainer reported that step 3 also hits healthy subprocesses. A host page fault that takes more than 30 seconds, for example on a slow FUSE server, looks the same as a lost signal, because the stub cannot take the interrupt until the fault returns. &lt;a href=&quot;https://github.com/google/gvisor/pull/15182&quot;&gt;gvisor#15182&lt;/a&gt; keeps the resend and the stack dump, and kills only when the stub thread no longer exists, not after 30 seconds. It is in review.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;My first version killed the entire sentry. gVisor maintainer Konstantin Bogomolov pushed for killing only the stuck subprocess, sparing healthy neighbors. He was right. The other reporter had raised the same concern earlier that day. Konstantin also caught a dropped continue, which led me to a race: a context that recovered during the wait could still get killed. Three review rounds in under a week took the PR from nobody having looked at it to merged on master.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The merged fix, staged from least to most disruptive. My first version jumped straight to killing the whole sandbox; review scoped it down.&lt;/em&gt;&lt;/p&gt;
&lt;h2 id=&quot;the-part-that-bothers-me&quot;&gt;The part that bothers me&lt;/h2&gt;
&lt;p&gt;After the merge, we found reports since 2021 from a Cloudflare engineer, a GKE user, Knative, Talos Linux, Arista and others on the containerd tracker. Roughly eight teams over five years. Causes differed, and some cases got fixes. But the pattern kept returning: pods stuck Terminating, debugging tools hanging, someone killing processes by hand, and more often than not the thread ends without a stack trace from inside the sandbox.&lt;/p&gt;
&lt;p&gt;Two teams filing in the same week broke that pattern: our production evidence and proposed fix met the other team&apos;s internal stack dump.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Five years of the same symptom, with different root causes underneath. Most threads ended with the manual workaround, until two half-reports completed each other.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The relay diagram near the top of this post still isn&apos;t fully green. The merged fix repairs the layer where the hang started. We filed gvisor#14548 to bound shim kill waits, then SIGKILL the sandbox tree, and containerd#14081 to SIGKILL the shim after repeated kill RPC timeouts. In review of the shim PR, gvisor#14549, the same maintainer asked to narrow it to bounded waits on Kill, Stats and Status, because killing the sandbox from the shim would take out healthy containers too, the same over-reaction he flagged in the first fix. That revision merged, and escalation stays open. Either escalation, as originally filed, would have capped our incident at minutes.&lt;/p&gt;
&lt;h2 id=&quot;lessons-learned&quot;&gt;Lessons learned&lt;/h2&gt;
&lt;ul class=&quot;contains-task-list&quot;&gt;
&lt;li class=&quot;task-list-item&quot;&gt;&lt;input type=&quot;checkbox&quot; disabled&gt; &lt;strong&gt;A watchdog that only barks is not a watchdog.&lt;/strong&gt; If code detects a should-never-happen state, it must act: retry, self-report, fail. Grep your own codebase for deadline checks whose only body is a log line. This one surfaced because a sandbox sat stuck for five days behind it.&lt;/li&gt;
&lt;li class=&quot;task-list-item&quot;&gt;&lt;input type=&quot;checkbox&quot; disabled&gt; &lt;strong&gt;File the issue even with half the evidence.&lt;/strong&gt; Our report had no stacks and said so plainly. Within a day a stranger&apos;s independent report supplied them, and their dump changed the fix. The half-report you are embarrassed to file is someone else&apos;s missing half.&lt;/li&gt;
&lt;li class=&quot;task-list-item&quot;&gt;&lt;input type=&quot;checkbox&quot; disabled&gt; &lt;strong&gt;Remediate the smallest thing that unblocks you.&lt;/strong&gt; Kill the subprocess, not the sandbox. The instinct under pressure is the big hammer. Review cut the kill from the whole sentry to one subprocess. The same restraint belongs in your incident runbooks.&lt;/li&gt;
&lt;li class=&quot;task-list-item&quot;&gt;&lt;input type=&quot;checkbox&quot; disabled&gt; &lt;strong&gt;Audit your escalation paths for circular dependencies.&lt;/strong&gt; Walk your stack with that diagram&apos;s two questions: bounded wait, and escalation that works without the target&apos;s cooperation. If the answer to the second bottoms out at &quot;a human with SSH&quot;, write that down, because that is your actual design.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The fix shipped in gVisor release-20260831.0. Until our fleet is on it, the SSM kill runbook stays, but this one has an ending, and the next report of this hang will arrive with its stacks attached.&lt;/p&gt;
&lt;p&gt;This post was written by Nahum Litvin.&lt;/p&gt;</content:encoded><category>gvisor</category><category>kubernetes</category><category>upstream</category></item><item><title>containerd part 3: the fix got reverted</title><link>https://www.catchkill9.dev/posts/containerd-part-3-reverted.html</link><guid isPermaLink="true">https://www.catchkill9.dev/posts/containerd-part-3-reverted.html</guid><description>People share their victories. A month later mine got reverted, and that was the right call.</description><pubDate>Fri, 25 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;People share their victories. A month ago I shared mine: my containerd fix got merged after 40 days and 30 review rounds.&lt;/p&gt;
&lt;p&gt;13 days later it got reverted. It never made it into a release.&lt;/p&gt;
&lt;p&gt;My fix put a timeout on the one call that hung on my nodes. It worked for my incident, and it felt like a win.&lt;/p&gt;
&lt;p&gt;So I opened a second PR to cover more of the calls that could hang the same way.&lt;/p&gt;
&lt;p&gt;That&apos;s where it came apart. Working through it with the maintainers, we saw the real problem was one layer out. The thing hanging was an external plugin, and it can hang on any of its calls. Patching them one at a time would never end.&lt;/p&gt;
&lt;p&gt;The right fix is a single timeout where containerd talks to the plugin, covering every call at once. That made my first fix redundant, so with a release about to be cut, Michael Brown reverted it. I agreed.&lt;/p&gt;
&lt;p&gt;In Part 2 I said the review is the distance between a patch that&apos;s correct for your incident and one that&apos;s correct for everyone. Turns out it doesn&apos;t stop at merge.&lt;/p&gt;
&lt;p&gt;A new PR is on its way. Round 31. Not giving up.&lt;/p&gt;</content:encoded><category>containerd</category><category>upstream</category></item><item><title>containerd part 2: 17 lines, 40 days, 30 reviews</title><link>https://www.catchkill9.dev/posts/containerd-part-2-17-lines.html</link><guid isPermaLink="true">https://www.catchkill9.dev/posts/containerd-part-2-17-lines.html</guid><description>Writing the fix was easy. The review was where the engineering happened.</description><pubDate>Tue, 25 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;The fix was 17 lines. Getting it into containerd took 40 days and 30 review rounds, and taught me more than writing it did.&lt;/p&gt;
&lt;p&gt;Part 2 of the bricked-node story: the resilience patch got merged upstream this week. My biggest open source contribution to date.&lt;/p&gt;
&lt;p&gt;Writing the code was the easy part. The review was where the real engineering happened.&lt;/p&gt;
&lt;p&gt;The timeout key got renamed twice, because naming in a project with hundreds of config keys is an API commitment, not a variable name.&lt;/p&gt;
&lt;p&gt;The default went from 30 minutes to 5 and back to 30. I pushed hard for 5, a maintainer pushed back with an argument I hadn&apos;t considered: the old behavior was unbounded, so on upgrade the least surprising default is the big one, and anyone running remote snapshotters can tune it down. He was thinking about thousands of clusters upgrading. I was thinking about mine.&lt;/p&gt;
&lt;p&gt;My test got rewritten around Go&apos;s synctest, so it runs in milliseconds on a fake clock instead of mutating global state.&lt;/p&gt;
&lt;p&gt;And when CI wedged in the merge queue, Michael Brown restarted the buckets, herded the reviewers, and queued it again until it landed. Maintainers do this invisible work daily, for free, for strangers.&lt;/p&gt;
&lt;p&gt;A patch that&apos;s correct for your incident and a patch that&apos;s correct for everyone running containerd are two different things. The review is the distance between the two.&lt;/p&gt;
&lt;p&gt;What&apos;s the hardest review you&apos;ve been through, and what did it fix that the code didn&apos;t?&lt;/p&gt;</content:encoded><category>containerd</category><category>upstream</category></item><item><title>A node that reported Ready and started nothing: four parts, one containerd fix</title><link>https://www.catchkill9.dev/posts/containerd-part-1-the-bricked-node.html</link><guid isPermaLink="true">https://www.catchkill9.dev/posts/containerd-part-1-the-bricked-node.html</guid><description>From a 4.5 hour outage to a client-side timeout in containerd, through a merge, a revert and a second PR.</description><pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;h2 id=&quot;what-happened&quot;&gt;What happened&lt;/h2&gt;
&lt;p&gt;In July one of our Kubernetes nodes stopped starting pods for 4.5 hours. It reported Ready the entire time. The scheduler kept placing pods on it, every new container hung in creation, and nothing on the node&apos;s health path looked wrong.&lt;/p&gt;
&lt;p&gt;We run untrusted user code at Wix, one site per sandbox, so pods come and go all day. A node that silently stops starting them is worse than a node that dies: the scheduler keeps feeding it.&lt;/p&gt;
&lt;p&gt;The easy answer was a file descriptor limit. The snapshotter logs were full of &quot;too many open files&quot;. Raise the limit, restart, move on. I almost shipped that.&lt;/p&gt;
&lt;p&gt;This is the whole story in one place: the outage, the two bugs, the fix that got reverted, and the one that stays.&lt;/p&gt;
&lt;h2 id=&quot;why-it-happens&quot;&gt;Why it happens&lt;/h2&gt;
&lt;p&gt;Two bugs lined up.&lt;/p&gt;
&lt;p&gt;The first was in containerd. During snapshot garbage collection, containerd calls the snapshotter&apos;s Remove while holding that snapshotter&apos;s write lock in containerd&apos;s metadata layer, with no deadline. It even drops the caller&apos;s cancellation. If the snapshotter never answers, the lock is held forever, and every snapshot operation on the node queues behind it, including the one that starts a new pod sandbox. Nothing on the readiness path takes that lock, which is why the node kept saying Ready.&lt;/p&gt;
&lt;p&gt;Readiness is a separate path. Kubelet asks the runtime for its status, the runtime answers from memory, and nobody touches the snapshot metadata. So the node was healthy by every check we had, and useless by the only one that mattered: can it start a container.&lt;/p&gt;
&lt;p&gt;The second was the trigger, in AWS&apos;s soci-snapshotter, the plugin that lazily loads container images. Every file verified on first access opened a reader from the span cache and never closed it. The sibling function about 45 lines up closes the exact same reader. One leaked file descriptor per file, on every node running lazy-loaded images. Under a burst of pod starts the snapshotter hit its limit of 65,535 and stopped answering. containerd was waiting on it, holding the lock.&lt;/p&gt;
&lt;p&gt;A proxy snapshotter like soci runs in its own process, and containerd talks to it over gRPC. That boundary is where the real problem lived, though it took me two more PRs to see it.&lt;/p&gt;
&lt;p&gt;The goroutine dumps told the story once we knew where to look. Hundreds of goroutines parked on that snapshotter lock, and one GC goroutine holding it, waiting for a gRPC header from the snapshotter that was never coming. The node was not broken. It was waiting, politely and forever.&lt;/p&gt;
&lt;h2 id=&quot;what-we-got-wrong-first&quot;&gt;What we got wrong first&lt;/h2&gt;
&lt;p&gt;I got the diagnosis wrong twice before I got it right. After the storm the file descriptor count looked low, 151 out of 65,535, so I concluded the limit theory was wrong. Then I flip-flopped back. The leaked descriptors had been reclaimed once Go&apos;s finalizers ran, so the steady state hid the peak.&lt;/p&gt;
&lt;p&gt;I did most of that investigation with Claude, reading goroutine dumps from the wedged node and two unfamiliar codebases side by side. It didn&apos;t write the fix. It made the investigation cheap enough to do properly, including catching my own wrong conclusion.&lt;/p&gt;
&lt;p&gt;The bigger mistake came later. My first containerd fix put a timeout on the one call that hung on my nodes, the GC&apos;s Remove. It was correct for my incident. It got merged after 40 days and 30 review rounds, and I wrote about it as a win.&lt;/p&gt;
&lt;p&gt;Those 30 rounds were the real engineering. The timeout&apos;s config key got renamed twice, because in a project with hundreds of keys a name is an API commitment. The default went from 30 minutes to 5 and back to 30. I pushed hard for 5. A maintainer pushed back: the old behavior was unbounded, so on upgrade the least surprising default is the big one, and anyone running remote snapshotters can tune it down. He was thinking about thousands of clusters upgrading. I was thinking about mine.&lt;/p&gt;
&lt;p&gt;My test got rewritten around Go&apos;s synctest, so it runs in milliseconds on a fake clock instead of sleeping and mutating global state. And when CI wedged in the merge queue, Michael Brown restarted the buckets, herded reviewers and queued it again until it landed.&lt;/p&gt;
&lt;p&gt;Then I opened a second PR to cover another call that could hang the same way. Working through it with the maintainers, we saw that I was patching symptoms one call at a time. The thing that hangs is an external plugin, and it can hang on any of its calls. With a release about to be cut, Michael Brown reverted my first fix 13 days after it merged, so it never shipped. I agreed with the call.&lt;/p&gt;
&lt;p&gt;The suggestion that changed the plan came in that second PR&apos;s thread: stop bounding components one by one, and put the bound where containerd talks to the plugin. Michael agreed. Reverting my first fix was the price of doing it that way, because two mechanisms bounding the same call would have been worse than one.&lt;/p&gt;
&lt;h2 id=&quot;the-fix&quot;&gt;The fix&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;soci-snapshotter: close the reader in file Verify. One line, defer r.Close(), merged upstream in July.&lt;/li&gt;
&lt;li&gt;containerd, first attempt: bound the GC path&apos;s Remove with a configurable timeout. Merged, then reverted before release, because it fixed one call out of many.&lt;/li&gt;
&lt;li&gt;containerd, the fix that stays: a default_timeout setting on proxy plugins. When it is set, every call containerd makes to a proxy snapshotter gets that deadline if the incoming context has none. A caller that brings its own deadline keeps it. Leave it empty and nothing changes on upgrade.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The proxy snapshotter has 10 gRPC methods and none of them had a deadline unless the caller already set one. Putting the bound at the client covers all 10 at once, including the GC path my first fix targeted. It also covers paths I never hit in production, which is the point.&lt;/p&gt;
&lt;p&gt;Streaming calls needed a decision. Walk gets one deadline for the whole stream, not one per message, so time spent in callbacks counts against it. The content store proxy has the same gap but a different shape: a writer outlives the call that created it, so a per-call deadline would cut an ingest that is still being written. That one needs its own setting and its own PR.&lt;/p&gt;
&lt;p&gt;The review on the final PR was shorter, and it still taught me something. I had added an options argument to the exported constructor. Compatible, I thought. The maintainer pointed out it still breaks anyone who stores that function as a value, so the old constructor stays untouched and a new one takes the options.&lt;/p&gt;
&lt;p&gt;The SOCI side was the quick part. The PR went in on July 15 and AWS merged it five days later. The containerd side took three PRs and two months. The difference is not the size of the change. A one-line fix to an obvious leak has one right answer. A timeout in a runtime that thousands of clusters run has many plausible answers, and the review is where you find out which one the project can live with.&lt;/p&gt;
&lt;p&gt;It doesn&apos;t fix everything, and the PR says so. A GC pass over many slow snapshots can still add up while it holds the lock, because each call finishes just under the limit. That needs the lock itself reworked, and it is the next piece.&lt;/p&gt;
&lt;p&gt;Four posts, three PRs, one revert. The fix that finally lands is smaller than the one that got reverted, and it covers more. That is usually how it goes when the review is doing its job.&lt;/p&gt;
&lt;h2 id=&quot;lessons&quot;&gt;Lessons&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;A patch that is correct for your incident and a patch that is correct for everyone running the project are different things. The review is the distance between them.&lt;/li&gt;
&lt;li&gt;Review doesn&apos;t stop at merge. Mine got reverted, and the revert was the right engineering.&lt;/li&gt;
&lt;li&gt;Put the bound at the boundary. One deadline where containerd talks to the plugin covers every call; one deadline per call never ends.&lt;/li&gt;
&lt;li&gt;The steady state lies. Look at the peak, not the numbers after the storm.&lt;/li&gt;
&lt;li&gt;Maintainers do invisible work daily, for free, for strangers. Say so when they do it for you.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;credits-and-links&quot;&gt;Credits and links&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Michael Brown and the containerd maintainers, for 30 review rounds, a revert I agreed with, and a better fix&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/containerd/containerd/pull/13799&quot;&gt;containerd/containerd#13799&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/containerd/containerd/pull/14187&quot;&gt;containerd/containerd#14187&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/awslabs/soci-snapshotter/pull/2043&quot;&gt;awslabs/soci-snapshotter#2043&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded><category>containerd</category><category>kubernetes</category><category>upstream</category></item><item><title>Krew is K8s Game Changer!</title><link>https://www.catchkill9.dev/posts/krew-k8s-game-changer.html</link><guid isPermaLink="true">https://www.catchkill9.dev/posts/krew-k8s-game-changer.html</guid><description>The kubectl plugins that changed how I work with Kubernetes.</description><pubDate>Thu, 11 Jul 2024 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Kubernetes can be quite complex, but it doesn&apos;t have to slow you down. Here&apos;s how I simplified my K8s management with some awesome Krew plugins.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://krew.sigs.k8s.io/docs/user-guide/setup/install/&quot;&gt;installation guide for Krew&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://krew.sigs.k8s.io/plugins/&quot;&gt;list of all Krew plugins&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://artifacthub.io/packages/search&quot;&gt;alternate list of all plugins&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;discovering-krew-plugins&quot;&gt;Discovering Krew Plugins&lt;/h3&gt;
&lt;p&gt;When I first started using Kubernetes, I found it overwhelming to manage all the resources efficiently. That&apos;s when I stumbled upon &lt;strong&gt;Krew&lt;/strong&gt;, a package manager for kubectl plugins, which transformed my workflow. Here are the plugins that made a huge difference for me:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/ahmetb/kubectl-foreach/tree/master&quot;&gt;&lt;strong&gt;foreach&lt;/strong&gt;&lt;/a&gt; This plugin is a game-changer. Imagine being able to execute commands across multiple resources with just one line. It&apos;s perfect for bulk operations, saving me a ton of time when managing multiple pods or services.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href=&quot;https://github.com/corneliusweig/ketall&quot;&gt;get-all&lt;/a&gt;&lt;/strong&gt; Getting a comprehensive view of all resources in a namespace used to be a hassle. Not like using kubectl get all With get-all, I can actually retrieve every resource type in one go.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/vladimirvivien/ktop&quot;&gt;&lt;strong&gt;ktop&lt;/strong&gt;&lt;/a&gt; Monitoring cluster resources is crucial, and ktop makes it so much easier. It&apos;s like having a real-time dashboard for your cluster&apos;s health and performance. I check it regularly to ensure everything is running smoothly.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/nilic/kubectl-netshoot&quot;&gt;&lt;strong&gt;netshoot&lt;/strong&gt;&lt;/a&gt; This one really changed my life. it allows me to easily debug any networking issues, attaching debug container to any pod with ease.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/kvaps/kubectl-node-shell&quot;&gt;&lt;strong&gt;node-shell&lt;/strong&gt;&lt;/a&gt; Sometimes, you just need direct access to the nodes. Node-shell allows me to get a shell on any node in the cluster. It&apos;s invaluable for troubleshooting and performing maintenance tasks directly on the node. and doesn&apos;t require ssh,ssm,ec2-connect or any of those. it is a very privileged container and might cause your security team to call you to ask WTF?&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/stern/stern&quot;&gt;&lt;strong&gt;stern&lt;/strong&gt;&lt;/a&gt; Debugging applications can be tedious, especially when you need to tail logs from multiple pods. Stern makes this process seamless by streaming logs from all relevant pods in real-time. It&apos;s a lifesaver when tracking down issues across pods.&lt;/li&gt;
&lt;/ol&gt;</content:encoded><category>kubernetes</category><category>kubectl</category></item><item><title>Running kubectl commands on multiple clusters</title><link>https://www.catchkill9.dev/posts/running-kubectl-on-multiple-clusters.html</link><guid isPermaLink="true">https://www.catchkill9.dev/posts/running-kubectl-on-multiple-clusters.html</guid><description>One shell function that runs any kubectl command on every cluster.</description><pubDate>Mon, 01 Jul 2024 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Managing multiple Kubernetes clusters can be a real headache. I know firsthand how frustrating it can be to juggle multiple contexts, keep deployments consistent, and maintain stability across different environments.&lt;/p&gt;
&lt;p&gt;so I found this awesome repo: &lt;a href=&quot;https://github.com/ahmetb/kubectl-foreach&quot;&gt;https://github.com/ahmetb/kubectl-foreach&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Run a kubectl command in one or more contexts (clusters) in parallel (similar to GNU parallel/xargs).&lt;/p&gt;
&lt;h3 id=&quot;usage&quot;&gt;Usage&lt;/h3&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;background-color:#fff;--shiki-dark-bg:#24292e;color:#24292e;--shiki-dark:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;bash&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#6F42C1;--shiki-dark:#B392F0&quot;&gt;kubectl&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt; foreach&lt;/span&gt;&lt;span style=&quot;color:#24292E;--shiki-dark:#E1E4E8&quot;&gt; [OPTIONS] [PATTERN]... -- [KUBECTL_ARGS...]&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;and created a small helper function to use it. It&apos;s a simple way to run kubectl commands across multiple clusters in one go, making life a bit easier.&lt;/p&gt;
&lt;p&gt;Here&apos;s the function you can add to your bashrc or zshrc:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;background-color:#fff;--shiki-dark-bg:#24292e;color:#24292e;--shiki-dark:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;bash&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#D73A49;--shiki-dark:#F97583&quot;&gt;function&lt;/span&gt;&lt;span style=&quot;color:#6F42C1;--shiki-dark:#B392F0&quot;&gt; kfe&lt;/span&gt;&lt;span style=&quot;color:#24292E;--shiki-dark:#E1E4E8&quot;&gt;() {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#005CC5;--shiki-dark:#79B8FF&quot;&gt;  eval&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;kubectl foreach -q dc1 dc2 dc3 -- &lt;/span&gt;&lt;span style=&quot;color:#005CC5;--shiki-dark:#79B8FF&quot;&gt;$*&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt;&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#24292E;--shiki-dark:#E1E4E8&quot;&gt;}&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Just use kfe followed by your kubectl command, for example:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;background-color:#fff;--shiki-dark-bg:#24292e;color:#24292e;--shiki-dark:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;bash&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#6F42C1;--shiki-dark:#B392F0&quot;&gt;kfe&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt; get&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt; pods&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;and it will execute across dc1, dc2, and dc3 without any hassle. No more repetitive tasks or constant context switching, just a smoother workflow.&lt;/p&gt;
&lt;p&gt;I hope this helps lighten the load for anyone dealing with the complexities of Kubernetes. We&apos;re all in this together, and sharing little tools like this can make a big difference.&lt;/p&gt;</content:encoded><category>kubernetes</category><category>kubectl</category></item><item><title>Fixing EKS overprivileged cluster settings</title><link>https://www.catchkill9.dev/posts/fixing-eks-overprivileged-cluster-settings.html</link><guid isPermaLink="true">https://www.catchkill9.dev/posts/fixing-eks-overprivileged-cluster-settings.html</guid><description>The cluster creator gets admin by default. Now you can take it back without a rebuild.</description><pubDate>Sun, 28 Jan 2024 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Are you aware that your Amazon EKS cluster could be at risk due to a default permission setting? By default, the IAM role that creates an EKS cluster is granted the &lt;strong&gt;&lt;em&gt;system:master&lt;/em&gt;&lt;/strong&gt; RBAC role, providing broad administrative access. This could be a significant security concern.&lt;/p&gt;
&lt;p&gt;In the past, rectifying this issue was a cumbersome process, often requiring the recreation of the entire cluster. But now, AWS has streamlined this with new cluster access management APIs. With a simple command:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;background-color:#fff;--shiki-dark-bg:#24292e;color:#24292e;--shiki-dark:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;bash&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#6F42C1;--shiki-dark:#B392F0&quot;&gt;aws&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt; eks&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt; delete-access-entry&lt;/span&gt;&lt;span style=&quot;color:#005CC5;--shiki-dark:#79B8FF&quot;&gt; --cluster-name&lt;/span&gt;&lt;span style=&quot;color:#D73A49;--shiki-dark:#F97583&quot;&gt; &amp;#x3C;&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt;CLUSTER_NAM&lt;/span&gt;&lt;span style=&quot;color:#24292E;--shiki-dark:#E1E4E8&quot;&gt;E&lt;/span&gt;&lt;span style=&quot;color:#D73A49;--shiki-dark:#F97583&quot;&gt;&gt;&lt;/span&gt;&lt;span style=&quot;color:#005CC5;--shiki-dark:#79B8FF&quot;&gt; --principal-arn&lt;/span&gt;&lt;span style=&quot;color:#D73A49;--shiki-dark:#F97583&quot;&gt; &amp;#x3C;&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt;IAM_PRINCIPAL_AR&lt;/span&gt;&lt;span style=&quot;color:#24292E;--shiki-dark:#E1E4E8&quot;&gt;N&lt;/span&gt;&lt;span style=&quot;color:#D73A49;--shiki-dark:#F97583&quot;&gt;&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;you can adjust these permissions efficiently without the need to recreate the cluster.&lt;/p&gt;
&lt;p&gt;This development is a game-changer in EKS cluster security management, aligning with the principle of least privilege access. It&apos;s crucial for EKS users to be aware of these changes and implement them to ensure their clusters are secure.&lt;/p&gt;
&lt;p&gt;For more detailed steps and insights on this topic, make sure to read the full AWS blog post &lt;a href=&quot;https://aws.amazon.com/blogs/containers/a-deep-dive-into-simplified-amazon-eks-access-management-controls/&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;</content:encoded><category>aws</category><category>eks</category><category>security</category></item><item><title>My first open source PR was 3 lines</title><link>https://www.catchkill9.dev/posts/first-oss-pr-apparmor.html</link><guid isPermaLink="true">https://www.catchkill9.dev/posts/first-oss-pr-apparmor.html</guid><description>An unhandled AppArmor case, millions of log lines, and a CLA nobody at work could pick.</description><pubDate>Mon, 04 Dec 2023 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;We ran the &lt;a href=&quot;https://github.com/kubernetes-sigs/security-profiles-operator&quot;&gt;Security Profiles Operator&lt;/a&gt; to ship AppArmor profiles to our Kubernetes nodes. It worked. It also logged the same error over and over:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;background-color:#fff;--shiki-dark-bg:#24292e;color:#24292e;--shiki-dark:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;text&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;getting owner profile: the node status owner is of an unknown kind&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Millions of lines of it. The profiles still loaded, so nothing was broken. Just loud.&lt;/p&gt;
&lt;p&gt;The operator keeps a status object per node for every profile, and a reconciler looks up which profile owns it. That lookup was a &lt;code&gt;switch&lt;/code&gt; on the owner&apos;s kind. Seccomp was in it. SELinux was in it, twice. AppArmor wasn&apos;t.&lt;/p&gt;
&lt;p&gt;So on 30 November 2023 I opened &lt;a href=&quot;https://github.com/kubernetes-sigs/security-profiles-operator/issues/1993&quot;&gt;issue #1993&lt;/a&gt; and asked three honest questions: is this important, does it mean something is wrong, and can someone guide me to a fix. An hour later I sent the fix myself as &lt;a href=&quot;https://github.com/kubernetes-sigs/security-profiles-operator/pull/1994&quot;&gt;PR #1994&lt;/a&gt;:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;background-color:#fff;--shiki-dark-bg:#24292e;color:#24292e;--shiki-dark:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;go&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#D73A49;--shiki-dark:#F97583&quot;&gt;	case&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;AppArmorProfile&quot;&lt;/span&gt;&lt;span style=&quot;color:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#24292E;--shiki-dark:#E1E4E8&quot;&gt;		prof &lt;/span&gt;&lt;span style=&quot;color:#D73A49;--shiki-dark:#F97583&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;color:#D73A49;--shiki-dark:#F97583&quot;&gt; &amp;#x26;&lt;/span&gt;&lt;span style=&quot;color:#6F42C1;--shiki-dark:#B392F0&quot;&gt;apparmorapi&lt;/span&gt;&lt;span style=&quot;color:#24292E;--shiki-dark:#E1E4E8&quot;&gt;.&lt;/span&gt;&lt;span style=&quot;color:#6F42C1;--shiki-dark:#B392F0&quot;&gt;AppArmorProfile&lt;/span&gt;&lt;span style=&quot;color:#24292E;--shiki-dark:#E1E4E8&quot;&gt;{}&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Plus the import. Three lines.&lt;/p&gt;
&lt;p&gt;Writing it took minutes. Getting it merged took 4 days, and none of the delay was the fix itself:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The linter failed on import order. Then on a whitespace change I couldn&apos;t see. My actual comment on the PR was &quot;I have no idea what fails, can any1 explain?&quot;&lt;/li&gt;
&lt;li&gt;The CLA bot blocked it. Corporate or individual? Nobody on our internal Slack knew which one I should sign, so the PR sat until someone decided.&lt;/li&gt;
&lt;li&gt;The e2e tests needed a maintainer to type &lt;code&gt;/ok-to-test&lt;/code&gt; and &lt;code&gt;/retest&lt;/code&gt; for me, more than once.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Sascha Grunert did all of that: pasted the exact linter diff, retriggered the tests, approved it. It merged on 4 December.&lt;/p&gt;
&lt;p&gt;The lesson I took from it: a first PR is mostly about learning the project&apos;s machinery, the linter, the CLA, the bots, the slash commands. The code is the small part.&lt;/p&gt;</content:encoded><category>kubernetes</category><category>security</category><category>upstream</category></item><item><title>Awesome K8s Using Bash</title><link>https://www.catchkill9.dev/posts/awesome-k8s-using-bash.html</link><guid isPermaLink="true">https://www.catchkill9.dev/posts/awesome-k8s-using-bash.html</guid><description>Republished from Medium.</description><pubDate>Mon, 27 Nov 2023 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;stop suffering when using kubectl&lt;/p&gt;
&lt;h2 id=&quot;introduction&quot;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;Efficient Kubernetes management often entails repetitive kubectl commands. To streamline this, I&apos;ve developed a set of Bash functions, collectively accessible via a &lt;a href=&quot;https://gist.github.com/NahumLitvin/e8e67bb33a0fc0f8f61b28c64303aea1&quot;&gt;gist&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;All those functions will find a pod by partial name: &lt;em&gt;partial_pod_name&lt;/em&gt; and run common k8s commands on the found pod.&lt;/p&gt;
&lt;p&gt;we effectively scan through the pods in the current namespace and retrieves the name of the first pod that matches a given partial string. Those functions becomes particularly handy when dealing with dynamically generated pod names, where you typically know the deployment name but not the unique identifier of each pod.&lt;/p&gt;
&lt;h2 id=&quot;the-functions-and-theirusage&quot;&gt;The Functions and Their Usage&lt;/h2&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;background-color:#fff;--shiki-dark-bg:#24292e;color:#24292e;--shiki-dark:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;bash&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#6F42C1;--shiki-dark:#B392F0&quot;&gt;podNameSeeker&lt;/span&gt;&lt;span style=&quot;color:#24292E;--shiki-dark:#E1E4E8&quot;&gt; [partial_pod_name] &lt;/span&gt;&lt;span style=&quot;color:#6A737D;--shiki-dark:#6A737D&quot;&gt;# find pod by partial name&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#6F42C1;--shiki-dark:#B392F0&quot;&gt;shpod&lt;/span&gt;&lt;span style=&quot;color:#24292E;--shiki-dark:#E1E4E8&quot;&gt; [partial_pod_name] &lt;/span&gt;&lt;span style=&quot;color:#6A737D;--shiki-dark:#6A737D&quot;&gt;# Access a pod (by partial name) using /bin/sh&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#6F42C1;--shiki-dark:#B392F0&quot;&gt;bashpod&lt;/span&gt;&lt;span style=&quot;color:#24292E;--shiki-dark:#E1E4E8&quot;&gt; [partial_pod_name] &lt;/span&gt;&lt;span style=&quot;color:#6A737D;--shiki-dark:#6A737D&quot;&gt;# Access a pod (by partial name) using /bin/bash&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#6F42C1;--shiki-dark:#B392F0&quot;&gt;commandpod&lt;/span&gt;&lt;span style=&quot;color:#24292E;--shiki-dark:#E1E4E8&quot;&gt; [partial_pod_name] [command] &lt;/span&gt;&lt;span style=&quot;color:#6A737D;--shiki-dark:#6A737D&quot;&gt;# Execute a command inside a pod (by partial name).&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#6F42C1;--shiki-dark:#B392F0&quot;&gt;logpod&lt;/span&gt;&lt;span style=&quot;color:#24292E;--shiki-dark:#E1E4E8&quot;&gt; [partial_pod_name] &lt;/span&gt;&lt;span style=&quot;color:#6A737D;--shiki-dark:#6A737D&quot;&gt;# Fetch logs from a pod (by partial name).&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#6F42C1;--shiki-dark:#B392F0&quot;&gt;debugbashpod&lt;/span&gt;&lt;span style=&quot;color:#D73A49;--shiki-dark:#F97583&quot;&gt; &amp;#x3C;&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt;partial_pod_nam&lt;/span&gt;&lt;span style=&quot;color:#24292E;--shiki-dark:#E1E4E8&quot;&gt;e&lt;/span&gt;&lt;span style=&quot;color:#D73A49;--shiki-dark:#F97583&quot;&gt;&gt;&lt;/span&gt;&lt;span style=&quot;color:#6A737D;--shiki-dark:#6A737D&quot;&gt; #attach netshoot debug container to pod and bash into it&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#6F42C1;--shiki-dark:#B392F0&quot;&gt;debugcommandpod&lt;/span&gt;&lt;span style=&quot;color:#D73A49;--shiki-dark:#F97583&quot;&gt; &amp;#x3C;&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt;partial_pod_nam&lt;/span&gt;&lt;span style=&quot;color:#24292E;--shiki-dark:#E1E4E8&quot;&gt;e&lt;/span&gt;&lt;span style=&quot;color:#D73A49;--shiki-dark:#F97583&quot;&gt;&gt;&lt;/span&gt;&lt;span style=&quot;color:#D73A49;--shiki-dark:#F97583&quot;&gt; &amp;#x3C;&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt;comman&lt;/span&gt;&lt;span style=&quot;color:#24292E;--shiki-dark:#E1E4E8&quot;&gt;d&lt;/span&gt;&lt;span style=&quot;color:#D73A49;--shiki-dark:#F97583&quot;&gt;&gt;&lt;/span&gt;&lt;span style=&quot;color:#6A737D;--shiki-dark:#6A737D&quot;&gt; # attach netshoot debug container to pod and run command in it&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h2 id=&quot;the-debug-functions&quot;&gt;The Debug Functions&lt;/h2&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/nicolaka/netshoot&quot;&gt;netshoot&lt;/a&gt; is an awesome project , a container with all the network tools you will ever need. iperf , nc, curl, wget and more. you can use kubectl debug to attach and ephemeral container with it to your pod to debug it see &lt;a href=&quot;https://kubernetes.io/docs/tasks/debug/debug-application/debug-running-pod/#ephemeral-container&quot;&gt;Debugging with an ephemeral debug container&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;think about a container named &quot;bob&quot; trying to access &quot;myserver.address.com&quot; on port &quot;myport&quot; and falining but you have no networking tools installed on the pod&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;background-color:#fff;--shiki-dark-bg:#24292e;color:#24292e;--shiki-dark:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;bash&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#6F42C1;--shiki-dark:#B392F0&quot;&gt;debugcommandpod&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt; bob&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt; nc&lt;/span&gt;&lt;span style=&quot;color:#005CC5;--shiki-dark:#79B8FF&quot;&gt; -vz&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt; myserver.address.com&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt; myport&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;will attach netshoot to the bob container and run NC to myserver.address.com and check if myport is open&lt;/p&gt;
&lt;h2 id=&quot;the-codehttpsmediumcommedia9dd483e9b2a7fb251de1a40bc598d401href&quot;&gt;The Code&lt;a href=&quot;https://medium.com/media/9dd483e9b2a7fb251de1a40bc598d401/href&quot;&gt;https://medium.com/media/9dd483e9b2a7fb251de1a40bc598d401/href&lt;/a&gt;&lt;/h2&gt;
&lt;h2 id=&quot;integrating-into-yourshell&quot;&gt;Integrating into Your Shell&lt;/h2&gt;
&lt;p&gt;To add these functions to your shell environment, you can run the following command in your terminal:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;background-color:#fff;--shiki-dark-bg:#24292e;color:#24292e;--shiki-dark:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;bash&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#6F42C1;--shiki-dark:#B392F0&quot;&gt;curl&lt;/span&gt;&lt;span style=&quot;color:#005CC5;--shiki-dark:#79B8FF&quot;&gt; -sL&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt; https://gist.githubusercontent.com/NahumLitvin/e8e67bb33a0fc0f8f61b28c64303aea1/raw/d35e95e1c408c382739258dd7048bb866277946e/podseeker.sh&lt;/span&gt;&lt;span style=&quot;color:#D73A49;--shiki-dark:#F97583&quot;&gt; &gt;&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt; ~/podseeker.sh&lt;/span&gt;&lt;span style=&quot;color:#24292E;--shiki-dark:#E1E4E8&quot;&gt; &amp;#x26;&amp;#x26; &lt;/span&gt;&lt;span style=&quot;color:#005CC5;--shiki-dark:#79B8FF&quot;&gt;\&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#005CC5;--shiki-dark:#79B8FF&quot;&gt;echo&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &apos;[ -f ~/podseeker.sh ] &amp;#x26;&amp;#x26; source ~/podseeker.sh&apos;&lt;/span&gt;&lt;span style=&quot;color:#D73A49;--shiki-dark:#F97583&quot;&gt; &gt;&gt;&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt; ~/.&lt;/span&gt;&lt;span style=&quot;color:#24292E;--shiki-dark:#E1E4E8&quot;&gt;${SHELL&lt;/span&gt;&lt;span style=&quot;color:#D73A49;--shiki-dark:#F97583&quot;&gt;:&lt;/span&gt;&lt;span style=&quot;color:#24292E;--shiki-dark:#E1E4E8&quot;&gt;t}&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt;rc&lt;/span&gt;&lt;span style=&quot;color:#24292E;--shiki-dark:#E1E4E8&quot;&gt; &amp;#x26;&amp;#x26; &lt;/span&gt;&lt;span style=&quot;color:#005CC5;--shiki-dark:#79B8FF&quot;&gt;\&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#005CC5;--shiki-dark:#79B8FF&quot;&gt;source&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt; ~/.&lt;/span&gt;&lt;span style=&quot;color:#24292E;--shiki-dark:#E1E4E8&quot;&gt;${SHELL&lt;/span&gt;&lt;span style=&quot;color:#D73A49;--shiki-dark:#F97583&quot;&gt;:&lt;/span&gt;&lt;span style=&quot;color:#24292E;--shiki-dark:#E1E4E8&quot;&gt;t}&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt;rc&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This one-liner will append the functions to your .bashrc or .zshrc and reload the configuration.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Conclusion&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;These custom Bash functions are designed to make Kubernetes pod management more efficient and user-friendly. By incorporating them into your workflow, you can reduce time spent on routine commands and focus on more complex tasks.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://medium.com/_/stat?event=post.clientViewed&amp;#x26;referrerSource=full_rss&amp;#x26;postId=59a767d441e4&quot; alt=&quot;&quot;&gt;&lt;/p&gt;</content:encoded><category>kubernetes</category><category>ssh</category><category>bash</category></item><item><title>Effortless Git: Introducing compush, a One-Stop Git Command</title><link>https://www.catchkill9.dev/posts/compush-one-stop-git-command.html</link><guid isPermaLink="true">https://www.catchkill9.dev/posts/compush-one-stop-git-command.html</guid><description>Republished from Medium.</description><pubDate>Sun, 19 Nov 2023 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;
&lt;p&gt;&quot;One command to add them all, one command to commit them,
One command to push them all, and in the remote branch compush them;
In the Land of Bash, where the Git repos lie.&quot;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;tldr&quot;&gt;TL;DR&lt;/h2&gt;
&lt;p&gt;&quot;One command to rule them all: Add, commit, and push in a single stroke. Simply integrate it into your bash configuration, and embark on your coding quest with ease.&quot;&lt;/p&gt;
&lt;h3 id=&quot;what-iscompush&quot;&gt;What is compush?&lt;/h3&gt;
&lt;p&gt;compush is a clever shell function that transforms the usual multi-step process of staging, committing, and pushing changes into a single, streamlined operation. Its standout feature is its ability to work from any subdirectory within your Git repository, eliminating the hassle of navigating back to the root directory. This functionality is a real time-saver, allowing you to concentrate more on coding and less on command-line intricacies.&lt;/p&gt;
&lt;h3 id=&quot;how-compush-enhances-yourworkflow&quot;&gt;How compush Enhances Your Workflow&lt;/h3&gt;
&lt;p&gt;When you call compush into action, it kicks off a series of operations:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Finds the Git Repository&apos;s Root: It intelligently locates the top-level directory of your Git repository, or uses the current directory if it&apos;s not in a Git repo.&lt;/li&gt;
&lt;li&gt;Stages Changes: Automatically adds all changes from the top-level directory to the staging area.&lt;/li&gt;
&lt;li&gt;Commits Changes: Commits the staged changes, using either your provided commit message or an empty message if none is specified.&lt;/li&gt;
&lt;li&gt;Pushes to Remote: Effortlessly pushes the commit to the origin remote of the current branch and sets up tracking.&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id=&quot;setting-upcompush&quot;&gt;Setting Up &lt;strong&gt;compush&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;To incorporate &lt;strong&gt;compush&lt;/strong&gt; into your workflow, add this function to your shell configuration file (.zshrc for Zsh users or .bashrc for Bash users):&lt;a href=&quot;https://medium.com/media/e6e7c5168b04ddaa25f18e4f9ffba6fe/href&quot;&gt;https://medium.com/media/e6e7c5168b04ddaa25f18e4f9ffba6fe/href&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Activating compushTo activate compush, source your configuration file:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;For Zsh: source ~/.zshrc&lt;/li&gt;
&lt;li&gt;For Bash: source ~/.bashrc&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Using **compush **, within any directory of your Git repository, execute:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;background-color:#fff;--shiki-dark-bg:#24292e;color:#24292e;--shiki-dark:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;bash&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#6F42C1;--shiki-dark:#B392F0&quot;&gt;comush&lt;/span&gt;&lt;span style=&quot;color:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;Your insightful commit message here&quot;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;To a commit without a message:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;background-color:#fff;--shiki-dark-bg:#24292e;color:#24292e;--shiki-dark:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;bash&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#6F42C1;--shiki-dark:#B392F0&quot;&gt;compush&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;img src=&quot;https://medium.com/_/stat?event=post.clientViewed&amp;#x26;referrerSource=full_rss&amp;#x26;postId=6188197563e4&quot; alt=&quot;&quot;&gt;&lt;/p&gt;</content:encoded><category>git</category><category>devops</category></item><item><title>Git Takes Sha Instead Of Remote Branch</title><link>https://www.catchkill9.dev/posts/git-takes-sha-instead-of-remote-branch.html</link><guid isPermaLink="true">https://www.catchkill9.dev/posts/git-takes-sha-instead-of-remote-branch.html</guid><description>Republished from Medium.</description><pubDate>Thu, 26 Nov 2020 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Was debugging a weird production issue. the version deploy was absolutely wrong.&lt;/p&gt;
&lt;h2 id=&quot;premise&quot;&gt;Premise:&lt;/h2&gt;
&lt;p&gt;When our CI build the product it creates a branch. with the build number. in our CI it was a running integer but in the &lt;a href=&quot;https://github.com/NahumLitvin/GitTakesShaInsteadOfRemoteBranch&quot;&gt;repo&lt;/a&gt; I created for the example it will be &lt;a href=&quot;https://github.com/NahumLitvin/GitTakesShaInsteadOfRemoteBranch/tree/2c4f227&quot;&gt;2c4f227&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;by coincidence it was starting the same way as a random commit on the main branch.&lt;/p&gt;
&lt;h3 id=&quot;on-the-deploy-stage-we-clone-the-repo-on-main-branch-bydefault&quot;&gt;on the deploy stage we clone the repo (on main branch by default)&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;git clone &lt;a href=&quot;mailto:git@github.com&quot;&gt;git@github.com&lt;/a&gt;:NahumLitvin/GitTakesShaInsteadOfRemoteBranch.git&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;git checkout 2c4f227&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3 id=&quot;and-thats-the-result-igot&quot;&gt;and thats the result I got:&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Note: checking out &apos;2c4f227&apos;.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;You are in &apos;detached HEAD&apos; state. You can look around, make experimental
changes and commit them, and you can discard any commits you make in this
state without impacting any branches by performing another checkout.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;If you want to create a new branch to retain commits you create, you may
do so (now or later) by using -b with the checkout command again. Example:&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;git checkout -b &lt;new-branch-name&gt;&lt;/new-branch-name&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;HEAD is now at &lt;em&gt;&lt;strong&gt;&lt;em&gt;2c4f227&lt;/em&gt;&lt;/strong&gt;&lt;/em&gt; Initial commit&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;git-checked-out-the-commit-_2c4f227-instead-of-the-branch_2c4f227&quot;&gt;&lt;strong&gt;git checked out the commit&lt;/strong&gt; &lt;strong&gt;_2c4f227 instead of the branch _&lt;/strong&gt;2c4f227!&lt;/h2&gt;
&lt;h2 id=&quot;solution&quot;&gt;Solution:&lt;/h2&gt;
&lt;blockquote&gt;
&lt;p&gt;git checkout origin/2c4f227&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;will checkout the correct branch&lt;/p&gt;
&lt;p&gt;another more concise way is to use the new &lt;a href=&quot;https://git-scm.com/docs/git-switch&quot;&gt;git switch&lt;/a&gt; command available from git 2.29.0&lt;/p&gt;
&lt;h2 id=&quot;git-switch&quot;&gt;git switch&lt;/h2&gt;
&lt;blockquote&gt;
&lt;p&gt;git switch 2c4f227
warning: refname &apos;2c4f227&apos; is ambiguous.
Switched to branch &apos;2c4f227&apos;
Your branch is up to date with &apos;origin/2c4f227&apos;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a href=&quot;https://git-scm.com/docs/git-switch&quot;&gt;&lt;strong&gt;git switch&lt;/strong&gt;&lt;/a&gt; will _&apos;Switch to a specified branch. The working tree and the index are updated to match the branch. All new commits will be added to the tip of this branch.&apos;
&lt;em&gt;&lt;strong&gt;when: **
&lt;a href=&quot;https://git-scm.com/docs/git-checkout&quot;&gt;&lt;strong&gt;git checkout&lt;/strong&gt;&lt;/a&gt; _&apos;Updates files in the working tree to match the version in the index or the specified tree. If no pathspec was given, _&lt;/strong&gt;&lt;em&gt;git checkout&lt;/em&gt;**&lt;/em&gt; will also update _&lt;em&gt;HEAD to set the specified branch as the current branch.&apos;&lt;/em&gt;&lt;/p&gt;
&lt;h2 id=&quot;additional-reading&quot;&gt;additional reading&lt;/h2&gt;
&lt;p&gt;see this stack overflow Q&amp;#x26;A for a better explanation: &lt;a href=&quot;https://stackoverflow.com/questions/57123031/git-checkout-commit-hash-vs-git-checkout-branch/&quot;&gt;https://stackoverflow.com/questions/57123031/git-checkout-commit-hash-vs-git-checkout-branch/&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://medium.com/_/stat?event=post.clientViewed&amp;#x26;referrerSource=full_rss&amp;#x26;postId=a905ab866559&quot; alt=&quot;&quot;&gt;&lt;/p&gt;</content:encoded><category>git</category><category>software-development</category><category>devops</category></item></channel></rss>