Editing in Prod: A Love Letter to Every SRE Who's Ever Broken Glass
Some panic only shows up at 2AM. The alerts start blaring, Slack lights up, and somewhere in the shadows of your terminal, a Kubernetes deployment is eating itself alive. You know the drill: open Git, file a PR, wait for approval, let ArgoCD sync, maybe light a candle for good luck. But the fire's already spreading, and tonight you're not waiting. You kubectl edit in prod, and guess what? That's okay.
In the world of GitOps, where the infrastructure is code and deployments are supposed to be pristine, repeatable rituals, there's an unspoken understanding that sometimes you've gotta break the glass. You break it because the system didn't account for reality, the one where people sleep, where the only engineer with access to ArgoCD is unreachable, and where the sync policy you carefully set up is now undoing your fixes faster than you can type them in.
The DevOps dream meets the 2AM reality
GitOps is the holy grail of modern Kubernetes management. It promises consistency, auditability, and self-healing infrastructure. Tools like ArgoCD and Flux continuously watch Git repos, ensuring what's running matches what's version-controlled. Sounds great, until the person who can approve your hotfix is halfway through REM sleep and your cluster is screaming.
One user summed it up best: "Edit in prod while you wait for the PR to get approved. Sometimes you just gotta put the fire out." That's the core conflict, a dreamy CI/CD pipeline versus the cold, lonely battlefield of real-world ops. When it's your name on the pager, ideals don't extinguish incidents.
The sync loop under pressure
ArgoCD, for all its beauty, can feel like a vengeful ghost when you're trying to make an emergency fix. It detects your manual changes and gleefully reverts them within milliseconds, because it's doing exactly what it was told. "With ArgoCD set up to autoheal," one engineer wrote, "you can edit manually as often as you want. It will always go back."
It's like playing whack-a-mole with your own YAML files.
The clever ones know the workaround: disable auto-sync temporarily. Even that assumes you've got the right access, which in more than one horror story is locked down to a single senior engineer who happens to be unreachable. One poor soul watched ArgoCD undo their kubectl changes for eight straight hours because no one else could stop the sync. That's a systems failure, and a frustrating one.
The trouble is rigid process
The tension is between rigid processes and the humans they're supposed to support. Nobody's saying you should be cowboy coding in prod every day. But when your options are "wait for the PR to merge" or "lose customer data," suddenly kubectl edit doesn't look so evil.
A principal SRE chimed in with wisdom: "Access to prod should require a breakglass account. Not something onerous, just monitored, logged, and requiring a postmortem." That's the balance: make it easy to act, but hard to forget. You shouldn't need a prayer and a Slack rant to do your job.
The bigger crime is building systems that leave your junior SRE holding the pager alone at 2AM without support or tools.
Cowboy culture vs. guardrails that work
A lot of orgs are still clawing their way out of cowboy DevOps. You know the type: no approvals, no audits, just vibes and root access. Then they swing the other way, wrapping every action in red tape and mandatory sign-offs that don't scale under pressure.
The healthiest teams build for both. They assume incidents will happen and give engineers safe, documented, reversible paths to act fast. That might mean toggling auto-sync off, pointing ArgoCD at a temporary patch branch, or (gasp) doing a manual edit with clear rollback instructions. What matters is that it's a deliberate choice.
GitOps, but make it human
One of the more nuanced takes from the field: "GitOps is not just Git pushing to the cluster. It's also reconciliation, automation, and, when needed, control." What the GitOps purists sometimes miss is that the best infra is designed for the humans who run it.
That means making room for controlled chaos, and understanding that self-healing can be self-defeating if you don't also have self-awareness. It also means treating "edit in prod" as a signal that your system needs a better escape hatch.
A love letter, with logging
So here's to the ones who stayed up: the juniors who got paged because nobody else could, the seniors who gave them the tools and trust to act, and the ones who disabled sync, made the fix, then wrote the postmortem that taught everyone what to do next time. You're contingency plans in action.
Keep your kubectl handy, just don't forget to tell Git about it when the fire's out.