When GitOps Meets Emergency Fixes: ArgoCD Operational Lessons
GitOps sounds pristine on paper. You define your entire infrastructure and deployments in Git, your pipelines handle the rest, and tools like ArgoCD keep everything in your cluster aligned to the version-controlled truth. It's the kind of DevOps dream that gets keynote time at every cloud-native conference.
Real-world infrastructure is a graveyard of trade-offs, late-night alerts, and fire drills, though. And when GitOps breaks bad, usually at 2AM, the person waking up is rarely the senior engineer giving the talk at KubeCon. It's the junior SRE sweating over kubectl edit, hoping their changes stick before ArgoCD notices.
This is painfully, hilariously real, and the internet has stories.
GitOps vs. prod fires: who wins?
One of the more common frustrations with GitOps workflows is that they're rigid by design. That rigidity is the point, since we want to eliminate manual intervention, reduce drift, and ensure repeatability. But sometimes, repeatability is the enemy.
Imagine this: production is on fire, the pipeline is slow, and approval queues are jammed. Meanwhile ArgoCD, the loyal enforcer of your declared state, keeps reverting your emergency fixes because they weren't done "the GitOps way."
That's what happened to one junior SRE who got paged in the middle of the night. Faced with a deployment issue, they used kubectl edit to try and patch things up, only to watch ArgoCD reset their changes every few minutes. To make it worse, the only person with access to the ArgoCD platform was out. The result was eight hours of downtime because nobody could stop the automation from undoing the emergency work.
It's far from a one-off story. This happens more often than most teams want to admit.
Drift detection: savior or saboteur?
The idea behind drift detection is noble: any changes made outside of Git should be flagged or reverted. In emergency scenarios, though, it becomes a double-edged sword. As one engineer put it, "Sometimes you just gotta put the fire out."
To do that, many teams have to either temporarily suspend drift detection, fix the issue, and re-enable it after the PR gets merged (if it ever does), or hack their way into the system with breakglass permissions (if they exist at all).
Even then, you're racing the ArgoCD reconciliation loop. In some setups, it'll undo your fix in milliseconds. In others, you might have a three-minute window before it reverts your cluster back to broken.
The ArgoCD catch-22
ArgoCD deserves its own section, because it shows up in almost every one of these war stories.
It's a solid tool that does what it promises. But it doesn't care that your approval chain is asleep or that your CI pipeline takes forever. It will enforce the last declared state, no matter how out of date it is.
This has led to some... creative solutions:
- Turning off auto-sync by editing the ArgoCD application object manually (kubectl edit app argocd -n argocd)
- Redirecting ArgoCD to a temporary PR branch to bypass the slow merge process
- Just deleting the ArgoCD server pod entirely, hoping that when it comes back, it's learned its lesson (spoiler: it hasn't)
Nobody would call these best practices; they're desperation tactics. But when prod is down, elegance takes a backseat to uptime.
Why do juniors have prod access?
One of the most hotly debated aspects of this mess is access control. Why does a junior SRE have production permissions in the first place? Is this a badge of trust, or a sign that your on-call rotation needs serious restructuring?
Opinions vary. Some believe juniors should never be first responders for critical incidents. Others argue that incident response is where you really learn, provided there's a senior available to shadow or lead.
Often the access itself is fine and the missing support structure is the problem. Giving a junior SRE keys to prod without guardrails, guidance, or a clear breakglass process is negligent, whatever the intent to empower them.
What a good breakglass setup actually looks like
Breakglass access is an uncontroversial idea that is supposed to exist for situations exactly like this, but it has to be designed right. Every use should be logged and reviewed, and access should expire automatically. Every use must come with a reason, and often a postmortem. And it has to be accessible: it shouldn't require waking five people or waiting an hour to use.
If your "emergency access" takes longer to unlock than the PR pipeline, you've built a trap instead of a solution.
One team described using AWS IAM roles where you can assume a special "breakglass" role via SSO, but only after jumping through an extra confirmation step. It's quick and visible, and it forces engineers to think twice before making changes without blocking them when minutes matter.
Is GitOps actually the problem?
A lot of these issues come from bad implementation more than from GitOps itself.
If your GitOps flow is too slow to be useful during an incident, the pipeline design has failed, and the paradigm is fine. If you're relying on a single engineer for ArgoCD access, that's a people/process problem and has nothing to do with tool limitations.
And if engineers feel they have to break process just to fix things quickly, it probably means your process wasn't built for real-world chaos in the first place.
Lessons from the front lines
The engineers in these situations are resourceful, far from clueless. They know that saving the system sometimes means bending the rules, and they're also frustrated that they have to.
Here's what teams could actually do to reduce GitOps pain in production:
- Document the emergency playbook. Don't assume everyone knows how to pause ArgoCD or reroute a sync. Write it down.
- Create fast paths for prod hotfixes. If your PR takes hours to merge, what you have is continuous frustration rather than continuous deployment.
- Set up webhooks in ArgoCD. Stop relying on polling every 3 minutes. Push mode is faster, less confusing, and more efficient.
- Rotate access. Don't bottleneck critical permissions with one person. If you need someone to edit Argo at 2AM, they shouldn't be unreachable.
Most importantly, have postmortems. If something breaks bad, make sure the next 2AM responder has more than hope to work with.
The bottom line
GitOps isn't broken, though your team might be.
Automation is only as good as the humans behind it, and when the humans are tired, under-trained, or unsupported, the best tools in the world can't save you.
So yes, ArgoCD will auto-heal. It won't heal your broken processes, your lack of documentation, or the absence of a support net for your most junior engineers. Fix those, and GitOps stops breaking bad and just works.