2,000+ Service Accounts, No Owners: The Multi-Cloud IAM Problem
Large cloud environments produce their own particular kind of panic. It doesn't happen on day one, or even in year one. It happens around year three, after the third acquisition and the fifth Kubernetes cluster, when someone asks, "Wait… how many service accounts do we actually have?"
In this case the answer was 2,000+ machine identities across AWS, Azure, and GCP. There was no centralized inventory and no consistent rotation policy, just a mix of IAM roles, service principals, workload identities, and Kubernetes pod identities. That's a governance time bomb, well beyond "a little messy."
And the constraints are brutal:
- Automated discovery of machine identities
- Rotation without app downtime
- Least privilege recommendations based on actual usage
- CI/CD integration (Jenkins, GitHub Actions)
- API-first architecture
- No agents
- No code changes
- No six-month science projects
This is also explicitly about machine identity lifecycle at scale, with PAM for humans out of scope. Most enterprises stall out at this point without saying so, which makes it worth working through what's actually realistic.
There is no magic unified button
You've already seen it in your evaluations. CyberArk is powerful and expensive, and it feels like bringing a tank to a knife fight.
HashiCorp Vault is solid, but someone has to run it: HA clusters, storage backends, secret engines, and policy drift add up to work for a whole team.
Cloud-native secrets managers like AWS Secrets Manager and Azure Key Vault are functional but fragmented, and you end up managing three different control planes. Fragmentation is exactly what got you here.
The industry still hasn't produced a clean, vendor-neutral "multi-cloud machine identity governor" that requires zero operational lift. So what are enterprises actually doing? They're solving it with architecture, and no single product does the job.
The pattern that's winning: kill static credentials
Look at the strongest signal in the discussion:
Workload identity + federation for cross-cloud calls. Avoid storing keys or static credentials.
That's a strategy more than a product recommendation. Instead of rotating long-lived access keys, tracking secret sprawl, and managing password lifecycle, teams are eliminating the problem entirely.
In Kubernetes, this usually means using OIDC federation, letting workloads assume roles dynamically, and exchanging short-lived tokens across clouds. AWS supports this with IAM roles for service accounts, GCP supports workload identity federation, and Azure supports federated credentials for service principals.
When you wire them together correctly, pods don't store credentials. They request them, the credentials expire, and they refresh automatically. There's no rotation downtime, no agent, and no code changes if you're already using SDK-default credential providers. That's why this approach scales.
But what about discovery?
This is where things get real. You don't have centralized inventory, and before you optimize rotation, you need visibility. Enterprises are handling this in three main ways.
1. Cloud-native inventory + aggregation
Pull identity data from AWS IAM, Azure Entra ID, GCP IAM, and the Kubernetes API, and feed it into a central data platform (even something as simple as scheduled exports into a warehouse).
It's not glamorous, but it gives you a list of machine identities, role bindings, last used timestamps, and attached policies, and you can build least-privilege recommendations on top of it. It isn't turnkey, but it is faster than deploying a massive PAM platform.
2. Policy-as-code + drift monitoring
Enterprises leaning into API-first models are treating IAM like infrastructure. Terraform state + cloud logs + usage metrics = effective permissions map.
From there you can identify unused actions, detect over-privileged roles, and automatically open pull requests with reduced policies. It's not fully autonomous, but it scales better than manual review, and it fits your "API-first" requirement.
3. Identity graph tools (emerging category)
There's a newer class of tools building identity graphs across clouds. They ingest role assignments, trust policies, federation relationships, and usage telemetry, then surface excess privilege, cross-cloud trust misconfigurations, and dormant service accounts.
These are lighter weight than full PAM suites and often agentless. The tradeoff is that they're governance visibility tools first, and lifecycle automation sometimes comes second.
Why Vault feels heavy (and why it still wins sometimes)
Vault gets dismissed as "operational overhead," and it is. Some enterprises still choose it anyway, and what draws them is dynamic credentials more than secrets storage.
Vault can generate short-lived database credentials, issue temporary cloud access tokens, and enforce TTL and renewal policies. If you're willing to run it properly, it becomes a machine identity broker.
That's a commitment, though, and if you don't have staff for it, it will hurt. The dividing line is your appetite for staffing it, much more than any feature list.
The CI/CD integration reality
You need Jenkins and GitHub Actions integration. Most enterprises handle this by using OIDC federation from GitHub Actions into cloud roles, letting Jenkins assume cloud roles dynamically, and removing static CI credentials entirely.
That reduces credential leakage risk, manual rotation cycles, and pipeline secret sprawl. It also fits your "no code changes" constraint, because modern SDKs already support environment-based token injection.
What enterprises actually deploy in 2026
The honest answer is a combination:
- Workload identity federation everywhere possible
- Short-lived credentials instead of rotation
- Centralized identity inventory reporting
- Policy analytics tooling
- Minimal secret managers only where federation isn't possible
They don't unify everything under one mega-platform; they standardize patterns. The question they're answering is "how do we eliminate static machine credentials," and picking a vendor comes after that.
What you probably shouldn't do
Don't:
- Try to centralize all secrets into one vault in six months
- Force applications to change authentication patterns
- Deploy agents across 2,000 workloads
- Attempt to manually right-size every IAM policy
That's how timelines explode. Your constraints are clear, so the design has to respect them.
If this were my environment
With 2,000+ service accounts across three clouds, I'd prioritize:
Phase 1 (90 days):
- Inventory machine identities across clouds
- Enable workload identity federation for new workloads
- Remove new static credentials from CI/CD
Phase 2 (Next 90 days):
- Replace long-lived keys with federated tokens where possible
- Introduce usage-based policy trimming
- Build dashboards for identity ownership
Phase 3:
- Evaluate whether a lightweight identity governance platform adds value
Notice there's no massive platform rollout in there. Your real enemy is sprawl, more than any lack of tooling.
The uncomfortable conclusion
Multi-cloud machine identity governance is an architecture problem, and buying a product won't settle it. Enterprises that win here:
- Stop rotating secrets and start eliminating them
- Stop centralizing manually and start federating automatically
- Stop thinking "vault everything" and start thinking "trust exchange"
Better password rotation matters less than having fewer passwords in the first place. The teams that internalize that early are the ones who stop waking up at 3 a.m. wondering which forgotten service account still has admin access to production.