Enterprise Secrets Management: How to Solve Credential Problems at Scale

A two-person team knows all of the secrets that exist because the knowledge lives in everyone’s head. You can always ask somebody to tell you all the services using the Stripe key, which services read the database password, and what would break if you changed any given credential. A simple secrets manager and a few CI variables work because you know who to ask for everything.
At enterprise scale, hundreds of services run across several clouds and dozens of Kubernetes clusters. Secrets live in stores that originate with different teams and came in via acquisitions. And, the bigger your organization, the more demands compliance makes. Storing secrets has been solved so many times that the solutions themselves become the problem because nobody can tell you with certainty who uses, owns, or has changed any given secret.
What changes about secrets management at enterprise scale?
Volume doesn’t break an enterprise secrets program because a store that can hold 5 secrets can hold 500 secrets or 50,000. The bigger problem is that there are often disparate secret stores used by different teams and nobody can answer three questions with certainty:
- Who uses this secret? What a secret is used for (if at all).
- Who owns it? Who created it and what they need it for, and whether they set up automated rotations.
- Who has what access? Because access is granted one ticket at a time, it’s hard to reconstruct an exact audit trail.
Many enterprises can theoretically cobble together most of this with laborious processes (e.g. for a compliance audit), but would consume many resources in the process. While you may still have time to provide evidence to an auditor, secrets management failures can also cause direct operational issues like rotations that cause outages because the new credential doesn’t properly propagate.
Why do manual secrets workflows stop working at scale?
Manual secrets workflows stop working because their cost grows exponentially. Each new service, environment, identity, and piece of infrastructure multiplies the overhead. When that overhead causes complications engineering teams slow down as they wait on updated secrets or ticket resolutions.
This happens because many organizations reach enterprise scale without fundamentally rearchitecting early-stage secrets management:
- A new secret is requested in a ticket, created by hand in a cloud console days later, and pasted into CI variables for each environment.
- Rotation runs off a calendar reminder: someone changes the password, updates each consumer they know about, and restarts services.
- The inventory is a spreadsheet or a wiki page, accurate on the day it was last edited.
- Access reviews are an export pasted into a spreadsheet for managers to sign.
- Offboarding is a checklist of stores to remove someone from.
Each fails at scale:
- One secret used in three environments by four services is 12 places to update. An organization with 500 credentials on a 90-day rotation policy performs eight rotations per day, and each one means finding the consumers, changing the value, and confirming nothing broke.
- Manual steps create untracked copies. Pasting a secret into a CI variable, a ticket, a chat message, or a local
.envfile creates a copy the secrets store doesn’t know about. - Ticket-based secrets provisioning blocks someone's work until the ticket is resolved. At scale, those tickets accumulate.
- Manual rotations work because someone knows which services to restart, which pulls that engineer into every rotation, incident, and audit. That knowledge leaves with them if they do.
- Nothing is logged if it’s not manually recorded.
Manual workflows can’t scale without hiring more and more people to solve an increasingly complex problem. The solution is a different shape of workflow, not a faster version of the manual one:
- Consumers authenticate as themselves instead of receiving a pasted copy of a human’s credential.
- Rotation is automated and runs on a schedule.
- Temporary access expires without anyone remembering to remove it.
- Reviews run against the live permission model instead of an export.
- Every change leaves a record, because the system making the change writes it.
This type of workflow avoids the most common failure modes of secrets management.
What are the most common failure modes of enterprise secrets management?
Through our work with large companies like LG, Nvidia, UPS, and many other enterprises, we’ve seen that enterprise secrets management reliably surfaces similar challenges:
Rotations break production
Rotating a secret breaks production when the secret has consumers nobody knows about. Three things make it worse at scale:
- Shared credentials. One database password read by 12 services is 12 things that must switch at once, sometimes owned by different teams that need to know about the restart.
- Caches. Clients hold secret values to avoid calling the store on every request, which means the wrong secrets can seep into runtimes after a rotation.
- Single-credential rotation. If the old password stops working the moment the new one is set, every consumer has to switch in the same instant.
Repeated outages teach engineers to put rotation off, which harms the security posture. The US Cyber Safety Review Board found that Microsoft stopped rotating its consumer signing keys in 2021 after an outage linked to its manual rotation process. A 2016 key a hacker group used to forge authentication tokens was still valid in 2023 because of that pause.
Rotation becomes routine when each service can authenticate as itself. This means the identities allowed to read a secret are its consumers. Ideally, old and new credentials should overlap so a cache or a slow deploy has time to catch up. Where the target supports it, dynamic secrets go further and give each workload its own short-lived credential.
Nobody owns the secret
In a large organization, many secrets outlive their owners: the creator left, the team merged, or the integration was replaced and its key never revoked. Nobody rotates them, because nobody knows what that would break, and cleanup skips them because somebody assumes they are unused.
After Okta's 2023 breach, Cloudflare missed one service token and three service accounts, out of thousands, that its postmortem says "were not rotated because mistakenly it was believed they were unused." An attacker used them to reach Cloudflare's Atlassian systems.
Ownership has to come from structure rather than memory, so a secret in a team's project keeps an owner after its creator leaves. Last-read times make pruning safe: a secret nothing has read in 90 days can go, and one read yesterday by an unknown identity is a finding.
Secrets spread across a dozen stores
Most secret stores (especially cloud-native ones) stop at their own boundary: AWS Secrets Manager only works on AWS (and is scoped to one account and one region). Microsoft recommends one Azure Key Vault per application per region, so five applications in two regions is 10 vaults by design. Then come the stores nobody planned:
| Store | Where it stops | What it cannot see |
|---|---|---|
| Cloud secret manager | One account, subscription, or project, often one region | Other clouds, on-premises systems, and other accounts without explicit policy |
| CI/CD platform secrets | The CI system that holds them | Runtime, and every other CI system |
| Kubernetes Secrets | One cluster | Every other cluster, and anything outside Kubernetes |
| Self-managed vault | The teams that onboarded to it | The teams that found onboarding too slow |
| An acquired company's stack | The acquired company | The parent company's policies and audit |
| Team password manager | Humans | Workloads |
Each has its own access model, audit log, and console, so an engineer's first question about any credential is which store holds it and who can grant access. Uber's engineering team found 25 separate vaults owned by different teams before consolidating them into six.
Migrating everything into one vault can often be a long program that eventually grinds to a halt because the success condition is a fully exhaustive secret store, which is incredibly hard. Instead, good secrets management means having one source of truth for:
- The secret value
- The policy
- The audit trail
While existing stores can (at least temporarily) continue to exist as sync targets. An application reading from AWS Secrets Manager keeps reading from it, but another secrets manager like Infisical can live on top of it and sync to it.
The platform team becomes a ticket queue
Many enterprises grant or deny access based on tickets. But this turns them into a bottleneck. Waiting on a ticket resolution can block an engineer just to get access to an API key they need. This kind of policy either means humans resort to workarounds (e.g. sharing secrets in plaintext again) or engineering slowing down.
The fix is to separate defining the structure from working inside it:
| Who | Defines structure and policy | Reads and writes secrets | Can see |
|---|---|---|---|
| Platform or security owner | Across the organization | Not needed day to day | Everything |
| Service team lead | Within their own projects | Their own projects | Their own projects |
| Engineer | No | Development and staging for their services | Their own services |
| Deployment pipeline | No | The services and environment it deploys | Nothing interactively |
| Auditor | No | Nothing | Metadata and audit logs, not values |
The platform owner decides once what a project looks like (environments, roles, who approves production changes), and a project template lets new teams create their own without a ticket.
Approvals belong where a mistake is expensive: changes to production secrets, and people getting access they do not normally hold. Applied to everything, they become the same queue, and reviewers start approving on reflex.
- Change approvals gate edits on production paths, catching the mistyped connection string and recording who agreed to it.
- Access requests replace standing production access with access for a fixed period that expires on its own.
Pipelines cannot wait for a human at 3am, so scope their identities narrowly and exempt them deliberately. For emergencies, a bypass limited to named people that notifies approvers keeps a break-glass path open without hiding it.
Permissions drift with every reorg
Permissions in a secrets store are granted one request at a time and rarely removed, so they can drift if someone changes teams or gets promoted. There are four ways to prevent this:
- Structure around services and environments, not teams. A payments service's production secrets mean the same thing after the payments team splits.
- Grant access to groups synced from the identity provider over SCIM, so a team change or a departure updates secret access on its own.
- Scope roles with conditions instead of cloning them. "Write in development and staging, read-only in production" is one role for every team. For a one-folder exception, grant access on the folder, which is why we rebuilt folder-level permissions.
- Treat the right to edit roles as the most privileged permission, since anyone holding it can grant themselves everything else.
Shared service accounts in CI
A shared service account makes life simpler for engineers, but makes every action in the audit trail look the same and creates a single point of failure. Instead, each workload should have its own identity with which it authenticates to the secrets manager. This creates a more robust audit trail and makes issues easier to diagnose and debug.
Nobody can tell what a leak exposed
The first question in any secrets incident is which secrets were exposed and to which systems. Response often pulls engineers off of their planned work. After CircleCI's January 2023 breach, customers were told to rotate any and all secrets stored in the platform. If those workflows are manual, that causes a ton of work.
Rotating everything becomes easy when you have three things in place:
- Known reach. A CI system that receives secrets through per-pipeline identities has a known blast radius before the incident starts.
- Per-identity access logs that show which of those secrets were actually read during the exposure window.
- Routine rotation, so rotating in bulk does not cause a second incident.
Besides internal operational overhead, secrets management also quickly becomes a large part of compliance at enterprise scale.
What should an auditor be able to answer without asking anyone?
An auditor should be able to reconstruct:
- Who could access each secret (policy)
- Who did access what secret (audit logs)
- What changed from the system's own records
Ideally, you want to do this without pulling engineers off their work for weeks of screenshots and interviews.
Compliance frameworks typically want to see concrete evidence and make specific demands of how you manage secrets:
- PCI DSS 4.0 forbids hard-coded system and application passwords and requires them changed at a frequency set by the organization's risk analysis. It also requires periodic review of system account access
- and SOC 2's criteria expect credentials to be removed when no longer needed.
Whatever framework you follow, the records should answer directly:
- Which secrets exist, where, and which team owns each.
- Which identities can read or change each one, and through which role or group.
- Who read or changed each secret, and when.
- When each was last rotated, against the interval the risk analysis set.
- Which production changes were approved, and by whom.
- Who held temporary access, for how long, and who granted it.
With a tool like Infisical, you can stream those logs into the SIEM the security team already uses. Every question is scoped by the first one, so secrets in stores outside the inventory fall outside the audit.
How do you migrate secrets management quickly?
Moving secret values takes days, but rewiring sprawling infrastructure across multiple clouds, teams, and environments can become a large undertaking.
Because you need to keep everything running, this kind of project can take up to a year with several hundred pipelines. The steps to follow to ensure everything keeps running smoothly are simple in theory, but become complex in practice:
- Inventory every store (cloud managers, CI variables, cluster Secrets, the old vault) before designing the new structure.
- Import, then sync back to the same stores, so applications notice nothing and the old stores stop being edited by hand.
- Move one service in staging to find what the design got wrong while it is cheap.
- Move production one service at a time. The sync keeps old stores current, so each team switches in whichever sprint suits it.
- Retire each old path once its access logs show nothing still reads from it.
The best way to do this quickly and automate your secrets management moving forward is to use a tool like Infisical.
How Infisical solves enterprise secrets management
Infisical Secrets Management is built for organizations with complex secrets management needs, juggling several teams, clouds, and existing stores. This keeps engineers doing the work running it without a specialist team or a ticket queue.
- It combines its own secret store with existing stores. Secret syncs push to AWS Secrets Manager, AWS Parameter Store, Azure Key Vault, GCP Secret Manager, HashiCorp Vault, GitHub, Vercel, and dozens more, and can import what a destination already holds.
- Structure is delegated. Projects hold environments and folders, sub-organizations can separate business units or acquisitions, and project templates give new teams the standard layout.
- Permissions are scoped by condition. Roles can be limited by environment, secret path, and tags, folder-level access handles exceptions, groups sync over SCIM from Okta, Microsoft Entra ID, JumpCloud, or PingOne, and SSO can be enforced.
- Every workload gets its own identity. Machine identities authenticate with Kubernetes, AWS, Azure, GCP, OIDC (GitHub Actions, GitLab, CircleCI, Terraform Cloud), SPIFFE, TLS certificates, and more.
- Approvals come in both forms. Change policies require a set number of approvers on chosen environments and paths, with an optional bypass limited to named users or groups. Access requests grant time-bound access with multi-step approval.
- Rotation keeps consumers working. Dual-phase rotation keeps the previous credential valid for a full interval, and dynamic secrets issue short-lived credentials for databases, cloud IAM, Kubernetes, and more.
- Audit evidence goes where security already looks. Audit logs stream to Splunk, Datadog, Sumo Logic, Cribl, and more, and Secret Insights surfaces stale secrets and exports audit reports for SOC 2 and ISO 27001.
To bring your secrets, stores, and teams under one set of rules, sign up or talk to an expert about your setup.


Certificate Management for Compliance: What PCI DSS, ISO 27001, and SOC 2 Actually Require

How environment variables actually work (and why you should delete your .env file)
