A cloud API key was exposed
Rotate the secret, but first understand where it appeared, what it could do, how it was used, and what automation depends on it.
Scope
For teams responding to a cloud, SaaS, CI/CD, or application credential found in source code, logs, chat, tickets, or a public location.
Start with the situation, not the slogan
An incident rarely arrives with a neat label. It arrives as a forwarded screenshot, a worried telephone call, or a monitoring alert written by a machine with no sense of occasion. In this case “A cloud API key was exposed” is the working heading, but the label is only a starting hypothesis. The useful work is to establish what happened, what can still happen, and which decision cannot safely wait.
Keep three clocks in view: the attacker’s opportunity, the business interruption, and the lifetime of the evidence. They do not run at the same speed. A hasty change may interrupt access but erase context; a perfect investigation conducted at geological pace may leave the organisation exposed. Good response is the slightly unglamorous art of making the next reversible decision with the best evidence currently available.
Deleting the visible copy does not invalidate the credential. An exposed key may have been copied immediately, embedded in build artefacts, or used to create new persistence and resources.
Composite scenario
How this usually reaches the desk
Imagine the report begins with this observation: “Provider audit events, API calls, resource creation, policy changes, or downloads after the likely exposure time.” A second check returns another clue: “Secret-scanning alerts, public commits, build logs, package archives, container layers, or support attachments containing the key.” Neither fact alone tells the whole story. Together they justify a documented incident, a defined owner, and a deliberate containment decision. This is where a timeline beats a collection of heroic memories; memory is an excellent storyteller and a dreadful audit log.
This scenario combines common operational patterns; it is not presented as a report of one named incident.
What to look for
Begin with preserved, comparable evidence. One signal is rarely proof; use independent observations and a reliable timeline before declaring scope or intent.
Provider audit events, API calls, resource creation, policy changes, or downloads after the likely exposure time.
Record the exact time, source, identity, system and time zone. Compare it with a known-good baseline and with what the user or service owner expected. A surprising event is a lead, not a conviction.
Secret-scanning alerts, public commits, build logs, package archives, container layers, or support attachments containing the key.
Look for the control-plane event that made the visible activity possible: a changed credential, permission, rule, route, token or trusted device. Persistence often looks administrative because, technically, it is.
New credentials, users, roles, tokens, webhooks, or infrastructure created by the exposed identity.
Correlate the report with independent telemetry before deciding scope. User testimony, identity logs, endpoint evidence and service audit records are strongest when they agree on sequence rather than merely on mood.
Working examples
Run it, read it, decide what changes
These examples use documentation addresses, test identities and bounded targets. Replace placeholders only inside systems you own or are explicitly authorised to operate. Read the expected result and next action before running the command; a successful command is evidence, not yet a conclusion.
Locate a secret in Git history without printing every file
- Prerequisites
- A unique, non-secret identifier or a revoked token prefix; never paste a live secret into shared shell history.
git log --all --oneline -S 'REVOKED_TOKEN_PREFIX'
git grep -n 'REVOKED_TOKEN_PREFIX' $(git rev-list --all) -- ':!vendor' ':!node_modules'Commits and matching paths containing the identifier are listed.
Removing the current file does not remove earlier commits, forks, CI logs, artefacts or cloned copies.
Revoke first, inspect provider audit logs for use, replace the secret everywhere, then rewrite history only with repository-owner coordination.
Inventory environment-variable names without exposing values
- Prerequisites
- Run only on a system you are authorised to inspect.
env | cut -d= -f1 | sort | grep -Ei 'TOKEN|SECRET|KEY|PASSWORD|CREDENTIAL'Potential secret-bearing variable names appear without their values.
Names are hints; secrets may also live in files, keychains, CI variables or process arguments. Avoid commands that print values into terminals or tickets.
Map each credential to an owner, provider, permissions and rotation procedure; revoke suspected exposures before cleanup.
Prove the old credential no longer works
- Prerequisites
- A harmless read-only endpoint and the revoked credential held only in a private local variable.
OLD_TOKEN='replace-locally-never-commit'
curl -sS -o /dev/null -w '%{http_code}\n' -H "Authorization: Bearer $OLD_TOKEN" https://api.example.invalid/v1/me
unset OLD_TOKENA revoked bearer token should receive an authentication failure such as HTTP 401; the documentation host here must be replaced with the real provider endpoint.
HTTP 403 can mean the token is valid but lacks permission; a network error proves nothing. Use the provider’s documented harmless identity endpoint.
Record the revocation evidence, test the replacement through the application, and search audit logs from exposure time onward.
What to do
Read the whole sequence before starting. Several workstreams may run in parallel, but their evidence, authority and expected outcomes still need to be explicit. Every step below points back to a concrete example; use the example as implementation evidence, not as permission to operate outside the stated scope.
Operational judgement
Containment and recovery are different verbs. Containment limits the next harmful action; recovery returns a service to trustworthy operation. Between them sits eradication: removing the access path and persistence that would make the freshly restored service merely a cleaner target. For Cloud, Secrets and CI/CD, keep those decisions separate in the timeline even if a small team performs them minutes apart.
Communication is also a control. Tell affected people what is known, what remains uncertain, what they must do and when the next update will arrive. Avoid both melodrama and false reassurance. “We are investigating” is useful only when followed by an owner and a time. The goal is to reduce secondary harm without teaching a possible attacker exactly what the team has discovered.
Handover
Make the result useful to the next person
Maintain two views of the incident. The working timeline should contain detailed events, evidence locations, hypotheses and technical actions. The stakeholder update should contain confirmed impact, current containment, material uncertainty, decisions required and the next reporting time. Do not copy speculative indicators into executive statements. Equally, do not polish away uncertainty merely because it looks untidy. Both records should use absolute times with a declared time zone and should identify the source of each important fact.
At shift change, hand over the current scope, trusted administration path, preservation status, active controls, failed actions, business priorities and the next three decisions. Read back the most consequential assumptions. For this case, ensure the record begins with “Identify the exact credential, owner, permissions, environments, dependencies, exposure location, and earliest possible exposure time.” and does not finish until the team has addressed “Reduce privilege, shorten lifetime, move the replacement into a managed secret store, and enable automated detection and rotation.” A concise, accurate handover prevents the incoming team from repeating disruptive work or mistaking a quiet telemetry gap for successful containment.
Validate before you close
Confirm the old credential is rejected, all dependent services use the replacement, no unauthorised resources or identities remain, and detection covers future appearances.
Capture the test, the expected result and the observed result. Where a person or business owner must accept restored service, name them in the record. A green dashboard can confirm that a component is answering; it cannot confirm that invoices, identities or restored data are trustworthy.
Finish with a compact closure note: the original trigger, confirmed scope, evidence retained, controls changed, tests passed, known gaps, residual risk, and the people responsible for the remaining work. Schedule a review while the timeline is still fresh enough to challenge. The purpose is not to find a person to blame; computers already perform blame with admirable efficiency. The purpose is to make the next response faster, safer and less dependent on one person remembering where the useful log was hidden.
Common mistakes
- Resetting systems before preserving volatile evidence and audit logs.
- Treating the first visible symptom as the complete scope of the incident.
- Restoring service without verifying that the attacker’s access path is closed.
These errors usually come from haste, unclear ownership or misplaced confidence. Build the safeguard into the runbook: a required evidence field, a second-person review, a rollback test or a specific exit criterion.
Questions people ask when the clock is running
Does one suspicious event prove compromise?
No. Treat “Provider audit events, API calls, resource creation, policy changes, or downloads after the likely exposure time.” as a reason to investigate and preserve evidence. Confidence should rise when independent identity, service, endpoint or network records support the same sequence.
Should we reset everything immediately?
Reset or revoke what the evidence and risk justify, but preserve the state you will need to understand the incident. Broad, undocumented resets can interrupt the attacker, the business and the investigation in one impressively efficient stroke.
When can the incident be closed?
Confirm the old credential is rejected, all dependent services use the replacement, no unauthorised resources or identities remain, and detection covers future appearances. Closure also requires named owners for residual risk and follow-up work; “it seems quiet now” is an observation, not an exit criterion.
Safety boundary
Use these steps only on systems you own or are explicitly authorised to assess. Preserve evidence, follow your organisation’s legal and regulatory obligations, and prefer reversible actions when the situation is not yet understood.
Primary references
- SP 800-61 Rev. 3: Incident Response Recommendations and ConsiderationsNIST
- Software Supply Chain Security Cheat SheetOWASP
Editorial status: first edition. Review the linked vendor documentation for product- and version-specific changes before acting.