——
OperationalPLAYBOOKS

Playbook: ransomware containment and recovery

Coordinate business continuity, isolation, evidence, identity protection, external obligations, and clean recovery.

For incident command, IT, security, legal, communications, finance, insurance, and executive decision-makers.

RansomwarePlaybookRecovery

Start with the situation, not the slogan

A playbook is a decision aid for tired people working with incomplete information. Playbook: ransomware containment and recovery therefore names owners, evidence, containment and exit criteria instead of pretending every incident follows a flowchart. Read it before the incident, tailor the contacts and systems, and exercise the awkward branches.

The first objective is shared situational awareness: what is known, who owns the incident, what business process is at risk, and which decision is due next. A chat channel full of competent people is not the same as command. Assign a coordinator and a note taker, even when the organisation is small enough for both to use the same kettle.

Ransomware response fails when technical containment is separated from business priorities or when recovery begins before identity, management, and backup trust are established.

How this usually reaches the desk

The playbook should activate when the team can establish its entry conditions, including: Incident command, technical lead, business service owners, legal, and communications are reachable out of band. If the report remains ambiguous, open a low-severity record and time-box validation rather than ignoring it or declaring the end of civilisation. Escalate as evidence, privilege, spread or business impact increases.

This scenario combines common operational patterns; it is not presented as a report of one named incident.

What to look for

Begin with preserved, comparable evidence. One signal is rarely proof; use independent observations and a reliable timeline before declaring scope or intent.

01

Incident command, technical lead, business service owners, legal, and communications are reachable out of band.

Confirm this condition from an authoritative source and note who supplied it. Mark assumptions as assumptions; they have an unfortunate habit of becoming facts in copied status updates.

02

Affected segments, identity systems, remote management, hypervisors, backups, and essential services are mapped.

Identify the decision that depends on the information and the deadline for obtaining it. Evidence is most useful when it changes action rather than merely decorating the timeline.

03

Isolation and shutdown decisions account for active harm, evidence, safety, and recovery dependencies.

Check coverage and retention before containment. If a source is unavailable, record the gap and use independent evidence rather than silently treating absence as innocence.

Run it, read it, decide what changes

These examples use documentation addresses, test identities and bounded targets. Replace placeholders only inside systems you own or are explicitly authorised to operate. Read the expected result and next action before running the command; a successful command is evidence, not yet a conclusion.

Example 01

Find recently changed files without opening them

findLinux file server
Prerequisites
A known incident start time and read access to the affected filesystem.
shell
sudo find /srv/data -xdev -type f -newermt "2026-08-18 08:00:00 UTC"   -printf '%TY-%Tm-%TdT%TH:%TM:%TSZ %s %p\n' | sort | head -200
Expected result

A time-ordered sample of files modified after the incident start appears.

How to interpret it

A burst of renamed or similarly sized files can support the ransomware hypothesis; ordinary batch jobs can look similar. Do not use this list as the only scope measure.

Next action

Save the full output, compare with storage snapshots and endpoint telemetry, and identify the first host or account that wrote the files.

Example 02

Review active SMB sessions

PowerShell SMB cmdletsWindows file server
Prerequisites
Administrator PowerShell on the file server.
powershell
Get-SmbSession |
  Sort-Object ClientComputerName |
  Select-Object ClientComputerName,ClientUserName,NumOpens,SecondsExists
Expected result

Current SMB clients, users and open-object counts are displayed.

How to interpret it

A high open count or new client is a lead, not proof. Terminating the wrong session may interrupt recovery or evidence collection.

Next action

Record the list, confirm the suspected source with EDR and file audit logs, then isolate the endpoint or close only the authorised containment target.

Example 03

Test a restore into an empty directory

resticLinux, macOS or Windows
Prerequisites
A configured repository, its password supplied securely, and an empty destination outside production.
shell
restic snapshots
mkdir -p /restore-test
restic restore latest --target /restore-test
restic check --read-data-subset=5%
Expected result

The restore reports files restored and `restic check` completes without repository errors.

How to interpret it

A successful command proves only the sampled repository and selected snapshot were readable. Application consistency and business correctness still need testing.

Next action

Open the restored service offline, verify critical records with the owner, record RTO/RPO, and never overwrite encrypted production data during the test.

Response sequence

Read the whole sequence before starting. Several workstreams may run in parallel, but their evidence, authority and expected outcomes still need to be explicit. Every step below points back to a concrete example; use the example as implementation evidence, not as permission to operate outside the stated scope.

  1. 01

    Declare the incident, start the decision log, establish out-of-band communication, and notify required partners.

    Name incident command, technical lead, business owner, communications and evidence responsibility. Define a secure out-of-band channel if normal systems may be affected.

    Working example 01: Find recently changed files without opening them — find on Linux file server.

  2. 02

    Isolate affected paths with reversible controls and protect identity, backup, virtualisation, and remote-management administration.

    Capture volatile or short-retention evidence and a configuration snapshot before changes, when risk and safety permit. Record who collected what, when and from where.

    Working example 02: Review active SMB sessions — PowerShell SMB cmdlets on Windows file server.

  3. 03

    Preserve ransom artefacts, logs, endpoint evidence, network records, cloud audit, and backup changes.

    Choose containment that matches confirmed scope and business risk. State the expected effect, authority, rollback and verification method before execution.

    Working example 03: Test a restore into an empty directory — restic on Linux, macOS or Windows.

  4. 04

    Assess encryption, data theft, persistence, legal obligations, and business priorities before recovery decisions.

    Investigate related identities, systems, data and trust paths using shared indicators and behaviour. Keep confirmed facts separate from hypotheses and tasks.

    Working example 01: Find recently changed files without opening them — find on Linux file server.

  5. 05

    Recover clean priority services from verified sources, monitor closely, and rebuild trust in stages.

    Restore in controlled stages, verify security and business function, communicate residual risk, and assign every follow-up action to a named owner and date.

    Working example 02: Review active SMB sessions — PowerShell SMB cmdlets on Windows file server.

Operational judgement

Run the playbook as a checklist with judgement, not as a spell. For Ransomware, Playbook and Recovery, localise tenant names, log locations, provider contacts, legal thresholds, emergency credentials and business priorities. A generic document becomes operational only when the people on duty can find the required access without consulting the person currently on holiday.

Use a decision log separate from the task list. For each consequential action, note time, owner, evidence, choice, expected effect and observed result. This gives later reviewers something better than a reconstructed story and helps the next shift understand why a seemingly obvious action was delayed or rejected.

Make the result useful to the next person

Prepare the playbook package before an incident: primary and deputy owners, current contact routes, system inventory, evidence locations, delegated authority, provider escalation, legal and privacy triggers, secure communication, and emergency credentials. Keep a printable or otherwise independent copy where an identity or collaboration outage cannot hide it. Review contact details during exercises; telephone numbers appear to age faster than almost any cryptographic primitive.

For this playbook, the incident record should open with “Declare the incident, start the decision log, establish out-of-band communication, and notify required partners.” and drive deliberately towards “Recover clean priority services from verified sources, monitor closely, and rebuild trust in stages.” A shift handover must state confirmed facts, open hypotheses, business impact, containment already applied, evidence at risk, approvals pending and the next scheduled update. After closure, turn lessons into owned changes to controls, documentation and exercises. A lesson without an owner and a date is simply a well-written regret.

Validate before you close

Exit only when critical services operate from trusted systems, privileged access is rebuilt, persistence checks are clean, required notifications are complete, and follow-up risk is accepted.

Capture the test, the expected result and the observed result. Where a person or business owner must accept restored service, name them in the record. A green dashboard can confirm that a component is answering; it cannot confirm that invoices, identities or restored data are trustworthy.

Finish with a compact closure note: the original trigger, confirmed scope, evidence retained, controls changed, tests passed, known gaps, residual risk, and the people responsible for the remaining work. Schedule a review while the timeline is still fresh enough to challenge. The purpose is not to find a person to blame; computers already perform blame with admirable efficiency. The purpose is to make the next response faster, safer and less dependent on one person remembering where the useful log was hidden.

Common mistakes

  • Running the playbook without assigning a decision owner and note taker.
  • Containing too early or too late because evidence and business impact were not compared.
  • Closing the incident without documenting recovery checks and follow-up owners.

These errors usually come from haste, unclear ownership or misplaced confidence. Build the safeguard into the runbook: a required evidence field, a second-person review, a rollback test or a specific exit criterion.

Questions people ask when the clock is running

Who may activate the playbook?

Define that locally before an incident. Help desk or monitoring staff should be able to open a record and escalate; containment authority may sit with incident command, service ownership or an executive depending on impact.

Must every step be completed in order?

No. Evidence, containment, communication and recovery often run in parallel. Keep dependencies explicit and do not let a later checkbox imply that an earlier decision was actually verified.

What is the exit condition?

Exit only when critical services operate from trusted systems, privileged access is rebuilt, persistence checks are clean, required notifications are complete, and follow-up risk is accepted. Residual risk, communication and improvement actions still need named owners even after service is restored.

Safety boundary

Use these steps only on systems you own or are explicitly authorised to assess. Preserve evidence, follow your organisation’s legal and regulatory obligations, and prefer reversible actions when the situation is not yet understood.

Primary references

  1. StopRansomware GuideCISA
  2. Ransomware Protection and ResponseNIST
  3. SP 800-61 Rev. 3: Incident Response Recommendations and ConsiderationsNIST
  4. Restic Documentationrestic

Editorial status: first edition. Review the linked vendor documentation for product- and version-specific changes before acting.