——
HighINCIDENTS

The backup existed. The restore failed.

Treat a failed restore as an incident, then separate media integrity, credentials, dependencies, capacity, and runbook failures.

For infrastructure teams discovering during an outage or exercise that a required restore cannot complete or produce a usable service.

BackupRecoveryBusiness continuity

Start with the situation, not the slogan

An incident rarely arrives with a neat label. It arrives as a forwarded screenshot, a worried telephone call, or a monitoring alert written by a machine with no sense of occasion. In this case “The backup existed. The restore failed.” is the working heading, but the label is only a starting hypothesis. The useful work is to establish what happened, what can still happen, and which decision cannot safely wait.

Keep three clocks in view: the attacker’s opportunity, the business interruption, and the lifetime of the evidence. They do not run at the same speed. A hasty change may interrupt access but erase context; a perfect investigation conducted at geological pace may leave the organisation exposed. Good response is the slightly unglamorous art of making the next reversible decision with the best evidence currently available.

A successful backup job proves that data was written somewhere; it does not prove that the required systems, keys, dependencies, and recovery knowledge can recreate a working service within the business deadline.

How this usually reaches the desk

Imagine the report begins with this observation: “Restore jobs fail checksum, decryption, media, capacity, or permission checks.” A second check returns another clue: “The data restores but the application cannot start because identity, DNS, certificates, or databases are missing.” Neither fact alone tells the whole story. Together they justify a documented incident, a defined owner, and a deliberate containment decision. This is where a timeline beats a collection of heroic memories; memory is an excellent storyteller and a dreadful audit log.

This scenario combines common operational patterns; it is not presented as a report of one named incident.

What to look for

Begin with preserved, comparable evidence. One signal is rarely proof; use independent observations and a reliable timeline before declaring scope or intent.

01

Restore jobs fail checksum, decryption, media, capacity, or permission checks.

Record the exact time, source, identity, system and time zone. Compare it with a known-good baseline and with what the user or service owner expected. A surprising event is a lead, not a conviction.

02

The data restores but the application cannot start because identity, DNS, certificates, or databases are missing.

Look for the control-plane event that made the visible activity possible: a changed credential, permission, rule, route, token or trusted device. Persistence often looks administrative because, technically, it is.

03

Recovery time and recovery point objectives are exceeded even though backup dashboards appeared healthy.

Correlate the report with independent telemetry before deciding scope. User testimony, identity logs, endpoint evidence and service audit records are strongest when they agree on sequence rather than merely on mood.

Run it, read it, decide what changes

These examples use documentation addresses, test identities and bounded targets. Replace placeholders only inside systems you own or are explicitly authorised to operate. Read the expected result and next action before running the command; a successful command is evidence, not yet a conclusion.

Example 01

List snapshots and verify repository structure

resticAny restic-supported platform
Prerequisites
Repository access configured through a protected credential source.
shell
restic snapshots
restic check
Expected result

Snapshots are listed and repository metadata/data references are checked.

How to interpret it

A check can succeed even when the selected restore is application-inconsistent or credentials depend on the failed production system.

Next action

Record the latest usable snapshot time and perform a restore to a separate destination.

Example 02

Restore one snapshot without overwriting production

resticRecovery host
Prerequisites
Empty target storage and the snapshot ID selected from the inventory.
shell
mkdir -p /srv/restore-test
restic restore SNAPSHOT_ID --target /srv/restore-test
find /srv/restore-test -type f | wc -l
Expected result

Restic reports restored files and the count provides a simple completeness lead.

How to interpret it

File count is not business validation. Databases and directory services need supported consistency and recovery procedures.

Next action

Start the restored application isolated, run owner-approved integrity checks, and measure actual recovery time.

Example 03

Sample repository data on a schedule

resticBackup server
Prerequisites
A maintenance window and enough bandwidth for a bounded read.
shell
restic check --read-data-subset=5%
restic forget --dry-run --keep-daily 7 --keep-weekly 5 --keep-monthly 12
Expected result

A data sample is read; the retention command previews what would be kept and removed without deleting anything.

How to interpret it

Sampling changes the probability of finding corruption, not certainty. Dry-run output must be reviewed before any prune.

Next action

Rotate through subsets over time, alert on failed jobs, and keep at least one restore path isolated from production credentials.

What to do

Read the whole sequence before starting. Several workstreams may run in parallel, but their evidence, authority and expected outcomes still need to be explicit. Every step below points back to a concrete example; use the example as implementation evidence, not as permission to operate outside the stated scope.

  1. 01

    Open an incident record and identify the business service, required recovery point, deadline, dependencies, and decision owner.

    Write down who can authorise containment, who records the timeline and which business service is at risk. If nobody owns a decision, the decision will eventually be made by whichever system fails first.

    Working example 01: List snapshots and verify repository structure — restic on Any restic-supported platform.

  2. 02

    Preserve backup logs and catalogue state. Do not overwrite the last known-good recovery set while troubleshooting.

    Prefer a control that is fast, reversible and observable. Note its expected effect before applying it, then check that the effect occurred; clicking a red button is an action, not proof.

    Working example 02: Restore one snapshot without overwriting production — restic on Recovery host.

  3. 03

    Test the chain in layers: media readability, decryption keys, file integrity, operating system, application dependencies, then business transactions.

    Export or preserve the records most likely to expire, roll over or be changed by containment. Use original time stamps, document collection time and keep the untouched source alongside any working copy.

    Working example 03: Sample repository data on a schedule — restic on Backup server.

  4. 04

    Use an alternate recovery path or older verified recovery point when current media is suspect, recording the data-loss trade-off explicitly.

    Search for mechanisms that survive the obvious fix: alternate credentials, delegated access, scheduled activity, trusted applications, modified recovery details and management-plane changes.

    Working example 01: List snapshots and verify repository structure — restic on Any restic-supported platform.

  5. 05

    Correct the runbook, ownership, access, capacity, and monitoring gaps revealed by the failure; schedule a fresh restore exercise.

    Expand scope by shared infrastructure and behaviour, not by panic. Related identities, devices, recipients and services deserve review when evidence connects them to the same access path or campaign.

    Working example 02: Restore one snapshot without overwriting production — restic on Recovery host.

Operational judgement

Containment and recovery are different verbs. Containment limits the next harmful action; recovery returns a service to trustworthy operation. Between them sits eradication: removing the access path and persistence that would make the freshly restored service merely a cleaner target. For Backup, Recovery and Business continuity, keep those decisions separate in the timeline even if a small team performs them minutes apart.

Communication is also a control. Tell affected people what is known, what remains uncertain, what they must do and when the next update will arrive. Avoid both melodrama and false reassurance. “We are investigating” is useful only when followed by an owner and a time. The goal is to reduce secondary harm without teaching a possible attacker exactly what the team has discovered.

Make the result useful to the next person

Maintain two views of the incident. The working timeline should contain detailed events, evidence locations, hypotheses and technical actions. The stakeholder update should contain confirmed impact, current containment, material uncertainty, decisions required and the next reporting time. Do not copy speculative indicators into executive statements. Equally, do not polish away uncertainty merely because it looks untidy. Both records should use absolute times with a declared time zone and should identify the source of each important fact.

At shift change, hand over the current scope, trusted administration path, preservation status, active controls, failed actions, business priorities and the next three decisions. Read back the most consequential assumptions. For this case, ensure the record begins with “Open an incident record and identify the business service, required recovery point, deadline, dependencies, and decision owner.” and does not finish until the team has addressed “Correct the runbook, ownership, access, capacity, and monitoring gaps revealed by the failure; schedule a fresh restore exercise.” A concise, accurate handover prevents the incoming team from repeating disruptive work or mistaking a quiet telemetry gap for successful containment.

Validate before you close

A restore is proven only when an authorised business owner can perform an expected transaction and the team has measured actual recovery time, recovery point, and missing dependencies.

Capture the test, the expected result and the observed result. Where a person or business owner must accept restored service, name them in the record. A green dashboard can confirm that a component is answering; it cannot confirm that invoices, identities or restored data are trustworthy.

Finish with a compact closure note: the original trigger, confirmed scope, evidence retained, controls changed, tests passed, known gaps, residual risk, and the people responsible for the remaining work. Schedule a review while the timeline is still fresh enough to challenge. The purpose is not to find a person to blame; computers already perform blame with admirable efficiency. The purpose is to make the next response faster, safer and less dependent on one person remembering where the useful log was hidden.

Common mistakes

  • Resetting systems before preserving volatile evidence and audit logs.
  • Treating the first visible symptom as the complete scope of the incident.
  • Restoring service without verifying that the attacker’s access path is closed.

These errors usually come from haste, unclear ownership or misplaced confidence. Build the safeguard into the runbook: a required evidence field, a second-person review, a rollback test or a specific exit criterion.

Questions people ask when the clock is running

Does one suspicious event prove compromise?

No. Treat “Restore jobs fail checksum, decryption, media, capacity, or permission checks.” as a reason to investigate and preserve evidence. Confidence should rise when independent identity, service, endpoint or network records support the same sequence.

Should we reset everything immediately?

Reset or revoke what the evidence and risk justify, but preserve the state you will need to understand the incident. Broad, undocumented resets can interrupt the attacker, the business and the investigation in one impressively efficient stroke.

When can the incident be closed?

A restore is proven only when an authorised business owner can perform an expected transaction and the team has measured actual recovery time, recovery point, and missing dependencies. Closure also requires named owners for residual risk and follow-up work; “it seems quiet now” is an observation, not an exit criterion.

Safety boundary

Use these steps only on systems you own or are explicitly authorised to assess. Preserve evidence, follow your organisation’s legal and regulatory obligations, and prefer reversible actions when the situation is not yet understood.

Primary references

  1. Ransomware Protection and ResponseNIST
  2. Cybersecurity Framework 2.0NIST
  3. Restic Documentationrestic

Editorial status: first edition. Review the linked vendor documentation for product- and version-specific changes before acting.