Run a backup restore drill before an incident
Measure recovery point, recovery time, dependencies, and data integrity by restoring a synthetic service into an isolated environment.
Scope
For infrastructure teams with a test backup set and an isolated recovery target.
Start with the situation, not the slogan
This lab is designed to produce evidence and judgement, not a ceremonial screenshot of a tool running. Run a backup restore drill before an incident is successful when you can explain the question, predict the expected observation, collect it safely and distinguish a useful result from noise. The commands are the least interesting part, although they are traditionally the part everyone photographs.
Use systems you own or have explicit permission to test. Keep the exercise isolated from household, client and production networks, take a snapshot before deliberate breakage, and write the rollback step before the first change. A lab without a reset path is simply a future troubleshooting appointment.
Backup success is an operational claim until a restore proves that data, keys, configuration, dependencies, capacity, and people can recreate the service.
Composite scenario
How this usually reaches the desk
Set one modest objective for the session. Begin with the expected clue: The service owner defines one expected transaction and acceptable recovery point and recovery time. Then create or collect only enough benign activity to make that clue visible. If the observation does not appear, investigate the data path before adding more tools. Instrumentation that cannot see a known test event will not become more perceptive during a real incident.
This scenario combines common operational patterns; it is not presented as a report of one named incident.
What to look for
Begin with preserved, comparable evidence. One signal is rarely proof; use independent observations and a reliable timeline before declaring scope or intent.
The service owner defines one expected transaction and acceptable recovery point and recovery time.
Write the expected observation before the exercise. Include the source, destination, time and field that should carry it; this turns an interesting screen into a falsifiable test.
The recovery target cannot overwrite production or the only backup copy.
Confirm that clocks, names and identifiers line up across the lab. Time drift and ambiguous hostnames can turn three tidy events into an accidental detective novel.
Required keys, identities, DNS, certificates, storage, and application dependencies are listed.
Keep a known-good comparison. The aim is not merely to produce an alert or packet, but to explain how the test differs from ordinary activity and where false positives would arise.
Working examples
Run it, read it, decide what changes
These examples use documentation addresses, test identities and bounded targets. Replace placeholders only inside systems you own or are explicitly authorised to operate. Read the expected result and next action before running the command; a successful command is evidence, not yet a conclusion.
List snapshots and verify repository structure
- Prerequisites
- Repository access configured through a protected credential source.
restic snapshots
restic checkSnapshots are listed and repository metadata/data references are checked.
A check can succeed even when the selected restore is application-inconsistent or credentials depend on the failed production system.
Record the latest usable snapshot time and perform a restore to a separate destination.
Restore one snapshot without overwriting production
- Prerequisites
- Empty target storage and the snapshot ID selected from the inventory.
mkdir -p /srv/restore-test
restic restore SNAPSHOT_ID --target /srv/restore-test
find /srv/restore-test -type f | wc -lRestic reports restored files and the count provides a simple completeness lead.
File count is not business validation. Databases and directory services need supported consistency and recovery procedures.
Start the restored application isolated, run owner-approved integrity checks, and measure actual recovery time.
Sample repository data on a schedule
- Prerequisites
- A maintenance window and enough bandwidth for a bounded read.
restic check --read-data-subset=5%
restic forget --dry-run --keep-daily 7 --keep-weekly 5 --keep-monthly 12A data sample is read; the retention command previews what would be kept and removed without deleting anything.
Sampling changes the probability of finding corruption, not certainty. Dry-run output must be reviewed before any prune.
Rotate through subsets over time, alert on failed jobs, and keep at least one restore path isolated from production credentials.
Lab procedure
Read the whole sequence before starting. Several workstreams may run in parallel, but their evidence, authority and expected outcomes still need to be explicit. Every step below points back to a concrete example; use the example as implementation evidence, not as permission to operate outside the stated scope.
Operational judgement
Keep a lab notebook with four columns: time, action, expected evidence and observed evidence. Add screenshots only when they preserve information that text cannot. The notebook should allow another person to repeat the exercise without inheriting your browser history, shell history and particular relationship with luck.
The transfer-to-production question matters more than the demo. For Backup, Recovery and Exercise, consider data volume, retention, credentials, privacy, performance, ownership and failure behaviour. A successful lab proves that a mechanism can work under stated conditions; it does not prove that it can be deployed everywhere before lunch.
Handover
Make the result useful to the next person
Turn the exercise into a reusable lab card. Record the learning objective, isolation boundary, diagram, versions, seed data, expected observations, exact queries, screenshots that add real information, and the reset procedure. Mark which evidence was generated and which was supplied. If the lab uses a deliberately vulnerable image or sample, store its provenance and checksum. Future-you is a different operator and deserves better documentation than “it worked after I restarted something”.
End with a short teach-back. Explain why “Select a synthetic or sanitised recovery set and record its expected timestamp and integrity values.” matters, demonstrate the observation that supports the conclusion, and show how the environment returns to baseline after “Destroy or secure the test environment, update the runbook, and schedule the next exercise based on the gaps.” Then name one production assumption the lab did not test. That last sentence keeps a useful experiment from turning into unjustified confidence and gives the next exercise a sensible place to begin.
Validate before you close
The drill passes only when the business transaction works, measured objectives are met, integrity is accepted, and another operator can follow the corrected runbook.
Capture the test, the expected result and the observed result. Where a person or business owner must accept restored service, name them in the record. A green dashboard can confirm that a component is answering; it cannot confirm that invoices, identities or restored data are trustworthy.
Finish with a compact closure note: the original trigger, confirmed scope, evidence retained, controls changed, tests passed, known gaps, residual risk, and the people responsible for the remaining work. Schedule a review while the timeline is still fresh enough to challenge. The purpose is not to find a person to blame; computers already perform blame with admirable efficiency. The purpose is to make the next response faster, safer and less dependent on one person remembering where the useful log was hidden.
Common mistakes
- Connecting an intentionally weak lab directly to a home or production network.
- Copying commands without recording the expected evidence and rollback step.
- Calling a test successful without comparing the result to a known-good baseline.
These errors usually come from haste, unclear ownership or misplaced confidence. Build the safeguard into the runbook: a required evidence field, a second-person review, a rollback test or a specific exit criterion.
Questions people ask when the clock is running
Can I run this against a public target for practice?
No. Keep the work to systems you own or are explicitly authorised to assess. An educational intention is not an access-control mechanism and will not improve the conversation with a provider or solicitor.
What should I save from the exercise?
Keep the topology, versions, raw evidence, exact filters or rules, expected result, observed result and rollback notes. Remove real secrets and personal data before sharing the notebook.
How do I know the lab worked?
The drill passes only when the business transaction works, measured objectives are met, integrity is accepted, and another operator can follow the corrected runbook. Repeat the key observation from a clean snapshot; repeatability is a stronger result than a single attractive screenshot.
Safety boundary
Use these steps only on systems you own or are explicitly authorised to assess. Preserve evidence, follow your organisation’s legal and regulatory obligations, and prefer reversible actions when the situation is not yet understood.
Primary references
Editorial status: first edition. Review the linked vendor documentation for product- and version-specific changes before acting.