The ransomware attack that became a four-hour non-event
A ransomware attack encrypted the primary patient records server. Drilled offsite backups and a written runbook turned what could have been a scramble into a full restore inside the recovery time objective.
The shape of the work
Healthcare & Dental Practices
Ongoing retainer
Ongoing retainer
Client name withheld under NDA. Engagement details are shown to the extent our agreement permits.
What went wrong, and when
- 01
Patient records across seven clinic locations depended on one primary server, which made an untested backup a real liability. If it failed to restore, appointments, medical histories and billing would all be unrecoverable.
An untested backup is a file. Nobody knew whether it would restore, how long it would take, or in what order the systems had to come back. With patient histories, appointments and billing all on the same server, the only plan left was finding out during an incident. The backups also shared credentials with the server they protected.
We set a recovery point objective of one hour and a recovery time objective of four hours, rebuilt the backup schedule and offsite redundancy to meet them, wrote a step-by-step recovery runbook, and ran quarterly restore drills. The third of those became a real recovery when a ransomware attack encrypted the primary server.
Phase by phase
Phase 1: Set objectives
What an hour is worth
Set recovery point and recovery time objectives from the real cost of an outage for this practice.
- RPO/RTO definition
- Impact analysis
Phase 2: Rebuild backups
Scope, frequency, redundancy
Rebuilt backup scope, frequency, and offsite redundancy to actually meet the stated objectives.
- Backup architecture
- Retention policy
Phase 3: Write it down
A runbook, not folklore
Wrote a step-by-step recovery runbook for the loss of the primary server.
- Recovery runbook
- Contact tree
Phase 4: Drill
Prove it quarterly
Ran quarterly restore drills to prove the backups and the runbook both worked, and one of them became the real event.
- Drill reports
- Post-incident review
Runbook step two, isolate everything: six clinics cut from the network inside six minutes with the seventh still connected, what each site holds and when its last clean copy was sealed, one of them half an hour older because that clinic had no writes since 08:04, so nothing was lost, and the first eight minutes of the incident in order.
The numbers, before and after
3h 40m (within 4h RTO)
Time to full restore
38 minutes (within 1h RPO)
Data lost
Under half a day
Clinic downtime across 7 locations
The restore time, data loss and downtime figures all come from the real ransomware recovery: the primary server was encrypted and the same steps ran for real. Both figures fell inside the objectives set at the start, which is the only validation this kind of work can really have.
Client name withheld under NDA. Figures are approximate, drawn from the engagement’s own reporting.
The engagement
Nightly backups were already running before we started working together, but nobody had ever tried restoring from one, and there was no written plan for what to do if the primary server went down.
Seven clinic locations depended on one primary server, with nightly backups that had run for years and had never once been restored from. There was no written recovery plan. The engagement began with the question that decides everything else: what does an hour of downtime actually cost, in appointments, records access and billing?
Backup Strategy & Disaster Recovery
How it was handled
- 01
Set recovery point and recovery time objectives against what an outage would cost
Pricing the outage first is what made every later question decidable, and it took one afternoon with the practice manager.
- 02
Rebuilt backup scope, frequency, and offsite redundancy to meet those objectives
Backups became immutable for their retention period and offsite on separate credentials, closing the failure mode where an attacker with admin access deletes them.
- 03
Wrote a step-by-step recovery runbook for a primary-server loss
The runbook was written as an ordered procedure for losing the primary server, and it reads like one: steps in order, not a description of the backup system.
- 04
Ran quarterly restore drills to prove the backups and runbook actually worked
Quarterly drills restore into an isolated environment and are timed, with the runbook corrected each time reality disagreed with it.
Objectives set against cost
A one-hour recovery point and four-hour recovery time, chosen from what an outage would actually cost.
The objectives were derived for this practice: an hour of lost work and four hours of downtime were priced against order volume and staff cost, and the backup design was built to hit those numbers. Setting them first made everything else decidable. Every question about frequency, retention and redundancy has an answer once you know what an hour is worth.
- RPO and RTO priced from order volume and staff cost
- Objectives set first, backup design second
- Every later decision resolved against those two numbers
The recovery objectives before any backup was designed: a one-hour recovery point and a four-hour recovery time, the impact analysis that priced them from appointments, records access, billing and staff time, and each later decision resolved against those two numbers.
Offsite and immutable
Runbook step three on the offsite store: seven hourly copies under an immutable retention lock and separate credentials, each read back on write, the newest rejected because it was already encrypted, and the 08:34 copy chosen as the restore point.
Backup scope, frequency, and redundancy rebuilt to meet those objectives.
Backups are immutable for their retention period, so an attacker with administrative credentials can't delete or encrypt them. That's the failure mode that turns a survivable incident into a company-ending one. Copies are held offsite on separate credentials, and integrity is verified on write, because a backup nobody has read is just a file.
- Immutable retention: deletable by nobody, admins included
- Offsite copies on separate credentials
- Integrity verified on write, never assumed
Drilled, not assumed
Quarterly restore drills, the third of which became a real recovery.
Quarterly drills restore into an isolated environment and are timed against the runbook, with the runbook corrected wherever reality disagreed with it. The first drill overran and the second met the target. The third wasn't a drill: a ransomware attack encrypted the primary server and the same steps ran for real, which is the only validation this work can really have.
- Quarterly restores into an isolated environment, timed
- Runbook corrected after every drill
- The third run was a real recovery
The drill history: the first quarterly drill overran, the second met the four-hour target, and the third was the real recovery at 3h 40m, with the runbook revised after each run.
Working inside their operation
A cross-functional team of 5 worked on an ongoing retainer, covering Backup strategy, Recovery planning, Restore drills. We ran daily standups with their own lead in the room, and a demo at the end of every sprint. Scope changed twice during the engagement, and both times the change was priced and agreed before work started.
This was a retainer, because the drills are the deliverable. Quarterly restores run into an isolated environment against the runbook, timed, with the runbook corrected wherever reality disagreed with it. The first drill overran the target, the second met it, and the third wasn't a drill.
What changed in the runbook
- 01
A backup you've never restored from is a hypothesis. A restored one is a safeguard.
The nightly backups had run successfully for years, and no one could have said whether they would restore. Writing a file successfully isn't the property that matters.
- 02
The runbook mattered as much as the backups. Under pressure, nobody improvises a restore order correctly.
In a real incident, the order of restoration is where improvisation fails, and the runbook is what takes the improvising away.
- 03
The attack was survivable because this was the fourth time the team had run the restore, and the first under real pressure.
Three rehearsals turned a four-hour objective into a four-hour outcome. A first attempt under real pressure wouldn't have met it.
The recovery, one step at a time
The order is the plan. Nobody improvises it under pressure.
The real ransomware recovery, replayed through the runbook: what happened at each step, why it had to come before the next, and where the clock stood against the objectives. Pick a step, use the arrow keys once one is focused, or press Replay.
09:38 → 09:49
Every clinic is cut from the network, starting with the primary server. Six sites were cut inside six minutes of 09:38 while the seventh was still connected; nothing is restored until all seven are off.
Why the order matters
A restore onto a network the attacker can still reach is encrypted again. Isolating first is what makes every later step worth doing.
The seven clinics
- NGNorthgateCut from network
- RSRiversideCut from network
- FVFairviewCut from network
- EPElm ParkCut from network
- SBStonebridgeCut from network
- HSHalstonCut from network
- WBWestburyStill connected
Data lost · recovery point objectivenot known yet / 1h 00m
Clock since declared · recovery time objective11m / 4h 00m
Step 1 of 4 · declared 09:38, restore point 08:34.
From the primary server to a restore that's already been proven
Each stage exists to meet one of the two objectives. Copies are made often enough to meet the recovery point, kept where an attacker can't reach them, and restored on a schedule, so the recovery time is a measurement and not a guess.
- 01 · SourcePrimary serverPatient records, appointments and billing for all seven clinics, protected to a one-hour recovery point.
- 02 · ScheduleBackup orchestrationScope, frequency and redundancy rebuilt from the objectives, so no more than an hour of work is ever at risk.
- 03 · CheckVerified on writeIntegrity is checked as each copy is written, because a backup nobody has read is only a file.
- 04 · StateImmutable offsite storeImmutable for the retention period and held on separate credentials, so admin access can't delete or encrypt it.
- 05 · ProofIsolated restoreRestored quarterly into an isolated environment, timed against the runbook and the four-hour objective.
If the attacker goes for the backups too
Backup survival & restore proof
Immutable for the retention period
Copies can't be deleted or encrypted by anyone until their retention period ends, including someone holding administrative credentials. That closes the failure mode where the attacker destroys the backups along with the server.
Offsite, on separate credentials
Copies are held offsite on their own credentials, separate from the primary server's. Before the engagement the backups shared credentials with the server they protected. They don't anymore.
Restores proven every quarter
Each quarter the copies are restored into an isolated environment and timed against the runbook, and the runbook is corrected wherever reality disagreed with it. Integrity is verified as copies are written.
Running your practices on backups nobody has ever restored from? Scope your build in 3 minutes.
Scope your buildNearby engagements
E-commerceA cart that won't check out until the prescription is real
An online pharmacy where a prescription is a first-class record with its own lifecycle, and prescription-only items simply can't leave the cart without one.
Healthcare & Dental Practices · 20 weeks
Web PlatformsThe right blood type isn't enough: it has to be someone who can get there
A donor register that matches each request on blood-group compatibility and on whether the donor can actually reach the hospital. One platform serves a web admin and a mobile client from the same records.
Healthcare & Dental Practices · 16 weeks
Web PlatformsA medicine isn't a SKU, so the catalog is generic, strength and company before it's a product
A pharmacy counter system whose catalog is modeled on how medicines actually differ (generic, strength and manufacturer as separate axes), so when a brand is out of stock the counter can find an equivalent instead of turning the customer away.
Healthcare & Dental Practices · 16 weeks
Let's talk
Running a large platform, shaping a first MVP, or getting a product ready for a funding round? Tell us where you are. We'll shape the process around it, and stay with you after launch.














