Skip to content

The ransomware attack that became a four-hour non-event

A ransomware attack encrypted the primary patient records server. Drilled offsite backups and a written runbook turned what could have been a scramble into a full restore inside the recovery time objective.

Step 5 · Restore in progress

Runbook rev 3 · from the 08:34 copy onto a clean server

Status for staffINC-071407/14/202612:36+2:58JRJonah ReyesRetainer on-call

Clinics back

5 / 7
1 restoring · 1 waiting

Since declared · RTO 4h

2h 58m of 4h 00m

Restore point · RPO 1h

08:34 · 38 min lost

Patient records counted

62,722 / 73,922

Clinics in runbook order

Counted against what each site held in its restore copy
ClinicStateCopyPatient recordsApptsBillingBack at
NGNorthgateClinic 1 of 7Back08:3414,212 / 14,212match2,10484610:58
RSRiversideClinic 2 of 7Back08:3411,847 / 11,847match1,68870211:21
FVFairviewClinic 3 of 7Back08:3412,390 / 12,390match1,81577711:43
EPElm ParkClinic 4 of 7Back08:349,603 / 9,603match1,32251912:04
SBStonebridgeClinic 5 of 7Back08:348,775 / 8,775match1,20748812:22
HSHalstonClinic 6 of 7Restoring08:345,895 / 10,16458%1,431611—
WBWestburyClinic 7 of 7Waiting08:04— / 6,931queued948402—
Westbury’s copy is 08:04, half an hour older: no writes since 08:04 · nothing lost. The clinic opens later; its first check-in is after 09:12.

Runbook

rev 3
  1. Confirm what is protected09:38
  2. Isolate everything09:49
  3. Pick the right copy10:02
  4. Stand up a clean server10:31
  5. 5Restore in runbook orderrunning
  6. 6Verify counts, reopen front desks—

Activity

12:36
  1. Halston · 5,895 of 10,164 records counted12:36
  2. Stonebridge back · 8,775 / 8,775 match12:22
  3. Elm Park back · 9,603 / 9,603 match12:04
  4. Fairview back · 12,390 / 12,390 match11:43
  5. Riverside back · 11,847 / 11,847 match11:21
  6. Northgate back · 14,212 / 14,212 match10:58
  7. Clean server up from the 08:34 copy10:31

The shape of the work

Industry

Healthcare & Dental Practices

Duration

Ongoing retainer

Cooperation model

Ongoing retainer

Services
Backup strategyRecovery planningRestore drills
Integrations
HL7 FHIRTwilioStripeDocuSign
Technologies
Immutable offsite backupsBackup orchestrationRecovery runbookRestore drill automationMonitoring
Team
1 Project lead1 Product designer2 Frontend engineers1 Backend engineer1 QA engineer

Client name withheld under NDA. Engagement details are shown to the extent our agreement permits.

The problem

What went wrong, and when

  1. 01

    Patient records across seven clinic locations depended on one primary server, which made an untested backup a real liability. If it failed to restore, appointments, medical histories and billing would all be unrecoverable.

    An untested backup is a file. Nobody knew whether it would restore, how long it would take, or in what order the systems had to come back. With patient histories, appointments and billing all on the same server, the only plan left was finding out during an incident. The backups also shared credentials with the server they protected.

    We set a recovery point objective of one hour and a recovery time objective of four hours, rebuilt the backup schedule and offsite redundancy to meet them, wrote a step-by-step recovery runbook, and ran quarterly restore drills. The third of those became a real recovery when a ransomware attack encrypted the primary server.

Process

Phase by phase

  1. Phase 1: Set objectives

    What an hour is worth

    Set recovery point and recovery time objectives from the real cost of an outage for this practice.

    • RPO/RTO definition
    • Impact analysis
  2. Phase 2: Rebuild backups

    Scope, frequency, redundancy

    Rebuilt backup scope, frequency, and offsite redundancy to actually meet the stated objectives.

    • Backup architecture
    • Retention policy
  3. Phase 3: Write it down

    A runbook, not folklore

    Wrote a step-by-step recovery runbook for the loss of the primary server.

    • Recovery runbook
    • Contact tree
  4. Phase 4: Drill

    Prove it quarterly

    Ran quarterly restore drills to prove the backups and the runbook both worked, and one of them became the real event.

    • Drill reports
    • Post-incident review
Step 2 · Isolate everything

Every clinic off the network before anything is restored

Cut remaining siteINC-071407/14/202609:46+0:08JRJonah ReyesRetainer on-call

Clinics cut

6 / 7
All six by 09:44 · 1 still going

Primary server

Encrypted
Isolated 09:39 · left untouched

Offsite store

Untouched
Retention lock on · separate credentials

Since declared

0:08 of 4:00
Declared 09:38 · RTO clock running

Network isolation

What each site holds, and its last clean copy
SiteHoldsClean copyNetworkCut at
Primary serverchd-primary-01Records, appointments, billing · all 7 clinicsEncryptedIsolated09:39
NGNorthgateClinic 114,212 records2,104 appts · 846 billing08:34sealed, verifiedCut09:39+1m
RSRiversideClinic 211,847 records1,688 appts · 702 billing08:34sealed, verifiedCut09:40+2m
FVFairviewClinic 312,390 records1,815 appts · 777 billing08:34sealed, verifiedCut09:41+3m
EPElm ParkClinic 49,603 records1,322 appts · 519 billing08:34sealed, verifiedCut09:42+4m
SBStonebridgeClinic 58,775 records1,207 appts · 488 billing08:34sealed, verifiedCut09:43+5m
HSHalstonClinic 610,164 records1,431 appts · 611 billing08:34sealed, verifiedCut09:44+6m
WBWestburyClinic 76,931 records948 appts · 402 billing08:04No writes since 08:04nothing lostStill up—
Nothing is restored until step 2 is complete. A restore onto a reachable network can be encrypted again.

The first eight minutes

09:38–09:46
  1. +0mMass file renames on the primary server09:38Monitoring alert on chd-primary-01 · incident declared
  2. +0mRunbook opened: Primary-server loss, rev 309:38Step 1 confirmed: 7 clinics, restore copies on the offsite store
  3. +1mchd-primary-01 and Northgate cut09:39Site firewall: all routes to the primary server dropped
  4. +2mRiverside cut09:40Site firewall: all routes to the primary server dropped
  5. +3mFairview cut09:41Site firewall: all routes to the primary server dropped
  6. +4mElm Park cut09:42Site firewall: all routes to the primary server dropped
  7. +5mStonebridge cut09:43Site firewall: all routes to the primary server dropped
  8. +6mHalston cut09:44Six sites cut inside six minutes of 09:38
  9. +7mOffsite store checked09:457 hourly copies present · retention lock on · separate credentials
  10. +8mWestbury: link still up09:46Site firewall not answering · on-call phoning the front desk
On screen

Runbook step two, isolate everything: six clinics cut from the network inside six minutes with the seventh still connected, what each site holds and when its last clean copy was sealed, one of them half an hour older because that clinic had no writes since 08:04, so nothing was lost, and the first eight minutes of the incident in order.

The numbers, before and after

3h 40m (within 4h RTO)

Time to full restore

38 minutes (within 1h RPO)

Data lost

Under half a day

Clinic downtime across 7 locations

The restore time, data loss and downtime figures all come from the real ransomware recovery: the primary server was encrypted and the same steps ran for real. Both figures fell inside the objectives set at the start, which is the only validation this kind of work can really have.

Client name withheld under NDA. Figures are approximate, drawn from the engagement’s own reporting.

Introduction

The engagement

Nightly backups were already running before we started working together, but nobody had ever tried restoring from one, and there was no written plan for what to do if the primary server went down.

Seven clinic locations depended on one primary server, with nightly backups that had run for years and had never once been restored from. There was no written recovery plan. The engagement began with the question that decides everything else: what does an hour of downtime actually cost, in appointments, records access and billing?

Backup Strategy & Disaster Recovery

How it was handled

  1. 01

    Set recovery point and recovery time objectives against what an outage would cost

    Pricing the outage first is what made every later question decidable, and it took one afternoon with the practice manager.

  2. 02

    Rebuilt backup scope, frequency, and offsite redundancy to meet those objectives

    Backups became immutable for their retention period and offsite on separate credentials, closing the failure mode where an attacker with admin access deletes them.

  3. 03

    Wrote a step-by-step recovery runbook for a primary-server loss

    The runbook was written as an ordered procedure for losing the primary server, and it reads like one: steps in order, not a description of the backup system.

  4. 04

    Ran quarterly restore drills to prove the backups and runbook actually worked

    Quarterly drills restore into an isolated environment and are timed, with the runbook corrected each time reality disagreed with it.

Objectives set against cost

A one-hour recovery point and four-hour recovery time, chosen from what an outage would actually cost.

The objectives were derived for this practice: an hour of lost work and four hours of downtime were priced against order volume and staff cost, and the backup design was built to hit those numbers. Setting them first made everything else decidable. Every question about frequency, retention and redundancy has an answer once you know what an hour is worth.

What shipped
  • RPO and RTO priced from order volume and staff cost
  • Objectives set first, backup design second
  • Every later decision resolved against those two numbers
Recovery objectives

Set before the backup design · priced with the practice manager

Impact analysis Signed off 12/09/2025JRJonah ReyesRetainer on-call

Recovery point objective · RPOSigned off

1 hour

The most work the clinics can afford to lose: everything written since the last clean copy.

Hourly copiesVerified on write

Recovery time objective · RTOSigned off

4 hours

The longest the seven clinics can be without records, appointments and billing.

Runbook restore orderQuarterly drills

Priced from

  • Appointmentsper hour down
  • Records accessper hour down
  • Billingper hour down
  • Staff time and costper hour down

Across all 7 clinics on one primary server

Impact analysis · what an hour down costs

Structure of the model
Cost driverWhat stopsPriced fromSets
AppointmentsChairs sit idle; booked patients are called and rebookedAppointment book, chairs per clinicRTO
Records accessMedical histories unreadable; treatment waits for themPatient recordsRTO
BillingClaims and payments are not postedBilling ledgerRTO
Lost workEvery change since the last copy is re-entered by handStaff time and costRPO

Cost of an outage by hours downShape of the model · no figures shown

0h1h2h3h4h5h6h7h8hRPO 1hRTO 4hdesign target

Every later decision, resolved

Against the two numbers
  1. 1How often to copyRPOHourlyNo more than an hour of work is ever at risk
  2. 2How to know a copy is goodRPOVerified on writeA copy nobody has read back is a file, not a backup
  3. 3How long copies surviveRPOImmutable for the retention periodAdmin credentials cannot delete or encrypt them
  4. 4Where copies liveRTOOffsite, separate credentialsLosing the server doesn't lose the copies
  5. 5How we know 4 hours is realRTOQuarterly timed drillsRestored into an isolated environment, runbook corrected
On screen

The recovery objectives before any backup was designed: a one-hour recovery point and a four-hour recovery time, the impact analysis that priced them from appointments, records access, billing and staff time, and each later decision resolved against those two numbers.

Step 3 · Pick the right copy

Seven hourly copies on the offsite store, newest first

INC-071407/14/202610:02+0:24JRJonah ReyesRetainer on-call

Immutable retention lockOn

No credential can delete or encrypt a copy before its retention period ends, admin included.

Offsite, separate credentialsOn

Written by an identity the clinic domain does not hold. The primary server has no key.

Integrity verified on writeOn

Each copy is read back and checksummed as it is sealed, then scanned for encrypted files.

Hourly copies · offsite store

7 of 7 read back on write
SealedSizeRead backContent scanLockDecision
09:34Newest212.6 GBChecksum matchEncrypted contentRenamed, unreadable filesRejected
08:341h older184.3 GBChecksum matchCleanEvery file opensRestore point
07:342h older184.1 GBChecksum matchCleanEvery file opensClean · older
06:343h older184.0 GBChecksum matchCleanEvery file opensClean · older
05:344h older184.0 GBChecksum matchCleanEvery file opensClean · older
04:345h older183.9 GBChecksum matchCleanEvery file opensClean · older
03:346h older183.9 GBChecksum matchCleanEvery file opensClean · older
The newest copy is not the cleanest. The 09:34 copy wrote and verified correctly. It faithfully copied files that were already encrypted, so restoring it would restore the attack.

Restore point

Confirmed by two people
08:34Hourly copy
184.3 GB · verified

38 min of data lostRPO 1h 00m

How the point was found

  1. 08:34Copy sealed, content clean
  2. 09:12First encrypted write (file log)
  3. 09:34Copy sealed, already encrypted
  4. 09:38Incident declared

Per-site restore point

6 clinics08:34
WestburyNo writes since 08:04 · nothing lost08:04
JRMGTwo-person confirm
10:02
Start restore

Offsite and immutable

On screen

Runbook step three on the offsite store: seven hourly copies under an immutable retention lock and separate credentials, each read back on write, the newest rejected because it was already encrypted, and the 08:34 copy chosen as the restore point.

Backup scope, frequency, and redundancy rebuilt to meet those objectives.

Backups are immutable for their retention period, so an attacker with administrative credentials can't delete or encrypt them. That's the failure mode that turns a survivable incident into a company-ending one. Copies are held offsite on separate credentials, and integrity is verified on write, because a backup nobody has read is just a file.

What shipped
  • Immutable retention: deletable by nobody, admins included
  • Offsite copies on separate credentials
  • Integrity verified on write, never assumed

Drilled, not assumed

Quarterly restore drills, the third of which became a real recovery.

Quarterly drills restore into an isolated environment and are timed against the runbook, with the runbook corrected wherever reality disagreed with it. The first drill overran and the second met the target. The third wasn't a drill: a ransomware attack encrypted the primary server and the same steps ran for real, which is the only validation this work can really have.

What shipped
  • Quarterly restores into an isolated environment, timed
  • Runbook corrected after every drill
  • The third run was a real recovery
Restore drills

Quarterly · isolated environment · timed against the runbook

Drill reports Next drill · next quarterJRJonah ReyesRetainer on-call
Drill 101/20/2026

Isolated environment

Against the 4h objective

Overran

Timed restore finished past the four-hour target. Every step where reality disagreed with the runbook was written down.

Runbook after rev 2 · corrected
Drill 204/21/2026

Isolated environment

Against the 4h objective

Met

Timed restore finished inside the four-hour target, on the corrected runbook. It was still corrected again afterward.

Runbook after rev 3 · corrected
INC-0714 · not a drill07/14/2026

Production · ransomware

Against the 4h objective

3h 40m

Full restore3h 40m of 4h 00m

Data lost38 min of 1h 00m

Runbook after rev 4 · review

Run log

Every run timed against the same runbook
RunDateEnvironmentOutcomeRunbook
Drill 101/20/2026Isolated environmentOverranrev 2 · corrected
Drill 204/21/2026Isolated environmentMetrev 3 · corrected
INC-071407/14/2026Production · ransomware3h 40m · metrev 4 · review
Drills restore into an isolated environment with no route to the clinics or the primary server, from the same offsite copies a real recovery would use. INC-0714 ran the same steps for real.

Runbook revisions

Primary-server loss
  1. rev 112/2025Written as an ordered procedure for losing the primary server
  2. rev 201/2026Corrected after Drill 1, where reality disagreed with it
  3. rev 304/2026Corrected after Drill 2 · the version run on 07/14
  4. rev 407/2026From the post-incident review
On screen

The drill history: the first quarterly drill overran, the second met the four-hour target, and the third was the real recovery at 3h 40m, with the runbook revised after each run.

Ways of working

Working inside their operation

A cross-functional team of 5 worked on an ongoing retainer, covering Backup strategy, Recovery planning, Restore drills. We ran daily standups with their own lead in the room, and a demo at the end of every sprint. Scope changed twice during the engagement, and both times the change was priced and agreed before work started.

This was a retainer, because the drills are the deliverable. Quarterly restores run into an isolated environment against the runbook, timed, with the runbook corrected wherever reality disagreed with it. The first drill overran the target, the second met it, and the third wasn't a drill.

What changed in the runbook

  1. 01

    A backup you've never restored from is a hypothesis. A restored one is a safeguard.

    The nightly backups had run successfully for years, and no one could have said whether they would restore. Writing a file successfully isn't the property that matters.

  2. 02

    The runbook mattered as much as the backups. Under pressure, nobody improvises a restore order correctly.

    In a real incident, the order of restoration is where improvisation fails, and the runbook is what takes the improvising away.

  3. 03

    The attack was survivable because this was the fourth time the team had run the restore, and the first under real pressure.

    Three rehearsals turned a four-hour objective into a four-hour outcome. A first attempt under real pressure wouldn't have met it.

The recovery, one step at a time

The order is the plan. Nobody improvises it under pressure.

The real ransomware recovery, replayed through the runbook: what happened at each step, why it had to come before the next, and where the clock stood against the objectives. Pick a step, use the arrow keys once one is focused, or press Replay.

09:38 → 09:49

Every clinic is cut from the network, starting with the primary server. Six sites were cut inside six minutes of 09:38 while the seventh was still connected; nothing is restored until all seven are off.

Why the order matters

A restore onto a network the attacker can still reach is encrypted again. Isolating first is what makes every later step worth doing.

The seven clinics

  • NGNorthgateCut from network
  • RSRiversideCut from network
  • FVFairviewCut from network
  • EPElm ParkCut from network
  • SBStonebridgeCut from network
  • HSHalstonCut from network
  • WBWestburyStill connected

Data lost · recovery point objectivenot known yet / 1h 00m

Clock since declared · recovery time objective11m / 4h 00m

Step 1 of 4 · declared 09:38, restore point 08:34.

Architecture

From the primary server to a restore that's already been proven

Each stage exists to meet one of the two objectives. Copies are made often enough to meet the recovery point, kept where an attacker can't reach them, and restored on a schedule, so the recovery time is a measurement and not a guess.

  1. 01 · Source
    Primary serverPatient records, appointments and billing for all seven clinics, protected to a one-hour recovery point.
  2. 02 · Schedule
    Backup orchestrationScope, frequency and redundancy rebuilt from the objectives, so no more than an hour of work is ever at risk.
  3. 03 · Check
    Verified on writeIntegrity is checked as each copy is written, because a backup nobody has read is only a file.
  4. 04 · State
    Immutable offsite storeImmutable for the retention period and held on separate credentials, so admin access can't delete or encrypt it.
  5. 05 · Proof
    Isolated restoreRestored quarterly into an isolated environment, timed against the runbook and the four-hour objective.

If the attacker goes for the backups too

Backup survival & restore proof

Immutable for the retention period

Copies can't be deleted or encrypted by anyone until their retention period ends, including someone holding administrative credentials. That closes the failure mode where the attacker destroys the backups along with the server.

Offsite, on separate credentials

Copies are held offsite on their own credentials, separate from the primary server's. Before the engagement the backups shared credentials with the server they protected. They don't anymore.

Restores proven every quarter

Each quarter the copies are restored into an isolated environment and timed against the runbook, and the runbook is corrected wherever reality disagreed with it. Integrity is verified as copies are written.

Running your practices on backups nobody has ever restored from? Scope your build in 3 minutes.

Scope your build
Have a project?

Let's talk

Running a large platform, shaping a first MVP, or getting a product ready for a funding round? Tell us where you are. We'll shape the process around it, and stay with you after launch.