Leading Ransomware Recovery

On a Saturday morning in September 2023, a Play ransomware variant encrypted virtually every Windows system across the company in about two hours. More than 500 endpoints and over 125 virtual servers were down before 8 a.m. I was the Service Desk Supervisor on paper. Within the first hour I was directing the technical response, and I ran it through recovery and the hardening that followed. We paid no ransom, lost no critical production data, and had roughly 80% of operations back within 24 hours. The work during and after the incident is what moved me into an infrastructure leadership role the following year.

What we were facing

The attackers had been in the network for days before they detonated, moving undetected on compromised privileged credentials and a legacy VPN appliance nobody was watching closely. When they pulled the trigger, the blast radius was most of the estate: around 500 endpoints across the plants, including shop-floor QC stations, and about 125 virtual servers. They had done their homework on recovery too, deleting backup control servers and production replicas on the way through. Behind all of it sat more than 5 million files and roughly 5 TB of file-server data that had to come back.

Taking command in the first hour

The title on paper did not matter much at 8 a.m. on a Saturday with production down, so I took operational command of the technical response and worked the active attack first. Two admin accounts were still logged into vCenter running destructive operations, deleting Veeam servers and VM replicas, so I killed those accounts across every system I could still reach, including the Microsoft 365 tenant. From there it was about stopping the spread: an emergency shutdown of the virtual infrastructure, legacy VPN access revoked, and on-site teams directed to physically isolate endpoints that were still being encrypted on the plant floor. I brought in a third-party DFIR firm to run forensics in parallel so the internal team could stay on recovery, and I kept leadership briefed on scope and options throughout, so the business decisions were being made on current information.

Why there was anything to recover

This was a recovery and not a total loss because of a decision made the year before. When we moved from Veeam 10 to Veeam 11, I owned the redesign, and I used it to rebuild the model rather than reuse it: new hardware end to end, the backup infrastructure pulled off the corporate Active Directory domain onto its own identity boundary, and immutability turned on, all audited against 3-2-1-1-0. The old Veeam 10 environment, built by the prior team and still joined to the corporate domain, was encrypted along with everything else. The new environment, sitting behind its own identity boundary, was never reached. Immutability mattered, but identity isolation was the deciding control, because it kept the threat actor from ever authenticating to the backup plane. That was the difference between the week we had and a catastrophic one.

Running the recovery

I stood up a temporary Veeam control server within hours to start VM restores while a hardened production backup environment was built in parallel. Domain controllers and configuration management came back first and were held isolated overnight, to confirm the threat actor was actually gone before anything broader came online. Early on, restoration was being driven ad hoc by business requests, which started creating dependency conflicts, so on day four I switched everyone to the pre-defined DR tier list. Throughput jumped, and we reached roughly 80% of restorable servers by the end of that day. Endpoints were re-imaged in waves, starting with the in-line quality systems so the plants could keep producing, and midway through I recreated the entire administrative account model as tiered accounts, endpoint, shop-floor, server, and domain controller, with no single account holding org-wide rights.

What it cost, and what it did not

All critical systems were restored within 48 hours, with about 80% back inside the first 24. No ransom was paid, and no production-critical data was lost. Because our email was cloud-hosted and untouched, customer service never went fully dark: orders kept being processed manually while the internal systems came back, so the business stayed in contact with its customers throughout. EDI, customer portals, and the other external services were back within roughly 10 days. The whole recovery was done with internal staff plus the DFIR partner, with no drawn-out outside rebuild, which kept the cost of the incident measured in days of effort rather than months of consulting.

Closing the doors that were open

Once operations were stable, I led the program that closed the gaps the attackers had used:

  • Identity and access: removed every legacy admin account, deployed the tiered admin model, put daily-MFA on VPN, rolled out LAPS across endpoints, rotated passwords on human and service accounts, and added alerting on privileged-group changes and interactive service-account logons.
  • Endpoint: replaced signature-based AV with a behavioral EDR/XDR platform across the whole environment.
  • Network: replaced the legacy firewalls and VPN with next-gen firewalls doing Layer 7 and TLS inspection, and segmented the shop-floor, office, and server networks from each other.
  • Management plane: moved backup, virtualization, and management infrastructure off the production AD domain for good.
  • Backups: brought every job under 3-2-1-1-0 and folded lab and shop-floor systems into the same immutable tier as the production VMs.

What I took from it

The recovery worked because of decisions made before anyone knew an attack was coming: isolated backup identity, immutable repositories, a tiered DR plan, and the discipline to keep it current. Under pressure, the value of those was not measured in features but in optionality. Every door closed in advance was one the threat actor could not open during the incident. Leading the response only reinforced it. Incident outcomes are mostly decided in the quiet quarters beforehand, not in the war room.