Skip to main content

Resilience is not just about cyber attacks!

  • July 14, 2026
  • 3 comments
  • 48 views

Stabz
Forum|alt.badge.img+9

Recently, one of our clients experienced a situation that perfectly illustrates why operational resilience must be at the heart of every IT strategy regardless of any external threat.

📖 The Scenario

We were in the middle of building a new infrastructure, which was not yet operational, when the client suffered the loss of both server rooms within 48 hours of each other.

The cause? Not ransomware. Not an intrusion. Simply the simultaneous end-of-life of SSD drives after nearly 8 years of operation.

The tiered storage architecture amplified the impact: as soon as the SSDs reached their endurance limit, the first site collapsed overnight. Investigation quickly confirmed the diagnosis. But the second site, running the same architecture with the same history, followed the same path two days later.

Bad timing? Absolutely.

⚙️ The Operational Response

The situation was all the more complex as the production infrastructure was end-of-life and out of support, we were precisely in the process of building its replacement when the crash occurred.

Here is how we handled the crisis:

✅ Site 1 → Rapid reconstruction of a backup production infrastructure using spare hardware
✅ Site 2 → Reuse of production servers to restart operations in degraded mode
✅ Full restoration of virtual machines across both environments using Veeam
📊 In total, nearly 20 TB of data were restored over this high-intensity weekend.

One key factor that made a difference: the Veeam server was physical. No need to reinstall it, no dependency on a failing infrastructure. It was immediately operational and saved us precious time during the restoration process.

💡 Key Takeaways

This experience highlighted several critical points:

Reconstruction procedures were not sufficiently mastered, particularly on the network side, where unexpected side effects slowed down communication between components.

Crisis documentation and reflexes were not sufficiently established

End-of-life hardware represents an operational risk in its own right, on par with any cyber threat.

Repository storage capacity is critical after a reconstruction: following the infrastructure change, some machines restarted their Veeam backup cycle with Active Full backups as soon as the jobs were rescheduled. Without sufficient storage headroom in the repositories, we could have gone from an infrastructure crisis directly into a backup retention crisis.

Having your infrastructure covered by an active hardware maintenance contract is absolutely essential. Without it, you lose access to manufacturer support, guaranteed response times and priority spare parts delivery. In a crisis situation, every minute counts and sourcing hardware without a support contract can dramatically extend your recovery time, turning what could have been a contained incident into a prolonged outage. This is a non-negotiable pillar of any serious resilience strategy.

The Fundamental Reminder

Backing up is good. Knowing how to restore is better. Having practiced it is essential. And planning the storage capacity behind it is just as critical.

Regularly testing your backups is a given. But simulating complete disaster scenarios including network procedures, communication chains, reconstruction sequences and repository capacity planning is what makes the difference between a managed crisis and a crisis that manages you.

💬 Have you ever simulated a complete infrastructure loss scenario in your organization?

Share your feedback in the comments real-world cases are often the best learning opportunities for the entire community.

3 comments

EdwinMoraal
Forum|alt.badge.img+1
  • Not a newbie anymore
  • July 17, 2026

Hi ​@Stabz, nice topic! Most people underestimate the consequences of simply letting hardware run. Simultaneously purchased equipment from the same vendor leads to a severely underestimated problem: simultaneous end-of-life of SSD drives and/or other components. In previous years, this was already the problem with the 7200 and 15000 RPM drives.


GerhardGibbs
Forum|alt.badge.img+4

Very good point - and I have been using it in my discussions with clients a lot lately where is seems everybody is talking resilience in terms of cyber attacks, yet only 49% of recoveries are because of it. And yes again - many do not pay attention to the surrounding ecosystem (hardware/storage/networking) as much as they should.


Jean.peres.bkp
Forum|alt.badge.img+8

Recovering 20 TB in such a challenging scenario is an impressive achievement and a great reminder that backup and recovery planning must address all types of operational risks, not just security incidents. Thanks for sharing this valuable lesson learned with the community! 👏🚀