How SoFi Cut Recovery to Minutes Across Five
AWS Regions and Made Backup Pay for Itself
2026.09.10 Eon-LandingPage-1540x660

Sponsored By:

output-onlinepngtools (16)
Thursday, September 10th

11 am ET

Recovery is the assumption every cloud team makes and rarely tests until the worst possible moment. In a 2026 survey of 583 cloud IT leaders, 92% were confident they could restore critical data during a major outage, yet most needed six or more hours to actually pull it off, and a majority found their gaps only after an incident, an audit, or a failed restore. Much of that gap traces back to recovery built on coarse, full-restore snapshots, an approach that strains at multi-region scale and buckles when the trigger is a bad deploy or an automated process gone wrong. For a regulated financial institution, every hour of that gap is exposed customer trust, mounting compliance risk, and lost revenue.

SoFi set out to close it. Running regulated workloads across five AWS regions, the team replaced its snapshot-dependent process with automated, granular recovery built alongside Eon and AWS, and did it in under four weeks with no agents to deploy or maintain. Restores that once took a full day now finish in minutes, precise enough to bring back a single row or file without rolling back everything around it. Retention and compliance rules apply automatically as data changes, so meeting PCI and financial recordkeeping requirements stopped being manual work. The savings came from the architecture itself, enough to cover the platform's cost in its first year.

Join us for this live session featuring Shannon Bradley, Manager of DevOps Engineering at SoFi, who will cover the resilience patterns that keep it working across regions. Shannon walks through the architecture decisions, the move off native snapshots, and the operational habits that turned recovery into something SoFi counts on rather than hopes for. Whether you lead a platform team, run SRE or cloud operations, or own infrastructure strategy, you'll leave with a blueprint for recovery that holds up when you need it most.

Key Takeaways:

- Why recovery built on full-restore snapshots sets a hidden ceiling on speed at multi-region scale, and what precise, targeted restore changes about incident response.

- How to let engineers recover on their own without weakening the retention and audit controls a regulated environment runs on.

- How to make the financial case for recovery, using infrastructure savings that come from architecture instead of added spend.

- What SoFi sequenced first, where the effort actually went, and what they'd tell a team attempting the same today.

Register Below:

We'll send you an email confirmation

Shannon-Bradley-modified

Shannon Bradley

Manager, DevOps Engineering - SoFi