An Introduction to RDMA for S3
One of the biggest challenges customers and partners bring to me is how to get their backups written to object storage faster. That challenge is only magnified when they’re pushing hundreds of terabytes of backup data into their object storage repositories every day. In most cases the bottleneck isn’t the disks, especially if the object storage is backed by SSD/flash. The culprit is most likely the TCP/IP network stack used to move data from the backup server to the object storage target. Every object written over HTTPS pays a TCP/IP penalty in CPU cycles, memory copies, and latency, and those penalties really become apparent when high volumes of backup data are being transferred during the backup window.
When we started our object storage journey with Veeam Backup & Replication (VBR) v9.5.4 back in early 2019, we expected object storage to become the predominant platform for secondary storage and it has. Adoption of object storage repositories has climbed steadily ever since and most customers choose Amazon Simple Storage Service (S3) or an S3-compatible platform. For most environments, standard HTTPS over TCP/IP is more than fast enough. But for our larger customers, the ones processing hundreds of TBs or more of backup data per day, TCP/IP overhead becomes a measurable bottleneck. This bottleneck adds time that increases their backup window.
What I typically see customers do is buy their way around it. They deploy 100 GbE networking and/or scale the object storage cluster with extra nodes just to add network interfaces and throughput. That works, but it’s expensive. They’re still paying the TCP/IP overhead. They just spread it across more hardware rather than eliminate it.
Remote Direct Memory Access (RDMA) is a technology that could help dramatically reduce those penalties. RDMA has been used for years in high-performance computing (HPC) and AI infrastructure, and it’s now showing up in object storage as RDMA for S3.
What is RDMA for S3?
Let’s start with “What is RDMA?”
RDMA had its origins in the 1990s, with the initial implementations focused on HPC. In the early 2000s, the industry began standardizing RDMA as a way to overcome performance bottlenecks between HPC systems. RDMA was a perfect fit because it lets one system read from or write to another system’s memory directly, bypassing the remote CPU, its operating system, and the much slower TCP/IP processing layer.
RDMA can use the Ethernet you already own
RoCE (RDMA over Converged Ethernet) lets RDMA run over standard Ethernet networks. It’s the most common way RDMA gets deployed in today’s data centers, because it works with the Ethernet switching most data centers already have. You’ll need RDMA-capable network cards and some switch configuration, but not a whole new network.
Systems transfer data directly from memory to memory across the network, skipping the operating system’s TCP/IP processing and the extra data copies that come with it. The result is lower latency, higher throughput, and less CPU load during large data transfers, like backups being sent to object storage repositories. And you get that performance on familiar Ethernet gear that most likely already exists in your data center, instead of buying into a specialized fabric like InfiniBand.
So what does RDMA for S3 mean?
Nearly every storage platform with an S3-compatible interface talks HTTPS over TCP/IP. This is great for compatibility and interoperability between object storage solutions and software like VBR. Configuring VBR to use object storage is really as simple as pointing the object storage repository to an S3 endpoint URL.
RDMA for S3 would let a backup application like VBR keep using the same S3 APIs it uses today, while changing how the data moves. The S3 requests themselves (the PUTs, GETs, and metadata) would still travel over HTTPS and TCP/IP, so VBR keeps speaking S3. The backup data itself would take a faster route, moving over RDMA straight from the backup server’s memory into the object storage platform’s memory and skipping most of the TCP/IP overhead.
The diagram below compares traditional networking with RDMA for S3 networking in an environment where VBR, or the Veeam Software Appliance (VSA), is writing to object storage:

What Has to Happen First
RDMA for S3 is real, but it’s early. A handful of object storage platforms support it today, each in its own way, and more are building it. From what I’ve seen, two things stand between where we are now and broad adoption.
Fragmented SDKs
Current implementations generally require the application to integrate a vendor-specific SDK, and those SDKs aren’t standardized. Each one is built for a single platform, so an application vendor that wants broad support has to integrate against every one of them separately. That doesn’t scale, for the vendor or for the customer trying to keep their storage options open.
What would really help is RDMA support in the widely used S3 SDKs, with a clean fallback to HTTPS when the target doesn’t support it. That would mean one integration that works across many platforms, with no penalty for storage that hasn’t caught up yet.
The fabric question
RDMA also needs a data center fabric that supports it. Depending on where the backup application sits relative to the storage the customer might need to add RDMA-capable switching in between and potentially additional switches. That adds cost and complexity to the environment. An architecture that shorten the path between the application and the storage reduce that requirement and lessens the cost. The lower cost provides the business case to implement RDMA for S3 as much as the reduction in backup times does.
What This Means for You
None of this changes how you design your backups today, but it’s worth being aware of RDMA for S3 since it will mature and should help with getting your backups to finish within their allotted windows.
The following are some indicators that RDMA for S3 might be a solution of interest for a customer:
- If backups are saturating your current network
- If you’re adding nodes or network capacity purely to hit a backup window
When evaluating object storage platforms, ask where RDMA sits on their roadmap and whether their implementation requires a proprietary SDK. The answer tells you how easily adopted their implementation of RDMA for S3 might be.
While RDMA for S3 does offer the potential to drastically reduce the times writing/reading backups to/from object storage, RDMA results published from high-performance computing don’t 100% translate directly to object storage. Several factors need to be taken into consideration, especially with on-premises S3-compatible solutions. They are:
- object storage architecture
- disk performance
- workload sizes
But RDMA for S3 does provide some benefits that I believe translate well to backups using object storage repositories:
- less CPU overhead
- lower latency
- more throughput on large backups
How much of that affects your backup window depends on your environment.
If you’re bumping into the TCP/IP ceiling in your own environment and have looked into RDMA for S3, I’d like to hear about it in the comments.
