The Network Is the Computer

When a training cluster is running, GPUs don't work in isolation. Every backward pass requires each accelerator to exchange gradient updates with its peers — a synchronization step called all-reduce that touches every node in the cluster simultaneously. At scale, this east-west traffic is enormous, continuous, and brutally latency-sensitive. A millisecond of stall, multiplied across thousands of GPUs billing at several dollars an hour each, is an expensive proposition. The fabric connecting those GPUs isn't peripheral infrastructure; it functionally defines the cluster's usable throughput.

That fabric debate has two main contenders: InfiniBand, the long-dominant interconnect borrowed from high-performance computing, and high-speed Ethernet augmented with RoCE (RDMA over Converged Ethernet), which borrows InfiniBand's low-latency transport semantics and runs them over a more familiar switching stack. Both camps have serious engineering behind them. Neither has decisively won.

The fabric connecting those GPUs isn’t peripheral infrastructure; it functionally defines the cluster’s usable throughput.

InfiniBand: Built for This Problem

InfiniBand's core properties — lossless delivery, very low latency, and high message-rate — were engineered decades ago for tightly coupled scientific computing, and they map almost perfectly onto AI training's communication patterns. The protocol carries RDMA natively, allowing a GPU on one node to write directly into the memory of a GPU on another without burdening the host CPU. Congestion is managed inside the fabric rather than left to transport-layer retransmits, which matters enormously when thousands of endpoints are communicating simultaneously.

The architecture dominates frontier AI clusters. NVIDIA's NVLink connects GPUs within a node; InfiniBand connects nodes across the cluster, with link speeds now at 400 Gb/s per port and a roadmap pushing to 800 Gb/s. Fat-tree topologies — the physical realization of the spine-leaf concept — give every GPU pair the same non-blocking bandwidth regardless of where they sit in the rack hierarchy, a property essential for all-reduce to complete predictably.

The cost is real. InfiniBand switches and host channel adapters command a significant premium over commodity Ethernet gear, the ecosystem is narrower, and at very large cluster sizes the cabling complexity is formidable.

Switch faceplate fully populated with 400G QSFP-DD transceivers, port LEDs in steady green
Built for this problem: lossless, low-latency, and unapologetically specialised.

Ethernet + RoCE: The Challenger Case

The appeal of Ethernet is obvious: a massive installed base, aggressive competition among switch vendors, and generations of operational tooling that network engineers already know. The objection has always been that Ethernet, left to its own devices, drops packets under congestion. Dropped packets force retransmission; retransmission introduces unpredictable latency; unpredictable latency stalls all-reduce. For AI training, this is intolerable.

RoCE addresses the dropping problem by requiring the underlying Ethernet to be lossless — achieved through Priority Flow Control and Explicit Congestion Notification. When the fabric is properly configured, RoCE delivers RDMA semantics over standard switches, narrowing the latency gap with InfiniBand to something many operators find acceptable. Meanwhile, Ethernet link speeds have kept pace: 400 GbE is deployable today, 800 GbE is arriving, and the same spine-leaf, non-blocking fat-tree topologies apply.

The operational catch is that "properly configured" does real work in that sentence. RoCE's lossless behavior depends on careful end-to-end tuning across switches, NICs, and the application stack. Misconfiguration can produce subtle, hard-to-diagnose head-of-line blocking. InfiniBand's losslessness, by contrast, is inherent to the protocol — it is simply what the fabric does.

The fabric decision
InfiniBandEthernet + RoCE
LineageHPC-native, designed losslessthe ubiquitous standard, hardened for AI
Loss behaviourlossless by designlossless when engineered — PFC, ECN, tuning
Ecosystemsingle-vendor gravitybroad merchant silicon, familiar tooling
Operational riskspecialist skillssubtle congestion tuning; misconfiguration hurts

Where the Industry Stands

Hyperscalers have invested heavily in bespoke Ethernet fabrics, partly to avoid single-vendor dependence and partly because at their scale, owning the full switch ASIC and software stack creates efficiencies no off-the-shelf solution can match. Several large AI labs run production training on RoCE. At the same time, the largest publicly known GPU clusters — including those built around NVIDIA's DGX SuperPOD architecture — continue to use InfiniBand as the inter-node fabric, precisely because its losslessness is guaranteed rather than configured.

The link-speed race has made the gap between the two smaller than it was five years ago, but the fundamental tradeoff remains: InfiniBand offers engineered losslessness and mature RDMA tooling at a premium; Ethernet offers cost, ecosystem breadth, and flexibility, in exchange for operational complexity that grows with cluster size. For the organizations spending hundreds of millions of dollars on GPU clusters, the fabric choice is not a procurement line item — it is an architectural commitment that shapes everything above it.