3 min read

Next-Gen Scale-Up Networking for AI Fabrics

Next-Gen Scale-Up Networking for AI Fabrics

AI infrastructure is exploding beyond single-server boundaries. As models grow and workloads demand tightly synchronized communication across larger accelerator pools, local PCIe, CXL, and proprietary interconnects are no longer sufficient as the primary scale-up mechanism. They remain important within the server, but can create constrained compute islands across accelerator generations, system designs, and vendors. Operators are moving toward rack-scale scale-up domains, interconnecting tens to hundreds of XPUs across enclosures with high-bandwidth, low-latency links to operate as a single, tightly coupled supercomputing instance.

Expanding scale-up domains to hundreds of XPUs fundamentally shifts connectivity demands and creates a new compute+network-centric paradigm. The network is no longer simply the fabric between servers; it becomes part of the compute system itself. Just as Ethernet has become the leading open foundation for AI scale-out fabrics, the industry is now advancing Ethernet-based scale-up solutions for performance, openness, interoperability, and ecosystem choice.

Evolution from Server I/O → Rack-Scale

Recognizing the critical need for an open scale-up ecosystem, Arista galvanized the Ethernet for Scale-Up Networks (ESUN) initiative within OCP as a founding member alongside industry leaders such as Broadcom, Meta, and Microsoft. ESUN is focused on open Ethernet switching and framing for tightly coupled AI fabrics, including lossless delivery, error resiliency, and scale. A rack-scale scale-up fabric must deliver:

    • High Radix & Single-Hop Topologies: Support scale-up domains spanning beyond 72 or 144 XPUs to hundreds and ultimately thousands of accelerators while minimizing hop count.
    • Lossless Transport for Low Latency: Incorporating congestion control ensures deterministic, in-order packet delivery for tensor-parallel collective operations without tail-latency spikes.
    • Reliable 200G/Lane Systems: Engineered specifically for 224G (and beyond) signaling, optimizing Signal Integrity (SI) and Power Integrity (PI) while deploying advanced liquid cooling for high-density.
    • Advanced Diagnostic and Telemetry: A robust hardware diagnostic and telemetry layer to guarantee bit-error correction, secure boot, and physical validation whether running Arista EOS® or an open NOS like SONiC.
    • Robust and Resilient NOS: Massive bandwidth demands and XPU uptime require modern, robust, secure, battle-tested software running atop the network.
Three Rack-Scale Options

To accommodate diverse datacenter footprints, thermal profiles, and accelerator architectures, Arista has pioneered three specialized liquid-cooled physical scale-up rack solutions in partnership with AMD, Arm, Broadcom, d-Matrix, Meta, Microsoft, and Qualcomm. Inspired by OCP’s open infrastructure direction, these designs provide a flexible path to integrate compute, networking, power, and liquid cooling across evolving high-density AI deployments.

1. Orthogonal Chassis Design: Using direct orthogonal connectivity between accelerator and switch blades, supporting accelerator density of 144 XPUs in 100 kW to 400 kW direct liquid cooling envelopes.

2. Cabled Backplane Design: Combines a high-density cabled backplane with a modular rack structure to enable a highly serviceable design.

3. Cross-Rack Design: A modular architecture providing maximum flexibility to scale compute and networking across multi-rack deployments, enabling larger scaling domains across scale-up and scale-out.

Scale-Up-AI-Fabrics-Blog-1

Figure 1: Illustration of three rack-scale options

Arista Etherlink for Scale-Up: Supporting 144 Accelerators and Beyond

Arista's rack-scale Etherlink network architecture, designated as Etherlink SU-144, supports unified scale-up domains of up to 144 accelerators. With a cross-rack scale-up architecture, the number of XPUs grows to 1024 using multi-rack interconnects. Arista’s open architecture stands in contrast to proprietary scale-up configurations based on a locked-in accelerator and interconnect ecosystem. ESUN provides network operators flexibility across accelerators, switching silicon, optics, physical rack designs, and network software while retaining the high bandwidth, low latency, and reliable in-order delivery required for collective communication.

The Powerful Role of Diagnostics

At rack scale, raw bandwidth only matters if the physical system can be validated, monitored, and serviced predictably. Arista Netdi (Network Diagnostics Infrastructure) provides a common diagnostics and telemetry foundation at both the switch and rack level. Built on more than 8,000 person-years of development, it provides deep hardware-level validation, signal integrity analysis, secure boot attestation, and Single Event Upset (SEU) resiliency, supporting rich telemetry for the switch, physical optics, power shelves, and liquid cooling infrastructure. Netdi is entirely NOS-agnostic, giving customers full operational insight no matter which NOS they run. While Arista champions the freedom of NOS choice, Arista’s EOS remains the industry's leading resilient, secure NOS across all AI fabric roles, including scale-up networks, especially in an era of AI-enabled adversaries looking to exploit software vulnerabilities.

Arista provides an open path for cloud and enterprise operators to scale heterogeneous AI supercomputing infrastructures without vendor lock-in. It achieves this by combining standards-based ESUN transport, flexible physical topologies spanning orthogonal, cabled backplane, and cross-rack architectures, Etherlink SU-144 scale-up domains, and Arista Netdi™ hardware diagnostics. Welcome to the new era of rack scale-up networking!

 References

AI Innovators video

The Many Facets of AI Fabrics

AI white paper

Netdi white paper

Infrastructure Security Webinar

Press Release

Visit us at OCP booth number: E61

Next-Gen Scale-Up Networking for AI Fabrics

Next-Gen Scale-Up Networking for AI Fabrics

AI infrastructure is exploding beyond single-server boundaries. As models grow and workloads demand tightly synchronized communication across larger...

Read More
The Next Frontier for AI Fabrics: Scale Across Networking

The Next Frontier for AI Fabrics: Scale Across Networking

As AI training workloads and frontier models continue their exponential climb to support trillions of parameters and millions of AI accelerators, a...

Read More
Racing Against Machine-Speed Threats: How Arista Is Using AI

Racing Against Machine-Speed Threats: How Arista Is Using AI

Two decades into our journey, quality remains Arista's absolute top priority: networking you can count on. Thus, product security is a first...

Read More