4 min read

The Next Frontier for AI Fabrics: Scale Across Networking

The Next Frontier for AI Fabrics: Scale Across Networking

As AI training workloads and frontier models continue their exponential climb to support trillions of parameters and millions of AI accelerators, a harsh reality is setting in. Physical space constraints and restrictive power availability are a scarcity. This introduces the new dimension of pragmatic reality to move from centralized vertical stacks to horizontal distributed scale-across AI fabrics. Basically, scale-across transforms the long-distance extension of the local scale-up and scale-out AI cluster to achieve high compute density across geographies.

Scale-Across with 7800 AI Spine

The scarcity of compute capacity and megawatts of power mandates that the AI infrastructure must be designed thoughtfully at scale. The Arista 7800 platform continues to be the ideal flagship spine for scale-across applications, providing traffic isolation, contextual routing and security. Scale-across AI innovations deliver many L2/L3 switching features for programmable, deterministic routing, SRv6 multiplane forwarding, multi-tenancy traffic engineering, and load-balancing across regions capable of providing near-instantaneous (tens of uSec) recovery in the event of transient congestion, packet loss, or a physical failure of AI clusters, independent of location. It uses SRv6, or segment routing, which isn't new, but using it to load-balance an AI fabric is the game changer. In SRv6, the sender tags each packet with a stack of SRv6 segment ID's, dictating the exact path the packet will take. The system then uses real-time congestion signaling to dynamically shift packets away from hotspots. Leveraging the reliable, state-sharing Arista EOS as a single, unified operating system, this SRv6 intelligence is supported all the way from the scale-out fabric to the long-distance, scale-across routing. Our customers now get the combination of high scale/performance, reliability, and operational rigor that Arista is known for while connecting to different forms of coherent optics such as ZR/ZR+, and DWDM transport.

Reliable Multi-Site High-Performant Foundation

Stretching an AI compute fabric across distributed geographies isn’t just a matter of provisioning a standard Data Center Interconnect (DCI) link. AI workloads demand massive, highly synchronized, and bursty collective communication flows. If long-haul communications aren’t explicitly architected for these patterns, it becomes a structural bottleneck that severely degrades AI performance. Scale-across AI fabrics are designed not only to optimize performance in best case scenarios, but also to react gracefully when plans go awry and provide highly reliable, secure and uncompromised communication, as shown in Figure1 below.

26-8-27, Slides for Scale-Across Blog2

Figure 1: Arista Scale-Across is based on foundational principles of uncompromised scale, security and reliability

A typical scale-across AI network is designed for reliable, consistent, lossless packet transport, paired with real-time analytics to measure and validate performance. It means optimizing for consistently low end-to-end latency, with intelligent traffic engineering to steer workloads to local vs. remote sites, while incorporating buffering insurance to protect latency by avoiding packet loss during transient congestion. Scale-across builds upon the “hope for the best but prepare for the worst” in demanding AI networks.

“REACH” with Scale-Across Fabrics

Arista enables optimized scale-across fabrics that extend REACH for AI workloads through a combination of foundational tenets that together deliver an essential suite of features for AI operators designing for consistent scale, performance and availability. Operating long-haul, scale-across networks without the protection of deep packet buffering is akin to riding a bike down a steep hill without a helmet. Arista’s REACH solution for scale-across AI fabrics includes:

  • Routing Intelligence: Modern AI fabrics require the intelligence of AI model routing at scale to handle ARP, FIB, and ACLs. The fabric must actively understand the latency and bandwidth profiles of your training jobs, intelligently steer workloads to optimize placement between local and remote domains. If you are pooling distributed compute resources, the scale-across network requires robust service separation to enforce per-tenant prioritization and policy control.
  • Encryption: Once AI workload traffic leaves the confines and security of your local data center, it becomes vulnerable to snooping eyes and malicious actors. Protection of scale-across fabrics with wire-speed, hardware-based encryption, built natively into every port with zero performance degradation must be ubiquitous.
  • Analytics & Availability: Workload-aware observability and availability are two sides of the same coin for intense AI traffic. The need for real-time visibility into ports, packet queues, flows, buffer utilization, and path latency across the entire scale-across AI fabric validates real-time performance. At the same time when an XPU gets hung up, immediate insight to the root cause as well as immediate recovery of the network with SSU (Smart system Upgrades) results in a closed-loop system to recover before training job completion times are impacted.
  • Congestion Protection: Intelligent features to avoid congestion combine with advanced recovery protocols and hierarchical, deep packet buffering to avoid packet loss when congestion does occur. Dropping packets has a far greater negative impact on net job completion time (JCT) than increased latency from transient packet buffering. It’s basic job (completion) insurance. This protects against packet loss under abnormal scenarios for the lowest JCT at global scale.
  • High Radix: Scale-across fabrics demand high density. If a local scale-out fabric might support up to hundreds of thousands of XPUs, scale-across fabrics can reach a rarefied scale of 1 million XPUs. Arista 7800 AI spines rise to the occasion by not only enabling local capacity but also seamlessly expanding AI fabrics to remote data centers. The 7800 switching fabrics guarantee optimal fair traffic distribution between all ports, so that scale-across destinations are first-class citizens alongside local hosts, as shown in Figure 2.

26-8-27, Slides for Scale-Across Blog

Figure 2: Arista REACH for Scale-Across AI Fabrics

Summary: Purpose-Built Network For AI Models

Not all routing is created equal. Legacy routers have many challenges with scale and recovery/convergence time of minutes, which do not meet the requirements of modern AI model routing. Arista’s modern AI routing was natively designed for AI/cloud-scale deployment using purpose-built software principles. It leverages merchant silicon designed for high-speed, lossless AI fabrics with fine-grained insights into AI performance and real-time availability of resources. Based on our philosophy of ONE operating system (EOS) across the entire Arista Etherlink AI portfolio for scale-up, scale-out and scale-across AI fabrics, the network brings consistent throughput for high utilization of expensive compute. Scale-across is designed to deliver uncompromised scale, security and reliability. Welcome to the new world of Arista’s scale-across REACH strategy for global scale without compromise.



 References

AI White Paper

Scale Across White Paper

AI Network Solution Guide

Powering Next-Gen AI Clusters: High-Performance Networking with AMD and Arista

Innovators Video

7800 AI Spine

The Next Frontier for AI Fabrics: Scale Across Networking

The Next Frontier for AI Fabrics: Scale Across Networking

As AI training workloads and frontier models continue their exponential climb to support trillions of parameters and millions of AI accelerators, a...

Read More
Racing Against Machine-Speed Threats: How Arista Is Using AI

Racing Against Machine-Speed Threats: How Arista Is Using AI

Two decades into our journey, quality remains Arista's absolute top priority: networking you can count on. Thus, product security is a first...

Read More
The Unified Edge for a Secure Branch

The Unified Edge for a Secure Branch

The industry has spent the last several years obsessed with securing the cloud. Secure Access Service Edge (SASE), as popularized by Gartner1, has...

Read More