Skip to main content

Building a 2,000-GPU AI Networking with Rail-Optimized Architecture and Enterprise SONiC

written by Asterfusion

September 16, 2026

Customer Background

S is a technology company focused on AI computing, providing GPUs and related computing products for AI and high-performance computing (HPC) applications. As the deployment of large-scale AI clusters continues to expand, the customer needed a data center network with high bandwidth, low latency, high reliability, and scalability to support computing and data transfer across large-scale GPU clusters.

Network Requirements

As the scale of large AI model training and high-performance computing (HPC) workloads continues to grow, the customer needed a high-performance data center network for a large-scale GPU cluster. Compared with traditional data center networks, AI clusters place higher demands on bandwidth, latency, congestion control, and network reliability.

First, the network needs to provide high bandwidth and low latency. In large-scale distributed AI training, GPUs frequently exchange gradients, parameters, and intermediate results. Network performance directly affects communication efficiency across the GPU cluster. The network therefore needs to deliver high-throughput, low-latency data transmission and minimize the impact of network communication on GPU utilization.

Second, the network needs to support high-performance Ethernet technologies such as RoCEv2. As the GPU cluster scales, the network architecture needs to further optimize congestion control, traffic scheduling, and lossless transport to ensure stable communication between large numbers of GPU nodes.

Third, the network needs to provide scalability and high reliability. For GPU clusters with thousands of GPUs or more, the network architecture must support horizontal scaling. Designs such as Spine-Leaf, ECMP, and Rail-Optimized can improve network resource utilization and path redundancy, while reducing the impact of individual link or device failures on the AI cluster.

In addition, compute, storage, and business management traffic needs to be effectively isolated. AI GPU servers typically carry compute traffic, storage access, and business and management traffic at the same time. Using independent network planes for traffic isolation reduces interference between different workloads and simplifies network operations and troubleshooting.

Finally, as the AI cluster and the number of network devices continue to grow, the customer also needs centralized network management and operations capabilities. These capabilities should support device status monitoring, link diagnostics, network connectivity testing, and fault isolation to improve operational efficiency in a large-scale AI data center.

Why Asterfusion?

Meeting the combined requirements of large-scale AI clusters for network performance, reliability, scalability, and operational efficiency, S selected Asterfusion for more than the high-performance forwarding capabilities of its switches. Asterfusion also provides an end-to-end solution covering open network operating systems, AI network architecture design, deployment, and operations.

1. Asterfusion builds enterprise-grade open networking solutions based on SONiC.

As the performance and scale requirements of large AI clusters continue to increase, the customer needed a more flexible and scalable network infrastructure. Asterfusion builds AsterNOS on the open-source SONiC network operating system, bringing an open networking architecture to AI data center environments while optimizing its features and performance for AI/HPC workloads. By decoupling network hardware and software, the infrastructure can be deployed and managed using a standardized open networking architecture, providing a flexible foundation for large-scale GPU clusters.

2.Open architecture provides greater flexibility for AI data centers.

Asterfusion adopts an open networking approach across hardware, network operating systems, and automated management. It supports standard network protocols, open interfaces, and automation tools. Based on the SONiC architecture, the customer can reduce reliance on a single proprietary networking stack and extend network functions, integrate automation, and manage operations based on its specific AI networking requirements. This openness also provides greater flexibility for future network expansion, architectural changes, and integration with other infrastructure.

3. Asterfusion has practical experience deploying AI networks.

Large-scale AI networks require systematic optimization of RoCEv2 congestion control, link reliability, and network performance from architecture design through deployment. Asterfusion uses RoCEv2 with technologies such as PFC and ECN to build high-performance networks for AI training. It has also accumulated deployment and validation experience in scenarios such as Rail-Optimized networks and AI Fabric, enabling SONiC and open networking architectures to be applied to large-scale GPU clusters.

4. Asterfusion can address the broader networking requirements across compute, storage, and business management networks.

In this project, the network included multiple network planes rather than a single AI compute fabric. Asterfusion used data center switches with different specifications together with a unified network management platform to build and manage each network independently. The platform also provides connectivity and link diagnostics through Ping, Traceroute, and PRBS, helping operations teams quickly identify link and device issues. This integrated hardware and software approach allows the customer to maintain an open network architecture while simplifying deployment and operations for a large-scale AI data center.

AI Networking Architecture Design

Device Deployment:

Device ModelQuantityNetwork Role
CX864E-N48AI Compute Network
CX732Q-N4Storage Network
CX664D-N30+Storage Network
CX308P-48Y-N-V218Business/Management Network

To meet the networking requirements of a large-scale GPU cluster, Asterfusion designed an AI data center network with separate compute, storage, and business/management networks. Different types of traffic run on independent network planes, reducing interference between compute, storage, and management traffic. This provides a stable, high-bandwidth environment for the GPU cluster while simplifying network operations and troubleshooting.

In the compute network, the project uses a Rail-Optimized architecture. High-speed network interfaces on the GPU servers are connected to their designated Rails, distributing GPU traffic across multiple Rails. Under a standardized cabling scheme, NICs with the same index connect to the corresponding Rail/Leaf. This allows GPU nodes to communicate efficiently through predictable network paths.

The compute network uses a Clos architecture for horizontal scaling. Leaf switches connect to GPU servers, while Spine switches provide interconnectivity between Leafs. This allows the compute network to scale horizontally as the number of GPU nodes increases.

In the storage network, an independent network plane carries data access traffic between the GPU cluster and storage systems. The project uses CX732Q-N and CX664D-N switches for the storage network. The storage network is physically and logically isolated from the compute network to prevent large-scale storage traffic from affecting GPU compute traffic.

The business/management network carries server business traffic, device management, and related operations traffic independently. The project uses CX308P-48Y-N-V2 to build this network, keeping it isolated from the compute and storage networks and providing an independent communication channel for business and management traffic across the AI cluster.

The overall architecture is designed for a single cluster with approximately 2,000 GPUs. The three independent network planes provide traffic isolation across compute, storage, and business/management workloads. Asterfusion also uses a unified AI data center management platform to centrally manage the network devices and provide connectivity testing and troubleshooting capabilities.

64-port OSFP 800GbE switch

800GbE Switch with 64x OSFP Ports, 51.2Tbps, Enterprise SONiC Ready

Please login to request a quote
32port 400G data center switch-1

32-Port 400G QSFP-DD Data Center Switch for AI/ML Enterprise SONiC Ready

Please login to request a quote
64-port 200G QSFP56 data center switch

64-Port 200G QSFP56 Data Center Switch Enterprise SONiC Ready

Please login to request a quote

Test Results

Before the formal deployment, the customer conducted practical tests to validate the overall AI network architecture and switch performance. The tests covered network connectivity, bandwidth performance, and high-speed link quality under actual business scenarios.

AI Networking Performance Test Topology
AI Networking Performance Test Topology

The test results showed that the measured network bandwidth met the customer’s business requirements. The compute network supported the data communication requirements of the large-scale GPU cluster, while the storage and business/management networks operated stably under the three-network separation architecture. Overall network performance and connectivity met the project requirements.

The customer also used the Asterfusion AI Data Center Management Platform to perform network connectivity and link diagnostic tests. Ping and Traceroute tests completed successfully, and the PRBS link test result was rated Excellent, indicating good link quality on the high-speed links.

Through the unified management platform, operations teams can perform network connectivity tests centrally and quickly identify potential link and network issues. This provides operational support for the subsequent deployment and maintenance of the large-scale AI cluster.

Compute_Backend-Configuration-Connectivity Diagnostics-Test Tasks
Compute_Backend-Configuration-Connectivity Diagnostics-History-View Summary

Based on the test results, Asterfusion switches met the customer’s requirements for performance, network connectivity, and link quality, providing a validated foundation for the subsequent deployment of the large-scale GPU cluster.

Project Value and Solution Evolution

Through this project, Asterfusion gained further practical experience in the deployment and validation of large-scale AI cluster networks. The project also drove continued improvements in its networking capabilities for next-generation AI infrastructure.

Based on the project’s experience with Rail-Optimized networking, RoCEv2, three-network separation, and large-scale GPU interconnects, Asterfusion further refined its network architecture and solution capabilities for Supernode and large-scale AI clusters. This provides a technical foundation for deploying larger and denser AI computing clusters in the future.

The project’s testing and deployment experience also further validated Asterfusion’s capabilities in SONiC, open networking architecture, AI Fabric, and unified network operations. These experiences continue to drive the evolution of related products and solutions toward higher bandwidth, lower latency, and larger-scale AI networking.

Conclusion

This project addressed the networking requirements of a large-scale GPU cluster. Based on SONiC and an open networking architecture, Asterfusion built a three-network AI data center architecture that separates compute, storage, and business/management traffic. The Rail-Optimized architecture provides the high-bandwidth, low-latency connectivity required by the GPU cluster.

Before deployment, the customer conducted practical tests covering the network architecture, switch performance, bandwidth, and high-speed link quality. The results showed that network performance and link quality met the project requirements, validating the applicability and stability of the Asterfusion solution in the AI workload environment. The unified AI data center management platform also provided centralized support for connectivity testing, link diagnostics, and ongoing network operations.

Through this project, Asterfusion met the customer’s networking requirements for a large-scale AI cluster while further strengthening its experience in AI Fabric, RoCEv2, Rail-Optimized networking, and large-scale GPU cluster deployment. These capabilities provide a foundation for further developing network solutions for Supernode and next-generation AI data centers.

Latest Posts