400G AI Switch Real-World Performance Test: Validating Lossless Ethernet with H100 GPUs
written by Asterfusion
Table of Contents
Ⅰ. Background and Test Objectives
In large-scale model distributed training and HPC scenarios, communication throughput and network latency between cluster nodes directly affect the upper limit of GPU utilization, measured by Model FLOPs Utilization (MFU). A traditional point-to-point (End-to-End) NIC direct-connect topology can provide near-peak bandwidth and ultra-low latency. However, it lacks the centralized scheduling and scalability of a switched network. As a result, it cannot meet the deployment requirements of large AI Pods with multi-node, high-density Leaf-Spine topologies. High-density 400G AI switches and an open lossless Ethernet (RoCE Fabric) therefore provide a practical approach to achieving both high performance and scalable deployment.
To determine whether introducing a switched network affects end-to-end communication performance, this test uses a NIC direct-connect topology as the baseline. It evaluates the actual forwarding overhead and traffic-handling capabilities of the CX732Q-N switch in a 400G NDR/RoCE environment. The test covers three progressive dimensions: physical link forwarding, NCCL collective communication efficiency, including All-Reduce and All-to-All, and real-world 7B large model training. The goal is to answer two key questions: how much performance loss is introduced by switch-based forwarding compared with direct NIC connectivity, and whether a RoCE AI Fabric can support high-load production training.
Ⅱ. AI Training Testbed and Network Topology for 400G AI Switches

The testbed uses CX732Q-N switch for end-to-end forwarding tests and NCCL tests. A 7B model was also deployed on the same testbed to evaluate large-model training performance and collect application-level performance data.
Hardware Configuration:
- Compute Nodes: 2× Supermicro SYS-821GE-TNHR, each equipped with 8× NVIDIA H100 GPUs
- Network Device: CX732Q-N, 400G AI switch with 32 ports
- NICs: Mellanox ConnectX-7 MCX75310AAS-NEAT (900-9X766-003N-SQ01)
- Optics/Cables: OSFP112 on the NIC side to QSFP-DD on the switch side, connected through the corresponding fiber links
- Test Scenarios:
- Scenario A: End-to-end testing with NIC direct connection and switch-based forwarding
- Scenario B: NCCL testing with a single node using NIC direct connection (Ring, Tree, and CollNet), and two nodes connected through the switch (Ring, Tree, and CollNet)
- Scenario C: Large-model application testing, covering forwarding performance, packet loss, and single-run training performance
Ⅲ. 400G RoCE AI Switch Test Results and Analysis
This validation focuses on the CX732Q-N 32×400G AI switch and covers three layers: end-to-end forwarding, NCCL collective communication, and 7B large-model training. The results show that the CX732Q-N provides near line-rate forwarding in a 400G RoCE AI network, while keeping the additional switching latency below one microsecond. In the two-node H100 NCCL All-Reduce test, cross-switch communication performance was nearly identical to the NIC direct-connect results. Under the 7B model training workload, the switch reached 97% bandwidth utilization with no packet loss observed. These results validate the platform’s end-to-end traffic-handling capability, from physical links and collective communication to real AI workloads.
The CX732Q-N is a 1U switch with 32×400G QSFP-DD ports and a total switching capacity of 12.8 Tbps. This port configuration is suitable for high-density 400G RoCE connectivity in Leaf/Spine networks for AI Pods and GPU clusters. The ConnectX-7 NICs used in the test support data rates of up to 400 Gb/s for InfiniBand and Ethernet. The results therefore reflect the high-speed interconnect capability between H100 nodes and a 400G network.
32-Port 400G QSFP-DD Data Center Switch for AI/ML Enterprise SONiC Ready
The Asterfusion CX732Q-N is designed for the super spine layer in the next-generation cloud data center CLOS network. It offers low latency and boasts a line-rate L2/L3 up to 12.8 Tbps switching performance by Marvell Teralynx chip.
1. End-to-End: Near Line-Rate 400G Forwarding
The end-to-end test verifies whether the CX732Q-N introduces an additional bandwidth bottleneck between GPU servers. The NIC direct-connect results are used as the baseline and compared with traffic forwarded through the switch.
| Test Scenario | Measured Bandwidth | Theoretical Port Rate | Port Utilization | Compared with Direct Connection |
| NIC Direct Connection | 391.95 Gbps | 400 Gbps | 97.99% | Baseline |
| Forwarded Through CX732Q-N | 391.96 Gbps | 400 Gbps | 97.99% | +0.01 Gbps |
| Bandwidth Difference | 0.01 Gbps | — | Approx. 0.003% | Negligible |
The measured bandwidth was 391.95 Gbps with NIC direct connection and 391.96 Gbps with traffic forwarded through the CX732Q-N. The difference was only 0.01 Gbps, or approximately 0.003%. Both scenarios achieved an effective throughput of about 98% of the 400G port rate.
Under the test conditions, the CX732Q-N introduced no measurable throughput loss on the end-to-end 400G link. Compared with a deployment that relies on direct NIC connections to achieve maximum performance, a switched network can maintain near line-rate bandwidth while providing greater port scalability, node density, and topology flexibility.
It is important to distinguish measured throughput from the nominal port rate. The 391.95/391.96 Gbps results represent effective test throughput, while 400 Gbps is the nominal line rate. The approximately 2% difference should not be directly interpreted as performance loss. Effective payload bandwidth is affected by Ethernet frame overhead, protocol headers, FEC, test-tool measurement methods, MTU settings, and the processing capabilities of the sender and receiver. The key finding is that under identical test conditions, forwarding through the CX732Q-N caused no observable reduction in effective throughput.
2. Low-Latency Forwarding: Only 560 ns of Additional Latency
In addition to throughput, latency is a key metric for AI Scale-Out networks. In multi-node H100 training, gradient synchronization, All-Reduce, parameter updates, and small-message control traffic are all affected by end-to-end latency. As the cluster scales, latency variation can translate into GPU idle time and ultimately affect Model FLOPS Utilization (MFU).
| Latency Metric | Measured Latency | Description |
| NIC Direct-Connect End-to-End Latency | 1.95 μs | Measured with two servers directly connected through their NICs |
| Total End-to-End Latency Through Switch | 2.51 μs | Server—CX732Q-N—Server path |
| CX732Q-N Forwarding Latency | 560 ns | 2.51 − 1.95 = 0.56 μs |
With NIC direct connection, the measured end-to-end latency was 1.95 μs. After adding the CX732Q-N as an intermediate forwarding device, the total end-to-end latency increased to 2.51 μs. The additional latency introduced by the switch was therefore:
2.51 μs − 1.95 μs = 0.56 μs = 560 ns
The additional latency of 560 ns shows that the CX732Q-N can provide 400G high-throughput forwarding while keeping the single-hop switching overhead below one microsecond.
For an AI cluster, the significance of this result extends beyond the single-hop latency itself. GPU nodes no longer need to rely on point-to-point connections. They can use a switched network to build a scalable Leaf-Spine or Clos Fabric. This allows the network to maintain low latency and high bandwidth as the number of GPUs increases, additional server racks are added, storage nodes are introduced, or multiple Pods are interconnected.
3. NCCL All-Reduce: Cross-Switch Performance Close to Direct Connection
The end-to-end link test verifies the basic forwarding performance of the switch. However, AI training depends more directly on the efficiency of collective communication libraries such as NCCL. This test therefore evaluates three levels: the single-node GPU communication baseline, two-node NIC direct connection, and two-node communication through the switch. Ring, Tree, and CollNet algorithms were tested under each applicable network path to compare direct and switched connectivity.
In the single-node test, the NCCL results were 478.57 Gbps and 478.64 Gbps, with a difference of only 0.07 Gbps. This indicates good test repeatability and a stable intra-node communication baseline for subsequent two-node testing.
Note: The measured throughput is higher than the rate of a single 400G port. This typically reflects the combined effects of GPU/NVLink/NVSwitch, PCIe, NIC, and the NCCL collective communication measurement methodology.
The test covers three NCCL algorithms: Ring, Tree, and CollNet:
- Ring: Suitable for large messages and stable link saturation during All-Reduce.
- Tree: Typically provides advantages for smaller messages or latency-sensitive operations.
- CollNet: More closely aligned with multi-node and hierarchical communication and can be used to evaluate the coordination between intra-node NVLink/NVSwitch communication and the inter-node 400G network.
Two-Node All-Reduce Results
| NCCL Algorithm | NIC Direct Connection | Forwarded Through CX732Q-N | Absolute Difference | Cross-Switch Performance Retention |
|---|---|---|---|---|
| Ring | 371.27 Gbps | 369.92 Gbps | -1.35 Gbps | 99.64% |
| Tree | 314.95 Gbps | 314.65 Gbps | -0.30 Gbps | 99.90% |
| CollNet | 370.07 Gbps | 370.38 Gbps | +0.31 Gbps | 100.08% |
The cross-switch performance retention is calculated as follows:
Performance retention = (NCCL throughput through CX732Q-N / NCCL throughput with NIC direct connection) × 100%
The test results show:
- Ring: Direct connection achieved 371.27 Gbps, while the switched path achieved 369.92 Gbps, with a performance retention of 99.64%.
- Tree: Direct connection achieved 314.95 Gbps, while the switched path achieved 314.65 Gbps, with a performance retention of 99.90%.
- CollNet: Direct connection achieved 370.07 Gbps, while the switched path achieved 370.38 Gbps. The 0.31 Gbps difference falls within the test variation range, resulting in a performance retention of 100.08%.
The results from all three NCCL algorithms show that the CX732Q-N forwarding path introduced no material All-Reduce throughput loss. Ring and Tree maintained more than 99.6% of direct-connect performance, while the CollNet results were nearly identical.
This indicates that the CX732Q-N can handle both traditional high-throughput traffic and the RoCE communication generated by H100 clusters during distributed training. For data-parallel training workloads that require frequent gradient synchronization, adding a switch hop did not materially reduce NCCL communication efficiency.
4. 7B Model Training: 97% High Bandwidth Utilization with No Packet Loss
To avoid evaluating network performance solely through synthetic traffic or microbenchmarks, this test also ran a 7B-parameter large-model training workload on the H100 platform. The model configuration included 32 layers, a hidden size of 4608, an FFN hidden size of 18432, and 36 attention heads. Group Query Attention (GQA) was enabled with 4 query groups.
During model execution, the CX732Q-N reported the following results:
| Metric | Measured Result | Description |
| AI Switch Bandwidth Utilization | 97% | Switch forwarding load during model training |
| Packet Loss | 0 | No packet loss observed during the test period |
| RoCE Training Performance | 5874.79 TGS | Tokens/GPU/second, measuring model training throughput. Under the 7B model configuration, each H100 GPU processed approximately 5874.79 training tokens per second. |
| InfiniBand Training Performance | 6819.27 TGS | Tokens/GPU/second, measuring model training throughput. Under the 7B model configuration, each GPU processed 6819.27 training tokens per second. |
| RoCE Relative Performance | 86.15% | Calculated with the InfiniBand result as the baseline |
The single-run training performance was 5874.79 TGS with RoCE and 6819.27 TGS with InfiniBand. Using the InfiniBand result as the reference, RoCE achieved approximately 86.15% of the training throughput:
(5874.79 / 6819.27) × 100% = 86.15%
During training, switch bandwidth utilization reached 97%, with no packet loss observed. This indicates that the network performed well under an actual model training workload, rather than only under low-load or idealized benchmark conditions. The switch maintained stable forwarding under traffic approaching full utilization.
For a RoCE AI Fabric, zero packet loss means more than basic link availability. RoCE relies on lossless or near-lossless transport mechanisms. Persistent packet loss, retransmissions, congestion propagation, or PFC pause propagation can cause NCCL communication latency variation, GPU idle time, and lower training throughput. No packet loss was observed at 97% utilization in this test. This shows that, under the tested two-node topology, 7B model workload, configuration, and test period, the CX732Q-N was able to stably carry high-intensity training traffic. Under the tested topology and workload conditions, the switch achieved zero observed packet loss at high bandwidth utilization.
There was still an approximately 13.85% difference in training throughput between RoCE and InfiniBand. This difference cannot be attributed to the switch alone. End-to-end training performance is affected by factors such as NCCL version and parameters, GPU parallelism, CPU/PCIe/NVLink topology, NIC drivers and firmware, RoCE PFC/ECN/DCQCN configuration, congestion conditions, storage and data-loading paths, as well as model batch size, sequence length, and mixed-precision strategy. Although the ConnectX-7 supports network speeds of up to 400 Gb/s, the final training TGS reflects the overall efficiency of the AI system.
For more 400G switch performance data in AI inference scenarios, please refer to AI Inference Network.
Ⅳ. NCCL All-Reduce and All-to-All Validation
The following section presents screenshots from the actual test results.
1. NCCL Test – Two-Node All-Reduce with NIC Direct Connection (NCCL_ALGO=ring)
This screenshot shows the actual terminal output from a two-node All-Reduce test using a NIC direct-connect topology with the NCCL_ALGO=ring algorithm.
The test used message sizes ranging from 512 B to 8.5 GB.
As shown in the terminal output, with an 8.5 GB message size (8,589,934,592 Bytes), the measured Out-of-place bus bandwidth (busbw) reached 371.27 Gbps. This corresponds to the two-node Ring direct-connect baseline data in Section 3.3. The test completed with no validation errors (#wrong = 0). This result is therefore used as the baseline for evaluating the forwarding performance of the CX732Q-N switch.

2. NCCL Test – Two-Node All-Reduce Across the Switch (NCCL_ALGO=ring)
This screenshot shows the actual terminal output from a two-node All-Reduce test with traffic forwarded through the CX732Q-N switch using the NCCL_ALGO=ring algorithm.
Key Results: As shown in the terminal output, with a 16 GB message size (17,179,869,184 Bytes), the measured Out-of-place bus bandwidth (busbw) reached 369.92 Gbps. This corresponds to the cross-switch Ring result in Section 3.3. Compared with the direct-connect baseline of 371.27 Gbps, the performance retention was 99.64%, confirming that the switched path introduced almost no bandwidth degradation.

3. NCCL Test – Two-Node All-to-All with NIC Direct Connection (NCCL_ALGO=ring)
In addition to conventional All-Reduce collective communication for gradient synchronization, All-to-All communication is frequently used in mainstream Mixture of Experts (MoE) models and sequence parallelism. The following results show the actual test performance of this communication pattern.
This screenshot shows the actual terminal output from the alltoall_perf microbenchmark running on two nodes (gpul-node2 and gpul-node3), with a total of 16 H100 GPUs. The nodes use a NIC direct-connect topology and run the test with NCCL_ALGO=ring.
As shown in the terminal output, with a 16 GB message size (17,179,869,184 Bytes), the measured In-place bus bandwidth (busbw) reached 79.29 GB/s (634.32 Gbps), while the Out-of-place mode reached 79.09 GB/s (632.72 Gbps). Since All-to-All is a fully interleaved one-to-one many-to-many communication pattern, its performance directly reflects the maximum traffic-handling capability when the physical links between nodes are fully utilized. These results are therefore used as the direct-connect baseline for evaluating cross-switch forwarding performance.

4. NCCL Test – Two-Node All-to-All with Cross-Switch Forwarding (NCCL_ALGO=ring)
This screenshot shows the actual terminal output from the alltoall_perf microbenchmark running on two nodes (gpul-node2 and gpul-node3), with a total of 16 H100 GPUs. The nodes are connected through the CX732Q-N switch and run the test with NCCL_ALGO=ring.
As shown in the terminal output, with a 16 GB message size (17,179,869,184 Bytes), the measured Out-of-place bus bandwidth (busbw) reached 79.40 GB/s (635.20 Gbps), while the In-place mode reached 79.39 GB/s (635.12 Gbps). Compared with the direct-connect result of 79.09 GB/s, the cross-switch result was nearly identical. Due to minor test variation, it was even approximately 0.39% higher. These results demonstrate that the CX732Q-N can maintain near line-rate forwarding with no packet loss observed under high-density, fully interleaved All-to-All traffic.

Note: NCCL uses different scaling factors when calculating All-to-All and All-Reduce bandwidth, which results in significantly different reported values. For this comparison, the key metric is the difference in busbw between the direct-connect and cross-switch configurations under the same All-to-All test. This comparison shows that the CX732Q-N did not introduce a performance bottleneck when handling fully interleaved All-to-All traffic.
Ⅴ. Summary
This test establishes a complete validation path from basic network forwarding and GPU collective communication to real-world model training workloads:
- Near line-rate end-to-end forwarding: The bandwidth through the CX732Q-N reached 391.96 Gbps, nearly identical to the 391.95 Gbps achieved with NIC direct connection. This represents approximately 98% of the nominal 400G port rate.
- Sub-microsecond additional latency: The switch introduced approximately 560 ns of forwarding latency, balancing scalable network connectivity with low-latency communication.
- NCCL performance close to direct connection: The cross-switch performance retention for two-node Ring, Tree, and CollNet reached 99.64%, 99.90%, and 100.08%, respectively.
- Stable performance under high-load training: During 7B model training, network bandwidth utilization reached 97%, with no packet loss observed during the test.
- Foundation for RoCE AI Fabric deployment: RoCE training achieved 5874.79 TGS, or 86.15% of the corresponding InfiniBand result. This result provides a baseline for further RoCE parameter optimization, NCCL tuning, and scaling to larger GPU clusters.
For AI architects starting with 8-GPU nodes and planning to scale to multiple Pods, the AsterFusion CX732Q-N demonstrates that an open Ethernet/RoCE architecture can maintain performance close to the limits of NIC direct connectivity while providing high-density connectivity and operational manageability. It provides a practical foundation for building a high-performance, cost-effective AI Fabric.
Explore More in Our 400G Technical Series
Building a resilient, high-performance 400G infrastructure requires a unified understanding of standards, interconnects, and hardware choices. Whether you are scaling an AI data center or upgrading campus backbones, explore our comprehensive 400G guide series to master every layer of the architecture:
- 1. Fundamentals: Learn core standards and deployment models in What Is 400G Ethernet? Standards, Technologies, and Deployment in Modern AI Data Centers.
- 2. Transceiver Specifications: Compare reaches, fiber types, and form factors in 400G Transceivers Explained: VR4, SR4, SR8, DR4, FR4, LR4, ER4 and ZR.
- 3. Coherent DCI Solutions: Explore how 400G ZR and coherent optics enable direct switch-to-switch data center interconnects in What Is 400ZR? How 400G Switches Enable Coherent DCI.
- 4. Switch Architecture: Solve bandwidth and latency bottlenecks by choosing the right platform in How to Choose a 64-Port 400G AI Switch: QSFP-DD vs. OSFP vs. QSFP112.
- 5. 400G QSFP-DD Cable Connection for AI Fabric: 400G QSFP-DD Cable Connectivity for AI Fabrics: DAC vs. ACC vs. AEC vs. AOC
- 6. AI Checkpoint Traffic: How 400G Switches Prevent Network Bottlenecks in GPU Training
- 7. 400G AI Switch Real-World Performance Test: Validating Lossless Ethernet with H100 GPUs



