Skip to main content

CX764QD-N

64×400G RoCE Switch, 25.6Tbps, Marvell Teralynx 10, For AI/ML/Cloud Data Center/HPC

  • Advanced AI features: WCMP, INT-driven Routing,Packet Spray and more
  • Deliver port-to-port latency of around 560ns.
  • 64 x QSFP-DD ports supporting high-density 400GbE deployment.
  • 25.6 Tbps switching capacity with non-blocking forwarding.
  • Large on chip buffer of 200+MB for better RoCE performance.
  • 10ns PTP and SyncE performance supports.
  • Open AsterNOS based on SONiC with the best SAI support.
  • Line-rate programmability to support evolving UEC standards.
  • ZR/ZR+ module supported
Please login to request a quote

64 x 400G QSFP-DD, 25.6Tbps, Enterprise SONiC, AI/ML/Cloud Data Center/HPC/Storage

Designed to power the most demanding AI training & inference fabrics, cloud data centers & DCI, high-performance computing (HPC), and distributed storage workloads, the Asterfusion CX764QD-N is a game-changing, high-density 400G RoCE switch. Powered by a single-chip 25.6 Tbps ASIC and featuring 64 x 400G QSFP-DD ports, it scales effortlessly to support ultra-large networks with line-rate, non-blocking performance. With an industry-leading ultra-low port-to-port latency of ~560ns, the switch guarantees deterministic, high-throughput connectivity critical for latency-sensitive applications. Furthermore, its integrated support for 400G Coherent ZR/ZR+ optics ensures high-capacity, cost-effective interconnectivity for modern long-reach cloud DCI architectures, eliminating the need for expensive external transponder systems

560ns

Ultra
Low Latency

Outperform IB

27.5% Higher TGR
20.4% Lower Latency

200+ MB

Packet
Buffer

KEY POINTS

Key Facts

560ns

Ultra Low Latency

Outperform IB

27.5% Higher TGR
20.4% Lower Latency

200+ MB

Packet Buffer

Optimized Software
for Advanced Workloads
AsterNOS

Pre-installed with Enterprise SONiC, the CX764QD-N packs cutting-edge features like RoCEv2 and EVPN multihoming, delivering lightning-fast forwarding and massive packet buffers for unmatched performance. Fully UEC-compliant, it offers rich open APIs to seamlessly integrate into diverse data center and HPC ecosystems.
As a vendor-neutral platform, it effortlessly supports heterogeneous GPUs and network cards from multiple manufacturers. By combining industry-leading low latency with expansive packet buffering, the CX764QD-N sets the gold standard for next-generation AI-era data centers.

Optimized Software for Advanced Workloads – AsterNOS

Fully compliant with UEC standards, it offers extensive open APIs for smooth integration into diverse data center and HPC environments. As a vendor-neutral platform, it supports heterogeneous GPUs and network cards from multiple manufacturers. Combining cutting-edge low latency with large packet buffers, the CX764QD-N sets the benchmark as the go-to switch for next-generation AI-era data centers.

RoCEv2

Easy RoCE
PFC
ECN
DCBX

Cloud
Virtualization

DCI
VXLAN
BGP EVPN

In-Band Network
Telemetry

INT Based Routing
Flowlet
WCMP
Packet Spray

High
Availability

ECMP
MC-LAG
EVPN Multihoming

Cloud Virtualization

DCI
VXLAN
BGP EVPN

RoCEv2

 PFC/ECN/DCBX
Easy RoCE

High Availability

ECMP
MC-LAG
EVPN Mulithoming

Ops Automation

REST API/NETCONF/gNMI
In-band Telemetry

Specifications

Ports64x 400G QSFP-DD, 2x 10G SFP+Switch ChipMARVELL TERALYNX 10
Switching Capacity25.6TbpsPort-to-port latency560ns
Packet forwarding rate 28800MppsPacket Buffer200+ MB
CPU Intel Xeon 4/Marvell OCTEON10 CN102
USB1 x USB2.0
RAM16GB/32GB DDR4, up to 96GBConsole1 x Console RJ45
SSD256GB M.2 SATAMGMT1 x MGMT GE RJ45
Hot-swappable Fans4 (3+1Redundancy)Hot-swappable Power Supplies
2 (1+1Redundancy)
Height2UDimensions
(WxHxD mm)
440x87x640
Input voltage200-240V AC
200V~320V DC
Operating temperature0 to 40℃(32 to 104 °F)
Maximum power consumption2100W (with full ports of 400G-SR4)
Relative humidity5% - 90% (non-condensing)
PTP (Optional)Class C (10ns)

Quality Certifications

ISO-9001-SGS
ISO-14001-SGS
ISO-45001-SGS
ISO-IEC-27001-SGS
fcc-icon
ce-icon

Features

Marvell Teralynx 10 Inside: Unleash Industry-Leading Performance & Ultra-Low Latency

With Marvell Teralynx 10 at its core, Asterfusion 400G switches
achieve blazing-fast 560ns latency—perfect for AI/ML, HPC, and NVMe where speed drives results.
Marvell Teralynx 10 chip powering Asterfusion 800G switches with ultra-low latency
25.6Tbps

Switching Capacity

560ns

End-to-end Latency

200+ MB

Buffer Size

51.2 Tbps

Switching Capacity

560ns

Ultra Low Latency

200+ MB

Buffer
Size

Full RoCEv2 Support for Ultra-Low Latency and Lossless Networking

Asterfusion 400G switches deliver full RoCEv2 support with ultra-low latency and
near-zero CPU overhead. Combined with PFC and ECN for advanced congestion control, it enables a true lossless network.
Diagram comparing TCP/IP stack vs RDMA stack, highlighting Asterfusion 800G switches' full RoCEv2 support for ultra-low latency and lossless networking with near-zero CPU overhead
Optimized Features for AI/ML/HPC
Asterfusion Enterprise SONiC is optimized for AI/ML/HPC scenarios, supporting dynamic intelligent traffic distribution (Packet Spray);
Efficient and resilient path selection via WCMP (Weighted Cost Multi-Path);
Real-time precise telemetry routing based on INT Driven Routing;
as well as technologies like Flowlet and Auto Load Balancing to enhance traffic management.
Asterfusion Enterprise SONiC features for AI/ML/HPC networking illustration

Asterfusion Network Monitoring & Visualization

AsterNOS supports Node Exporter to send CPU, traffic, packet loss, latency, and RoCE congestion metrics to Prometheus.
Paired with Grafana, it enables real-time, visual insight into network performance.
AsterNOS network monitoring with Prometheus and Grafana dashboard

O&M in Minutes, Not Days

Automate with Easy RoCE, ZTP, Python and Ansible,  SPAN / ERSPAN Monitoring & more—cutting config time, errors, and costs.
Asterfusion network O&M automation tools diagram

Asterfusion 400G Switch Outperforms InfiniBand in AI Inference with Higher TGR and Lower Latency

In AI inference networks, the Asterfusion 400G RoCE switch delivers higher TGR (Token Generation Rate) and lower P90ITL (90th Percentile Inter-Token Latency) compared to InfiniBand, demonstrating faster inference speed and greater overall throughput performance.
Bar chart comparing Token Generation Rate (TGR) between Asterfusion 800G RoCE switch and InfiniBand switch, showing up to 27.5% higher AI inference performance
Bar chart comparing P90ITL (90th percentile inter-token latency) between Asterfusion 800G RoCE switch and InfiniBand switch, showing up to 20.4% lower AI inference latency
↑27.5%

Token Generation Rate

↓20.4%

Inference Latency

↑27.5%

Token Generation Rate

↓20.4%

Inference Latency

Related Products

CX764QD-NCX764QO-NCX764QH-N
Port Speeds64x400G QSFP-DD, 2x10G SFP+64x400G OSFP, 2x10G SFP+64x400G QSFP112, 1x25G SFP28
Hot-swappable PSUs and FANs1+1 PSUs, 3+1 FANs1+1 PSUs, 3+1 FANs1+1 PSUs, 3+1 FANs
MC-LAG
BGP/MP-BGP
EVPN
BGP EVPN-Multihoming
RoCEv2
PFC
ECN
In-band Telemetry
Packet Spray
Flowlet
WCMP
INT-based Routing
Auto Load Balancing×
SRv6××

Panel

Asterfusion CX764QD-N 64x 400GE QSFP-DD data center switch front panel interface diagram.
Asterfusion CX764QD-N 64x 400GE QSFP-DD data center switch rear panel interface diagram.

What’s in box

800G switch with accessories

Warranty & Support

2-year hardware warranty

1-year free software upgrades & technical support

FAQs

What are the key advantages of selecting the CX764QD-N with its QSFP-DD architecture for AI training and inference fabrics?

The primary advantages of QSFP-DD are its backward compatibility and ecosystem maturity. The CX764QD-N provides 64 ports of 400G QSFP-DD within a compact 2RU space. This design allows infrastructure engineers to seamlessly interconnect with existing 100G/200G QSFP modules and DAC cables, simplifying integration with legacy server NICs and existing network deployments while offering a highly abundant supply chain. Despite its exceptional port density, its maximum power consumption (2100W) and physical dimensions (440x87x760 mm) remain well within the power and cooling thresholds of standard AI server racks.

In AI RoCEv2 networks, how does the CX764QD-N resolve the link imbalances and tail latency issues typically caused by traditional ECMP hashing?

This is a critical bottleneck in AI data center networking. AI workloads feature highly synchronized collective communication patterns (like All-Reduce) that generate few but massive “elephant flows.” Traditional hash-based ECMP frequently routes multiple elephant flows onto the same physical link, resulting in queue congestion and packet loss.

The CX764QD-N natively supports hardware-driven Flowlet-Based Adaptive Routing. Instead of routing based on the entire flow, the hardware detects minute inter-packet gaps to partition elephant flows into granular “flowlets”. It dynamically evaluates the real-time load of available paths and sprays traffic onto the least congested links, eliminating ECMP polarization and significantly reducing Job Completion Time (JCT).

Compared to Broadcom Tomahawk 5, what are Asterfusion’s advantages?

→ Latency: 560ns vs ~800ns
→ Packet Buffer: 200MB vs 165MB

Does it support Inband Telemetry / INT? Can it be used for traffic visualization and bottleneck analysis?

CX764QD-N supports INT. It can report metrics of BDC (Buffer Drop Collector) and HDC (High Delay Collector) to Prometheus, bottleneck analysis can be realized via Grafana.

What 400G optical module standards are supported ?

All OSFP-form 400G modules are supported, including ZR/ZR+ modules.

Why does the CX764QD-N integrate high-precision IEEE 1588v2 PTP, and what specific industry applications does it target?

The CX764QD-N supports IEEE 1588v2 PTP with Class C-level timing accuracy and integrates key industry-specific profiles like SMPTE 2059-2, AES67, ITU-T G.8275.1, and G.8275.2. This allows the 25.6T platform to extend beyond traditional switching into IP Media Broadcasting (ensuring perfect audio/video phase alignment over 400G fabrics) and Telecom Edge Computing (delivering precise clock delivery for carrier-grade synchronization). It flexibly operates across multiple clock roles—including GrandMaster (GM), Boundary Clock (BC), and Ordinary Clock (OC)—to guarantee deterministic, jitter-free timing performance across diverse specialized infrastructures.

Does the system support ISSU (In-Service Software Upgrade), NSR, or GR?

ISSU is not supported since it requires dual control plane module in hardware. AsterNOS supports hot-patch upgrades via containers to minimize service interruption.