Skip to main content

What Is a Border Leaf Switch? How to Choose Border Leaf Switches for AI Data Center Fabrics

written by Asterfusion

August 14, 2026

Introduction

In previous articles, we have discussed Spine and Leaf switches in Clos architecture. However, there is another important role: the Border Leaf switch.

In this article, we will explain what a Border Leaf switch is and why it is needed in a folded Clos architecture. We will also examine how to select the right Border Leaf for an AI data center fabric, including the hardware performance and software features it should provide.

What Is A Border Leaf Switch

A border leaf switch is a leaf-layer switch dedicated to providing connectivity between the internal data center fabric and external networks or service devices.

Like a normal leaf switch, a border leaf connects southbound to multiple spine switches. Unlike a server-facing leaf, its northbound-facing ports usually connect to devices outside the local fabric, including:

  • Data center border routers and WAN routers.
  • DCI routers or optical transport systems.
  • Internet edge routers.
  • Firewalls, NAT gateways, and security service chains.
  • Load balancers and application delivery controllers.
  • Shared services, such as DNS, DHCP, NTP/PTP, telemetry collectors, or authentication systems.
  • Legacy networks, enterprise networks, and external storage or backup networks.
  • Another data center fabric or a cloud on-ramp.

In an EVPN-VXLAN fabric, the border leaf commonly serves as the Layer 3 boundary between the VXLAN/EVPN domain and external IP routing domains. It exchanges EVPN control-plane information with the fabric while also running conventional IP routing—often eBGP—with external routers or service devices

The key point is that “border leaf” describes a network role, not necessarily a unique hardware category. It can be the same platform family as a regular leaf switch. What differentiates it is its placement, interface mix, routing scale, policy requirements, and responsibility for external connectivity.

Why Is a Border Leaf Switch Necessary?

In a large-scale network, when deploying Border Leaf switches, network architects may ask: Why not connect the WAN, firewall, or DCI directly to the Spine switches? Why not connect them to regular Compute Leaf switches? Why deploy a dedicated pair of Border Leaf switches?

There are several reasons to deploy dedicated Border Leaf switches.

Functional Domain Separation

The primary role of a regular Leaf is to connect endpoints to the Fabric, while the primary role of a Border Leaf Switch is to connect the Fabric to external networks, such as the WAN.

If all these connections are placed on regular Leaf switches, a single Leaf may gradually take on multiple roles, including Server/AI workload connectivity, Overlay/VTEP, and external routing. Dedicated Border Leaf switches separate these functions. The value is not simply “adding two more switches.” It is about clearly defining functional domains. The Compute Fabric can focus on high-performance East-West traffic forwarding, while the Border Leaf Swiches handle North-South traffic ingress and egress, including external connectivity and routing.

Simplified Configuration and Operations

If every Leaf connects to external networks, you may need to configure BGP peering, detailed Import/Export Policies, VRFs, inbound ACLs, and integrations with stateful firewalls and NAT across a large number of Leaf switches.

As the network grows, configuration and troubleshooting become significantly more complex.

This is especially important in AI data centers, where the number of GPU clusters and Leaf switches can grow rapidly. With external connectivity distributed across the Fabric, security and policy boundaries expand with the Fabric. A Dedicated Border Leaf consolidates external routing logic, security policies, and BGP sessions. This significantly reduces overall configuration and operational complexity.

Fault Domain and Blast Radius Control

Network failures are unavoidable in real-world deployments. External interfaces are exposed to more sources of uncertainty, including WAN route flapping, external DDoS traffic, aging transceivers, and configuration errors.

If external connections are placed directly on Compute Leaf switches that host critical GPU or server workloads, an external link failure can cause excessive CPU utilization or significant TCAM churn on the Leaf. This can affect the compute nodes connected to that Leaf.

Dedicated Border Leaf switches isolate external connectivity in a separate device group. Even if the Border Leaf fails, the impact is limited to North-South communication with external networks. Internal East-West traffic within the data center can continue to operate independently.

Prevent the Spine from Becoming a “Do-Everything” Device

The primary role of the Spine should be to provide a high-speed, non-blocking L3 IP underlay for forwarding traffic between Leaf switches.

If the WAN, campus core, or Internet is connected directly to the Spine, this role separation is weakened. The Spine may need to handle external BGP sessions, complex route redistribution, edge security filtering, or even VTEP encapsulation. These additional control-plane and overlay functions can increase the operational and resource burden on the data center backbone.

Deploying Border Leaf switches at the edge of the Fabric allows the Spine to focus on what it does best: line-rate, non-blocking underlay forwarding.

Control Cost, Device Count, Port Count, and BGP Session Scale

In theory, multiple or even all Leaf switches can connect directly to upstream routers and use BGP and ECMP for multipath forwarding. This model is technically feasible. However, the scale and cost need to be considered.

For example, assume a data center Fabric has 100 Leaf switches and a pair of upstream WAN routers. If every Leaf connects directly to the upstream routers, the routers need to provide hundreds or even thousands of high-speed interfaces. The resulting fiber cabling also becomes highly complex, while the routers must handle hundreds of redundant BGP sessions.

The cost of high-density router interfaces and long-distance optical transceivers can quickly exceed the hardware investment required for a pair of Dedicated Border Leaf switches.

Therefore, from both the physical topology and hardware cost perspectives, connecting every Leaf directly to upstream core routers is usually not the optimal design. A more practical approach is to designate a small number of Leaf switches as Fabric exit points.

Note: Whether a dedicated Border Leaf is required depends on the network topology and scale. In a small network, regular Leaf switches can perform the Border Leaf role. In some deployments, Spine switches can also function as Border Leaf switches. A dedicated Border Leaf is not required in every network.

How to Choose Border Leaf Switches

The right border leaf is selected by workload and topology requirements, not simply by choosing “the biggest switch.” Evaluate the following hardware and software factors.

Hardware Requirements

1. Fabric-Facing Bandwidth

The border leaf must have enough high-speed uplink capacity to connect to the spine layer without creating an unintended bottleneck.

For an AI fabric, consider:

  • The number of spine switches to which each border leaf connects.
  • Required uplink speed: 100GbE, 400GbE, 800GbE, or higher.
  • Whether all spine-facing links must operate at line rate.
  • Expected concurrent traffic to DCI, storage, inference ingress, and shared services.
  • Failure capacity after losing one border leaf, one spine, or one uplink bundle.

A practical requirement is N+1 capacity: after one border leaf fails, the surviving border leaf—or remaining border leaf group—should still support the required critical north-south traffic.

2. Port Flexibility and Breakout Support

Border leafs often need a more diverse port mix than GPU-facing leaf switches.

Look for support for:

  • 10/25/50/100GbE ports for legacy or service-facing devices.
  • 100/200/400/800GbE ports for modern routers, DCI, and high-speed storage.
  • Breakout modes such as 1×400GbE to 4×100GbE, or 1×800GbE to 2×400GbE or 8×100GbE, subject to platform support.
  • Both optical and DAC/AOC deployment options.
  • Sufficient port density to avoid using unnecessary aggregation layers.

This flexibility is particularly valuable where a high-speed AI fabric must attach to lower-speed firewalls, WAN routers, or existing enterprise infrastructure.

3. Packet Buffer and Congestion Behavior

Border traffic can be more bursty than GPU collective traffic. The switch should provide sufficient buffering and robust congestion-management capabilities for the expected mix of routed, encrypted, inspected, or load-balanced traffic.

Evaluate:

  • Shared and per-port buffer architecture.
  • Egress queue depth and scheduling.
  • WRED/ECN support.
  • Priority flow control support if lossless RoCE traffic must traverse the device.
  • Congestion telemetry and queue visibility.
  • QoS classification, remarking, policing, and shaping capabilities.

Do not assume that a deep-buffer switch is automatically best. Deep buffering can be useful for burst absorption, but it can also add latency. The correct choice depends on whether the border carries latency-sensitive AI flows, storage flows, Internet-facing traffic, or WAN/DCI traffic.

4. Throughput and Forwarding Scale

Confirm that the switch can forward at line rate with all required features enabled.

The specification review should include:

  • Total switching capacity.
  • Forwarding rate in packets per second.
  • IPv4 and IPv6 route scale.
  • ARP and ND scale.
  • MAC table scale, if Layer 2 extension is required.
  • VRF and VLAN/VXLAN segment scale.
  • ECMP next-hop scale.
  • ACL, QoS, NAT, and tunnel scale where relevant.
  • Multicast scale, if AI data distribution or other multicast services are involved.

The relevant question is not simply “How many routes can the switch store?” It is “Can it support the required scale while running the full policy, EVPN, telemetry, QoS, and security-adjacent feature set used in production?”

5. High Availability and Physical Resilience

A border leaf is a critical connection point. Choose platforms with:

  • Redundant, hot-swappable power supplies.
  • Redundant fan modules.
  • Front-to-back or back-to-front airflow compatible with the rack design.
  • Separate management interfaces.
  • Secure boot, hardware root of trust, and image integrity features.
  • Appropriate environmental and power specifications for the AI data center.

Deploying two border leaf switches is normally the baseline. For high-scale environments, multiple border leaf pairs can be used as border pods to scale external bandwidth and isolate services.

Software and Control-Plane Requirements

1. BGP, EVPN, and VXLAN Interoperability

Modern AI data center fabrics commonly use an IP underlay with BGP and may use EVPN-VXLAN for network virtualization. The border leaf should support the required control-plane functions, including:

  • eBGP or iBGP underlay peering.
  • MP-BGP EVPN.
  • EVPN Type-5 IP prefix routes, if external IP prefix distribution is required.
  • VXLAN routing and bridging functions, as needed.
  • VRF-aware routing.
  • Route-target import and export policies.
  • External route advertisement and filtering.
  • Graceful restart and BFD for rapid failure detection.

If the AI back-end fabric is a pure Layer 3 rail-optimized fabric without VXLAN, the border leaf may need only high-scale IP routing. But the platform should still match the fabric’s routing design and operational model.

2. Route Policy and Segmentation

A border leaf must enforce clear route boundaries. Important capabilities include:

  • Prefix lists and route maps/policies.
  • BGP communities and large communities.
  • Route filtering and maximum-prefix protection.
  • VRF route leaking only where explicitly required.
  • Default-route control.
  • Policy-based routing where justified.
  • ACLs for north-south segmentation.
  • IPv4 and IPv6 parity.

For example, GPU workload subnets may require access to a model repository or object-storage network but should not automatically receive unrestricted reachability to enterprise user networks or Internet egress.

3. QoS and RoCE Awareness

For AI fabrics using RoCE, ensure the switch supports the required QoS behavior end to end when relevant traffic crosses the border.

Key capabilities may include:

  • DSCP and 802.1p classification.
  • Priority queues and strict-priority scheduling.
  • ECN marking.
  • PFC configuration and watchdog mechanisms.
  • DCBX interoperability if used.
  • Per-queue counters and congestion telemetry.

However, treat PFC carefully. It is normally preferable to limit lossless behavior to the domain that genuinely requires it. Extending PFC indiscriminately across firewalls, routers, WAN links, or service chains can spread congestion and create operational risk.

4. Security and Observability

Because it sits at the fabric boundary, the border leaf should provide strong visibility and control.

Look for:

  • Streaming telemetry, such as gNMI or model-driven telemetry.
  • sFlow, IPFIX/NetFlow, or equivalent flow export.
  • ERSPAN, local mirroring, and packet-broker integration.
  • Syslog, SNMP, event streaming, and alerting hooks.
  • AAA integration with TACACS+ or RADIUS.
  • Role-based access control.
  • MACsec, IPsec, or support for encrypted uplinks where required.
  • Secure boot and signed software images.

The border leaf is often the best place to observe aggregate ingress and egress behavior, detect route anomalies, measure DCI utilization, and export traffic records to security or capacity-planning platforms.

5. Automation and Operational Consistency

For AI data center operations, manual configuration is not enough. Select a platform with an automation interface compatible with the rest of the fabric.

Prefer support for:

  • OpenConfig and native YANG models.
  • gNMI, NETCONF, RESTCONF, or well-documented APIs.
  • Ansible, Terraform, or Python automation workflows.
  • Configuration templates and Git-based change control.
  • Zero-touch provisioning.
  • Software image management and rollback.
  • Integration with SONiC or the organization’s chosen network operating model, where applicable.

The operational objective is to make a border leaf behave like a standardized fabric role—not a one-off, manually maintained edge device.

Border Leaf Deployment Models

In a data center Fabric, the Border Leaf does not have a single fixed form. Whether to deploy dedicated Border Leaf switches is essentially a trade-off between architectural separation, cost, and operational complexity. The common deployment models can be divided into two main types:

Dedicated Border Leaf

In the Dedicated model, the Border Leaf serves as a dedicated edge gateway. It is responsible only for connecting the Fabric to external networks, such as the WAN, external firewalls, the Internet, or the enterprise core network.

border leaf switch architecture

Advantages:

Separation of Concerns: East-West compute traffic within the data center is decoupled from North-South traffic, reducing interference between the two traffic domains.
Scalability: When external routing tables, such as the BGP Full Table, grow or North-South bandwidth requirements increase, the Border Leaf can be upgraded or scaled independently without modifying other Leaf switches in the Fabric.
Policy Centralization: Security policies, NAT, route filtering, and other control points are centralized on dedicated nodes. This provides clear configuration boundaries and better fault isolation.

Border Leaf + Leaf Convergence

In some small and medium-sized data centers or resource-constrained environments, Border Leaf functions are often co-located on regular Compute Leaf switches to reduce device count and hardware costs. In this model, selected Leaf switches serve both compute node connectivity and external network routing functions.

Although this converged model reduces CAPEX (Capital Expenditure), it also introduces several engineering challenges:

  • Role Convergence: The control-plane workload on these switches increases significantly. They must run internal Fabric overlay protocols, such as BGP EVPN/VXLAN, while also maintaining routing adjacencies with external networks.
  • Resource Contention:
    • Table Resource Contention: Hardware TCAM resources are limited. A converged node must store both internal host routes and external network routes, including potentially large numbers of external prefixes. This can quickly approach hardware table capacity.
    • Bandwidth and Performance Contention: East-West and North-South traffic share the same switching ASIC and uplinks. This can increase the risk of microbursts and congestion.
  • Operational Complexity: Changes to external routing policies can potentially affect server workloads connected to the same Leaf due to configuration errors. During troubleshooting, it can also be more difficult to determine whether an issue originates from the internal Fabric or the external network.

Whether to deploy Dedicated Border Leaf switches depends entirely on the actual network scale and business requirements.

  • Medium to Large / High-Reliability Data Centers: Dedicated Border Leaf switches are strongly recommended. Architectural separation provides greater resilience and stability.
  • Small / Edge Data Centers: The Convergence model can be used to control hardware costs. However, the hardware TCAM capacity should be carefully evaluated, and North-South traffic should be properly rate-limited and isolated using QoS.

Design Recommendations

Use these guidelines when positioning border leaf switches in an AI data center.

  1. Deploy border leafs in redundant pairs at minimum.
  2. Connect each border leaf to every spine in its fabric or pod whenever the topology and port budget allow.
  3. Use ECMP across the border leaf pair and across all available spine paths.
  4. Separate AI back-end, front-end, storage, management, and external-service traffic using distinct fabrics, VRFs, or clearly defined policy domains.
  5. Keep latency-sensitive GPU collective traffic off firewall, NAT, and load-balancer service chains whenever possible.
  6. Size border capacity for normal traffic and single-failure conditions, not only for aggregate interface speed.
  7. Use explicit BGP route policies to prevent accidental route leaks between tenant VRFs, AI fabrics, enterprise networks, and WAN domains.
  8. Validate interoperability among the border leaf, DCI router, firewall, load balancer, and external routing domain before production rollout.
  9. Monitor queue occupancy, ECN/PFC events, BGP state, route-table growth, packet drops, and per-link utilization continuously.
  10. Treat the border leaf as a dedicated fabric role, even if it uses the same hardware SKU and network operating system as ordinary leaf switches.

Asterfusion’s Border Leaf Switches for AI Data Centers

1. Native Border Leaf Support Across the Hardware and Software Portfolio

Asterfusion’s data center switch portfolio, covering speeds from 100G to 800G, provides the hardware capabilities and AsterNOS software features required for Border Leaf deployments. These include large routing table capacity, external routing protocols such as BGP/EVPN, and security policies.

Architects can select the appropriate switch model based on the actual network design and deploy it directly as a Dedicated Border Leaf.

edge-data-center-10g-800g
All Data Center Switches And Edge Router Platforms

2. Full Border Leaf Architecture for PD Disaggregation

    For Prefill-Decode Disaggregation (PD disaggregation) in LLM inference, Asterfusion has designed a Full Border Leaf architecture. In a PD-disaggregated architecture, Prefill and Decode nodes have different traffic patterns from traditional compute nodes. They frequently exchange data with external schedulers and across clusters. Enabling Border Leaf capabilities on all Leaf switches within the cluster can reduce multi-tier forwarding.

    • Reduce CAPEX by approximately 30%: Compared with traditional multi-tier Clos architectures, the Full Border Leaf design reduces the number of devices and ports required, lowering overall network hardware costs by approximately 30%.
    • More details coming soon.

    Conclusion

    A border leaf is the controlled boundary between an internal Clos fabric and the external networks and services it depends on. It connects the fabric to WAN, DCI, Internet edge, security appliances, load balancers, storage services, and management systems while centralizing routing, segmentation, observability, and policy enforcement.

    For an AI data center, the value of a border leaf goes beyond simple north-south connectivity. It protects the performance and predictability of the GPU fabric by separating external, service-chained, and potentially bursty traffic from critical east-west AI communication.

    When selecting border leaf switches, prioritize high-speed fabric uplinks, flexible interface options, sufficient forwarding and routing scale, resilient hardware, robust BGP/EVPN capabilities, QoS and RoCE-aware features where required, detailed telemetry, and automation support. The best border leaf platform is one that scales external connectivity without becoming a performance, availability, or operational bottleneck for the AI fabric.

    Not sure which switch is right for your border network? Submit a request and let our experts help you.

    Request a demo or need assistance ?

    Fill out the form, and we’ll reach out to you today !

    Latest Posts