How should cross-rail GPU traffic be routed in a Kubernetes SR-IOV network?

0
1
Asked By VelvetCactus42 On

I'm relatively new to GPU networking and am working on a 16-node setup with eight NVIDIA RTX GPUs per node. The GPUs do not have NVLink. The cluster uses vanilla Kubernetes, Multus, SR-IOV, and Spectrum-based Ethernet switches for the compute network.

This is a multi-rail design, with a separate /24 assigned to each rail. The SR-IOV networks, IP pools, and related configuration are already working, and traffic within the same rail is successful. The problem is routing traffic between rails, particularly for distributed inference workloads.

I believe the pods need a route for the overall rail-network supernet, with the appropriate compute-network switch as the next hop toward the other rails or the front-end network. I tried configuring routes through IPAM, but the route did not appear inside the pod. What is the recommended way to implement cross-rail routing in this setup? Should this be handled by the switches, Multus/SR-IOV configuration, source-based routing, or another mechanism?

3 Answers

Answered By HarborMosaic7 On

This may be a limitation of the current networking design rather than a missing IPAM option. A single-node GPU networking guide will not cover distributed inference across multiple nodes and rails. For cross-rail traffic, first confirm that the Spectrum switches have Layer 3 interfaces and routes for every rail subnet, then make sure the pod's SR-IOV interface uses the intended gateway. If the switches are not routing between those networks, adding a route inside the pod alone will not solve the problem.

Answered By QuietLemon8 On

Because the pods have the Kubernetes interface plus one or more SR-IOV interfaces, routing needs to be explicit. A source-based routing setup, such as the Multus source-based routing meta-plugin, can keep traffic originating from a particular rail address on that rail instead of sending it through the default Kubernetes route. This is especially useful when the pod has multiple gateways or when different destinations must use different interfaces.

CopperNectar31 -

In this case there is currently only one default route, through the Kubernetes network. The other interface is the SR-IOV compute NIC, so the main issue is getting routes for the rail supernets installed through the correct rail gateway rather than choosing between two default routes.

Answered By MapleOrbit56 On

IPAM route settings generally only work when the CNI and meta-plugin actually apply them to the pod network namespace. If the route is absent inside the pod, inspect the generated network attachment definition and the Multus delegate configuration, then check the pod's routing table and the gateway's reachability. You may need a CNI capability that explicitly enables route injection, or a dedicated routing/meta-plugin, rather than putting the route in the IP pool alone. Also verify that return routes exist on the switches; cross-rail communication requires both forward and reverse paths.

Related Questions

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.