I'm relatively new to GPU networking and need guidance on a 16-node Kubernetes cluster. Each node has eight NVIDIA RTX GPUs without NVLink, and the compute network uses Spectrum-based Ethernet switches. The cluster uses vanilla Kubernetes, Multus, Calico, and SR-IOV networking.
The setup has multiple rails, with a separate /24 subnet allocated to each rail. The SR-IOV networks, IP pools, and other required plumbing are already configured, and traffic within the same rail works correctly.
The problem is cross-rail communication for distributed GPU workloads. Each pod currently has the Kubernetes network as its default route, but it also needs routes for the other rail subnets through the appropriate compute-network switch or gateway. I tried configuring additional routes through IPAM, but the routes did not appear inside the pod network namespace. What is the correct way to install and manage these cross-rail routes?
2 Answers
First verify the expected topology: a route inside the pod can only forward packets to a next-hop gateway that is reachable through one of the pod’s attached interfaces. If each rail is a separate /24, configure a route for the remote rail supernet on the relevant SR-IOV attachment, with the compute switch’s rail-facing address as the gateway. Depending on the SR-IOV CNI and IPAM plugin, route fields may be ignored unless they are placed in the NetworkAttachmentDefinition or supported by the specific IPAM implementation. Check the pod’s network namespace directly with `ip addr`, `ip route`, and `ip rule`; if the route is absent, inspect the Multus and SR-IOV CNI logs and confirm that the selected IPAM plugin supports custom routes.
If the additional SR-IOV interfaces need to carry traffic based on their source addresses, look at Multus source-based routing rather than trying to add a second default route. The `sbr` meta-plugin can create policy-routing rules so traffic originating from a rail IP uses that rail’s interface and gateway, while the normal Kubernetes interface keeps the pod’s default route. The SR-IOV network definition also needs to provide the correct gateway and route information, and the gateway must have reachability to the other rail subnets.

I’m using one Calico interface and one SR-IOV interface per pod, with a single default route through the Kubernetes network. The missing piece is adding routes for the other rail networks through the appropriate compute-switch gateway, rather than creating another default route.