I run a bare-metal Kubernetes cluster with Cilium integrated with BGP and Cisco Nexus switches in an EVPN fabric. The cluster uses externalTrafficPolicy: Local for performance, and I assigned five worker nodes as dedicated load-balancer nodes that advertise service IPs externally.
The setup works until a load-balancer node fails. After some time, existing connections through the other load-balancer nodes also drop. Connections also reset when the failed node is repaired and added back into the fabric. My understanding is that the Nexus switch recalculates its ECMP next-hop hash whenever a BGP path is added or removed. Existing flows can then be sent to a different load-balancer node, which does not have the necessary connection state.
I tried enabling resilient ECMP with `hardware profile ecmp resilient`, but the issue remains. I understand the switch hashes Layer 3/4 packets without knowing the application connection state, but I expected existing flows to survive path changes.
Some suggested using Cilium Maglev with DSR, although I am concerned about asynchronous routing. Another possibility would be placing two or three Linux routers running FRR between the Nexus switches and the Kubernetes load-balancer nodes. Has anyone designed this type of architecture, and what is the recommended way to prevent established connections from being disrupted when load-balancer nodes are added or removed?
2 Answers
This is an ECMP rehashing problem. When a BGP next hop disappears or is added, the switch recalculates its flow distribution, and some established flows can move to a node that has no matching connection state. Resilient ECMP can reduce movement during individual failures, but it does not guarantee that flows will remain on the same node when the ECMP member set changes.
The usual choices are to use a dataplane with consistent service hashing, such as Cilium Maglev, or move the stateful load-balancing decision to a component that all traffic reaches as a stable destination. Simply inserting FRR routers may change the topology and next-hop behavior, but it will not inherently preserve connection state or eliminate rehashing if the final ECMP set still changes. Any design using intermediate routers would need a stable VIP or anycast service, consistent hashing, and a clear plan for failure and recovery behavior.
With dedicated load-balancer nodes, `externalTrafficPolicy: Local` may not provide the benefit you expect because traffic arriving at a load-balancer node can only be forwarded to backend pods local to that node. Consider `externalTrafficPolicy: Cluster` unless preserving the client source address and avoiding an extra hop are essential.
For Cilium, Maglev plus DSR is a common way to make backend selection more stable. Maglev gives nodes the same consistent-hash table, so if ECMP sends packets from an established flow to another node, that node can select the same backend. DSR preserves the original client IP and port instead of creating a different SNAT tuple on each load-balancer node, allowing the backend to recognize the existing connection.
DSR also reduces the performance concern behind `Local`: only the request takes the additional load-balancer path, while the response can return directly from the backend node to the client. It also preserves the original client address.
My concern is that adding another physical load-balancer node still changes the Nexus ECMP next-hop set. Would Maglev and DSR prevent the resulting connection resets, or would the switch still move established flows to different nodes? I am also considering two or three FRR-based Linux routers between the Nexus fabric and the Kubernetes nodes, but I am not sure whether that would actually solve the hashing issue.

The network uses EVPN, with anycast gateway addresses on the leaf switches, and the Kubernetes nodes peer with the leaves using BGP. I am trying to understand what a practical higher-layer load-balancing design would look like in this topology, especially if I want to avoid relying on asymmetric routing.