How can I prevent connection drops when scaling Cilium BGP load-balancer nodes behind Cisco Nexus ECMP?

0
0
Asked By MellowPine47 On

I run a bare-metal Kubernetes cluster using Cilium as the CNI, with service announcements handled through BGP peering to Cisco Nexus switches. For performance and source-IP preservation, my services use externalTrafficPolicy: Local. I currently have five dedicated worker nodes that advertise load-balancer service IPs to the outside network.

The setup works normally until one of the load-balancer nodes fails. After several hours, I observed existing connections dropping across the remaining nodes. When I repaired and re-added the failed node, connections dropped again. My understanding is that the Nexus switch recalculates its ECMP next-hop hashing whenever a BGP path is added or removed. Existing flows can then be sent to a different load-balancer node, which does not have the connection state for that flow.

I tried enabling the Nexus resilient ECMP profile, but it did not prevent the disruption. I understand that the switch hashes at the network-flow level rather than tracking application connections, but I expected established flows to continue working. Would Cilium Maglev with DSR solve this when nodes are added or removed? I am concerned about asymmetric routing with DSR. Another idea is to place two or three Linux routers running FRR between the Nexus fabric and the Kubernetes load-balancer nodes. Could that architecture prevent the ECMP rehash from breaking existing connections?

2 Answers

Answered By PacketTrail62 On

This is an ECMP rehashing problem. When a BGP next hop disappears or is added, the switch recalculates the hash and some established flows can move to a different node. That node does not have the original connection state, so the flow may be reset.

Resilient ECMP can reduce movement during a single path failure, but it generally cannot guarantee that flows will remain on the same path when the ECMP member set itself changes. FRR in front of the Kubernetes nodes would not automatically solve that; it would simply introduce another routing layer, and ECMP changes there could cause the same kind of flow movement.

To move the decision higher in the stack, use a stateful or connection-aware front end that presents one stable destination toward the fabric. Examples include a redundant pair of load balancers using a floating or anycast service address, or a dedicated proxy/load-balancer tier that owns the connection state and distributes traffic to Kubernetes. The network should then ECMP only toward that stable tier, rather than directly across independent nodes whose connection state is not shared.

CobaltMeadow3 -

The network fabric uses EVPN, with anycast gateway addresses on the leaf switches, and the Kubernetes nodes peer directly with the leaves over BGP. What would a practical design look like for placing a connection-aware load-balancing tier above the Kubernetes nodes in that topology?

Answered By QuietHarbor8 On

With dedicated load-balancer nodes, externalTrafficPolicy: Cluster is usually a better fit. With Local, a node can forward only to backend pods running on that same node, which largely defeats the purpose of having a dedicated pool of load-balancer nodes.

The usual Cilium approach for this situation is Maglev together with DSR. Maglev gives nodes the same consistent-hash table, so if ECMP sends a flow to another node, that node can still select the same backend. DSR preserves the original client and destination flow information instead of creating a node-specific SNAT tuple. That means the backend can continue recognizing the connection even if the ingress node changes.

DSR also avoids sending the response back through the load-balancer node. Only the request takes the extra hop, while the backend replies directly to the client, so the performance cost is often small. You also retain the real client source address.

MellowPine47 -

My concern is what happens when I add another dedicated node for capacity. The Nexus switch will still change its ECMP path set and may rehash existing flows. Would Maglev and DSR preserve those connections in that case? I am also considering FRR routers between the Nexus switches and the Kubernetes nodes because I would prefer to avoid DSR if possible.

Related Questions

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.