In a single Node.js process, a circuit breaker can track failures with a ring buffer, threshold, and timestamps. But when the same service runs on twenty replicas, each instance maintains its own state, meaning a failing downstream dependency may receive roughly twenty times the traffic before any breaker opens. Moving the failure window and breaker state into Redis can coordinate the replicas, but it also challenges assumptions made by the in-process design. What are the important issues to handle when building a distributed circuit breaker, and how should they be addressed?
1 Answer
Client-side exponential backoff and per-user request gating can reduce duplicate user actions, but they solve a different problem. A distributed circuit breaker protects a shared downstream dependency—such as a database, queue, or gRPC service—when many application instances are calling it at once. If twenty replicas independently decide that the dependency is healthy enough to probe, client or edge-level controls cannot coordinate those decisions. Shared breaker state lets the instances agree on when to stop sending traffic and limits the recovery probes across the whole fleet.

That distinction makes sense: client-side controls help suppress duplicate requests from one user, while the distributed breaker coordinates access to the dependency across all service instances.