I run an ECS/Fargate worker that consumes messages from an SQS queue. I want the service to scale down to zero when the queue is empty, with a minimum capacity of 0 and a maximum of 10 tasks. For normal scaling, I'm using a target-tracking metric based on ApproximateNumberOfMessagesVisible divided by RunningTaskCount. The problem is that RunningTaskCount has no datapoint when the service has zero tasks, so backlog-per-task becomes undefined. When messages arrive, I need the service to wake from 0 to 1 task, then let target tracking manage scaling from 1 to N and eventually back to 0. What is the recommended AWS pattern? Should I use a separate scale-out alarm, metric math that handles zero tasks, or raw queue depth for bootstrapping and backlog-per-task afterward? I would prefer not to keep one task running because the worker is idle most of the time.
4 Answers
Avoid a polling Lambda that constantly sets desired count unless you really need custom behavior. It effectively becomes a second autoscaler and can conflict with Application Auto Scaling during cooldowns or scale-in. A raw SQS backlog alarm for bootstrapping plus target tracking after a task exists is simpler and keeps the responsibilities clear. Keeping minimum capacity at zero is reasonable when idle Fargate cost matters, though the tradeoff is cold-start latency.
Use a separate CloudWatch alarm on SQS queue depth to wake the service from zero to one task. Target tracking cannot reliably make that transition because its denominator, RunningTaskCount, is missing when no tasks exist. Once the first task is running, backlog-per-task target tracking can handle scaling from one task upward and scaling back down. Treat the queue-depth alarm as a bootstrap mechanism, not as a competing autoscaling policy. Add a cooldown and keep a maximum task limit so a sudden or stuck backlog cannot cause uncontrolled fan-out.
You can also make the metric math explicit about the zero-task state with FILL or an IF expression. Conceptually, zero running tasks with a nonzero backlog should represent unbounded demand, which should trigger a wake-up. However, relying on metric math alone can be less predictable than having a simple queue-depth alarm for 0-to-1. Be careful with the false branch in the expression: returning the raw backlog can make a small queue spike request many tasks if the target value is low.
Make sure the metric represents unfinished work consistently. If messages are deleted before processing actually completes, visible messages alone may understate the workload; in-flight messages or an application-level work metric may be more accurate. Also use cooldowns and a sensible backlog-per-task target so intermittent traffic does not cause rapid scale-out and scale-in flapping.

Related Questions
Can't Load PhpMyadmin On After Server Update
Redirect www to non-www in Apache Conf
How To Check If Your SSL Cert Is SHA 1
Windows TrackPad Gestures