I run an ECS/Fargate worker that consumes messages from an SQS queue. I want the service to scale down to zero when the queue is empty, with a minimum capacity of 0 and a maximum of 10 tasks. For normal operation, I'm using target tracking based on backlog per task: ApproximateNumberOfMessagesVisible divided by RunningTaskCount.
The problem is that RunningTaskCount has no datapoint when the service has zero tasks, so the backlog-per-task metric becomes undefined. When messages arrive, I need the service to wake up from 0 to 1 task, then allow target tracking to scale between 1 and 10 tasks and eventually back down to zero.
Would the best approach be a separate scale-out-only CloudWatch alarm for the 0-to-1 transition, metric math that treats a missing or zero task count explicitly, or a raw SQS backlog alarm for bootstrapping combined with backlog-per-task target tracking afterward? I'd prefer not to keep one task running because the worker is idle most of the time and it runs on Fargate.
4 Answers
You can also make the metric math handle the missing task count by filling RunningTaskCount with zero. That makes the zero-task state explicit instead of leaving the expression undefined. However, dividing backlog by zero is effectively an indication of unbounded demand, so you still need to verify how Application Auto Scaling interprets the resulting expression and protect the service with its maximum capacity.
Be careful with an expression such as IF(running > 0, backlog / running, backlog). If the fallback is just the raw backlog, even a small queue spike could look like a large per-task target depending on your configured target value. Test the scaling behavior with realistic bursts and cooldowns.
A separate queue-depth alarm is generally clearer than using a Lambda that periodically changes desired count. The alarm acts only as a wake-up mechanism, while target tracking remains responsible for steady-state capacity. A polling Lambda can work, but then it becomes a second autoscaling controller and can introduce races, delayed reactions, and extra operational code.
Keeping min capacity at zero is reasonable when idle-cost savings matter, but expect cold-start latency while Fargate launches the first task. If messages need very fast processing, keeping one warm task may be simpler and more predictable. Also make sure scale-in is conservative enough that a task is not removed while it is still processing or while messages are arriving in short bursts.
The usual pattern is to make the 0-to-1 transition explicit. Add a CloudWatch alarm on SQS queue depth that scales the service from zero to one task. Once that task is running, let the backlog-per-task target-tracking policy handle scaling from 1 to N and scaling back down. Keep the bootstrap alarm scale-out-only so it does not compete with target tracking, and add a cooldown to avoid repeated wake-ups during intermittent traffic.

Related Questions
Can't Load PhpMyadmin On After Server Update
Redirect www to non-www in Apache Conf
How To Check If Your SSL Cert Is SHA 1
Windows TrackPad Gestures