I run a long-lived ECS/Fargate worker service that consumes an SQS queue. I want the service to scale between 0 and 10 tasks, with no tasks running while the queue is empty. For normal target tracking, I'm using ApproximateNumberOfMessagesVisible divided by RunningTaskCount as the backlog-per-task metric.
The problem is that RunningTaskCount has no useful datapoint when the service is already at zero, so the target-tracking metric becomes undefined. When messages arrive, I need a reliable 0-to-1 bootstrap, after which target tracking can manage scaling from 1 to N and eventually back to 0.
Would the recommended design be a separate scale-out alarm and step policy based on raw queue depth, while retaining backlog-per-task target tracking for the running service? Alternatively, can metric math safely handle the zero-task case with FILL or an IF expression?
I'd prefer not to keep min_capacity at 1 because the workers are idle much of the time and run on Fargate. The workers are long-lived consumers that may process many jobs, coordinate external API rate limits through Redis, and use a database as the source of truth. I'm also considering whether Lambda with bounded SQS concurrency would be a better fit, but for now I'm testing the Fargate approach.
2 Answers
You can use FILL in metric math, but be careful about what value you fill with. Filling RunningTaskCount with 1 makes the expression simple, but it can hide genuinely missing telemetry and produce misleading scaling decisions. Filling it with 0 and using an explicit condition is safer for the zero-capacity case.
A dedicated 0-to-1 step policy is usually easier to reason about: trigger it only when queue depth is positive and the service has no running tasks. Then leave all other scaling decisions to target tracking. This separates bootstrapping from proportional scaling instead of trying to make one target-tracking metric represent both situations.
Use two scaling mechanisms. Keep the backlog-per-task target-tracking policy for scaling once at least one task is running, but add a separate bootstrap alarm that watches raw SQS backlog and scales the service from 0 to 1. The alarm can trigger when visible messages are greater than zero, with a step policy that adds exactly one task. After that, target tracking can handle 1-to-N scaling and eventually scale back down.
For the target-tracking expression, make the zero-task behavior explicit, for example by filling missing RunningTaskCount data with zero and guarding the division with an IF expression. Including in-flight messages in the backlog can also prevent the service from scaling in while a task is still processing work, assuming messages are deleted only after successful processing.

That matches what I was looking for. I’ll keep the separate bootstrap policy and adapt the metric math in Terraform. Including in-flight messages should also better represent work that is still being processed.