How can an ECS service scale from zero tasks using SQS backlog per task?

0
0
Asked By MellowPine47 On

I run a long-lived ECS/Fargate worker service that consumes an SQS queue. I want the service to scale down to zero when the queue is empty, with a maximum of 10 tasks. The intended scaling metric is backlog per task, calculated as ApproximateNumberOfMessagesVisible divided by RunningTaskCount.

The problem is that RunningTaskCount has no useful datapoint when the service has zero tasks, so the target-tracking metric becomes undefined. When new messages arrive, I need the service to bootstrap from zero to one task, after which target tracking should manage scaling from one task up to the required capacity and back down to zero.

I am considering three approaches: adding a separate scale-out-only step policy for the zero-to-one transition, using metric math to handle a zero task count, or using a raw SQS backlog alarm for bootstrapping and backlog-per-task for normal target tracking. I would prefer not to keep one task running continuously because the worker is idle much of the time and runs on Fargate.

The worker is long-lived rather than one task per message. It can process many jobs during its lifetime, and a database provides job state and idempotency. Much of the work calls rate-limited external APIs, so a Redis-backed admission controller coordinates provider-specific concurrency across workers. I am also weighing Fargate against Lambda because of database connection pressure and the need to bound concurrency.

2 Answers

Answered By SilverMaple22 On

A separate scale-out-only policy is safer than trying to make target tracking infer a task when the denominator is missing. Use a metric expression that detects the specific bootstrap condition: the running task count is zero and the visible message count is greater than zero. The associated step policy should add exactly one task from zero; it should not compete with the normal target-tracking policy for other capacities.

You can also use FILL(running, 1) in the backlog-per-task expression to avoid an undefined value, but that can hide genuinely missing Container Insights data and produce surprising scaling decisions. FILL(running, 0) is more honest for the metric, while the separate bootstrap alarm handles the fact that target tracking cannot scale from an undefined metric. For the backlog itself, using an appropriate statistic and accounting for in-flight work is important.

Answered By CopperSparrow8 On

Use a separate bootstrap mechanism for the zero-to-one transition, then let target tracking handle the normal range. A practical design is an alarm based on SQS backlog that triggers a step-scaling policy to add one task when the service is at zero. Once a task exists, backlog-per-task can control scaling from one task upward and eventually back down.

For the metric math, make the task count explicitly usable when data is missing, for example with FILL(running, 0), and consider including both visible and in-flight messages if a task keeps a message invisible until processing finishes. Otherwise the queue could appear empty and the service could scale down while work is still being processed. Make sure the scale-in behavior and cooldowns are long enough for the worker's processing and shutdown behavior.

MellowPine47 -

That matches what I was looking for. The workers are long-lived and coordinate rate limits through Redis, so I will test a bootstrap alarm plus target tracking rather than forcing the minimum capacity to one. I can adapt the Terraform configuration.

Related Questions

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.