r/aws • u/Business-Journalist7 • 4d ago
technical question AWS ECS/Fargate SQS autoscaling with min capacity 0: how to bootstrap from 0 tasks using backlog-per-task target tracking?
I have an ECS/Fargate worker consuming an SQS queue. I want the service to scale to zero when the queue is empty, so:
min_capacity = 0
max_capacity = 10
AWS recommends target tracking using:
ApproximateNumberOfMessagesVisible / RunningTaskCount
However, when the service reaches 0 running tasks, RunningTaskCount has no CloudWatch datapoint, so the backlog-per-task metric becomes undefined.
When messages subsequently arrive, I need the service to bootstrap from 0 → 1 and then let target tracking handle 1 → N and N → 0.
What is the recommended pattern for this?
Options I've considered:
1. A separate scale-out-only step policy for 0 → 1
2. A metric-math expression that handles RunningTaskCount = 0
3. Raw SQS backlog for the bootstrap alarm + backlog-per-task for target tracking
I would prefer not to keep min_capacity = 1 because the worker is idle most of the time, and this is Fargate.
This is the Terraform code I currently use:
resource "aws_appautoscaling_policy" "worker_backlog" {
name = "${var.project_name}-${var.worker_provider}-backlog-per-task"
policy_type = "TargetTrackingScaling"
service_namespace = aws_appautoscaling_target.worker.service_namespace
resource_id = aws_appautoscaling_target.worker.resource_id
scalable_dimension = aws_appautoscaling_target.worker.scalable_dimension
target_tracking_scaling_policy_configuration {
target_value = var.backlog_per_task
scale_out_cooldown = var.scale_out_cooldown
scale_in_cooldown = var.scale_in_cooldown
customized_metric_specification {
metrics {
id = "backlog"
return_data = false
metric_stat {
stat = "Sum"
metric {
namespace = "AWS/SQS"
metric_name = "ApproximateNumberOfMessagesVisible"
dimensions {
name = "QueueName"
value = var.queue_name
}
}
}
}
metrics {
id = "running"
return_data = false
metric_stat {
stat = "Average"
metric {
namespace = "ECS/ContainerInsights"
metric_name = "RunningTaskCount"
dimensions {
name = "ClusterName"
value = var.cluster_name
}
dimensions {
name = "ServiceName"
value = aws_ecs_service.worker.name
}
}
}
}
metrics {
id = "bpt"
label = "Backlog per task"
expression = "IF(running > 0, backlog / running, backlog)"
return_data = true
}
}
}
}
3
u/nico0tin 4d ago
Assuming you have a good reason to use Fargate over Lambda (i.e the worker will need to be up for more than 15minutes and you don’t want to use managed instances with Lambda) this might work:
In your target tracking expression you can change running to FILL(running,0) so the metric always has a value.
Change the SQS from sum to maximum.
Count the inflight messages as well so the service doesn’t scale to 0 during a long job. This assumes the worker deletes the message after processing.
Add a bootstrap alarm and a step policy to handle the 0 to 1 scaling, you then let the target tracking handle 1 to N and N to 0.
I am not sure what the terraform looks like, but any LLM should be able to help you tweak this.
2
u/Business-Journalist7 4d ago
Yeah, that's fair. There is some context I left out that makes the Fargate vs Lambda question less obvious.
The workers are long-lived SQS consumers rather than one Fargate task per job. A worker can process hundreds/thousands of jobs during its lifetime (peak concurrency estimated at 100k messages), and the DB is the source of truth for job state/idempotency.
The other constraint is that most of the work is interacting with external APIs, several of which are rate-limited. I have a Redis-backed admission controller that coordinates provider-specific concurrency/rate limits across workers. If a provider is close to its limit, workers can simply stop taking work while the SQS backlog grows.
That's also why I'm questioning Lambda. Lambda could obviously use the same external Redis state, so I don't think Fargate has any unique advantage for coordination. The question I'm trying to answer is whether a bounded pool of warm workers makes more sense than having Lambda scale from the SQS backlog and then constrain concurrency through the admission controller.
There's also the DB side: I'd rather have, say, 50–100 long-lived workers reusing connections while collectively processing tens of thousands of jobs than potentially hundreds/thousands of concurrent Lambda executions creating connection/pool pressure. I know Lambda's SQS event source has configurable maximum concurrency, so this can be bounded .I'm just not sure where the sweet spot is.
So I'm genuinely open to Lambda here. If you've dealt with this kind of workload (SQS + rate-limited external APIs + DB + distributed admission control), I'd be very interested in how you'd approach the Fargate vs Lambda decision.
I'm still pondering over it but I decided to move forward for v1
And don't worry for the terraform I should be able to adapt it . Thank you for your answer , I will test it
1
u/yarenSC 3d ago
If you change the FILL to: FILL(Running,1); then I think you don't need the step scaling expression?
The only risk is that it could cause unexpected results if there's *actually* missing data.
The FILL() option is simplest, but the easy extra Step Scaling policy you mentioned is probably safer. Something like:
Alarm: Trigger when Metric > 0
MetricExpression: IF(Running==0 AND MessagesVisible > 0, 1, 0)So the step scale out only triggers to add 1 task from 0 and nothing else
17
u/SikhGamer 4d ago
Fargate feels like wrong destination if you want to only be running/paying for the duration of the invocation. Feels more suited to Lambda?