MLOps & Cloud

SageMaker async inference: the scale-to-zero trap

SageMaker async inference: the scale-to-zero trap

[!NOTE] TL;DR SageMaker async inference can reach 0 instances, but target tracking alone can leave small backlogs waiting. Pair backlog scaling with a wake-up alarm, then persist completion through SNS, SQS, and DynamoDB. Count billable instance-hours, including idle capacity; the calculations below are not production benchmarks.

The idle endpoint

The idle endpoint still costs money. In the AWS hosting cost analysis, a SageMaker async inference endpoint runs for 24 hours at 0.23 USD/hour, costing 5.52 USD/day. The model rests; the invoice doesn't.

That example dates from May 30, 2023. It isn't a current regional quote. I use its public billing evidence to examine a specific problem: removing idle capacity without leaving accepted requests unprocessed.

I treat the instance charge like a workshop electricity meter. Stopping the machine doesn't disconnect the meter; reducing provisioned capacity changes what it records.

The InvokeEndpointAsync API reference makes the acceptance boundary explicit:

Boundary Documented value Meaning Design consequence
Successful submission HTTP 202 Request accepted Not a completed prediction
Processing timeout default 900 seconds Container execution allowance Set explicitly for longer work
Processing timeout maximum 3600 seconds Maximum execution allowance Reject unsuitable jobs before submission
Queue TTL maximum 21600 seconds Maximum queue residence Reconcile expired requests

A question of waiting

My main question is how to remove idle inference capacity while keeping every accepted job recoverable. Not instantly completed. Recoverable.

I call the reference architecture ZED, for zero-endpoint dispatch. I separate 4 concerns: admission control, endpoint capacity, durable completion, and cost attribution. ZED is a proposed extension of patterns from my production work, not a deployed SageMaker benchmark.

I need 2 clocks on the workbench. One measures queue residence; the other measures execution, because the 21600-second queue allowance doesn't replace the 3600-second processing ceiling.

A reader-facing deadline is a separate application requirement. Set it deliberately. Neither service maximum promises that your job will finish within that time.

The sleeping autoscaler

One spontaneously imagines that setting MinCapacity=0 completes the optimization. AWS's asynchronous endpoint autoscaling guide describes the opposite trap: a small backlog can wait below the target-tracking threshold.

Its example target is 5 requests per instance. The additional HasBacklogWithoutCapacity policy handles the zero-capacity case. This metric equals 1 when requests exist but no instances serve the endpoint; otherwise, it equals 0.

The policies have different jobs. Target tracking adjusts capacity using ApproximateBacklogSizePerInstance. Step scaling requests an additional instance when the wake-up alarm fires. Neither eliminates provisioning latency.

A thermostat regulates an occupied workshop. It doesn't unlock the door after everyone leaves; the wake-up policy provides that separate entry mechanism.

Although target tracking remains useful under sustained traffic, I wouldn't use it alone for sporadic arrivals. Keep both policies. And don't copy a real-time InvocationsPerInstance dashboard: AWS explicitly excludes that metric from asynchronous endpoints.

ZED target architecture

My document-processing architecture uses 4 Lambda stages and 2 parallel Map states, each with a concurrency ceiling of 20. Intermediate artifacts live in S3; retries operate at the failed stage. Those are documented production choices, not throughput measurements.

I retain the separation of artifacts and orchestration for ZED. I don't transplant the concurrency setting. Inference capacity needs its own load test.

flowchart LR
    A[Validated job and S3 input] --> B[Step Functions admission]
    B --> C[Submission Lambda]
    C --> D[(DynamoDB job ledger)]
    C --> E[SageMaker internal request queue]
    E --> F[Async endpoint: 0 to 5 instances]
    F --> G[(S3 output)]
    F --> H[SNS success or error]
    H --> I[SQS completion queue]
    I --> J[Completion Lambda]
    J --> D
    I --> K[Dead-letter queue]
    L[CloudWatch backlog alarm] --> M[Application Auto Scaling]
    M --> F
    N[Scheduled reconciliation] --> D
    N --> G

SageMaker's internal request queue holds inference work. The SQS queue above holds completion notifications. They aren't interchangeable.

I use S3 as the artifact archive. DynamoDB holds its index cards, containing job state and object locations rather than the prediction payload itself.

Before invocation, the submission Lambda writes a job record with a caller-generated InferenceId. This establishes correlation before a fast completion arrives. Its 4 states are REGISTERED, SUBMITTED, COMPLETED, and FAILED.

Conditional updates must stop a late submission response from overwriting COMPLETED. The completion consumer records the terminal result before acknowledging its SQS message. A duplicate notification should leave that record unchanged.

Step Functions finishes admission after submission. It doesn't wait inside Lambda for inference. I deliberately keep admission success separate from prediction success, since HTTP 202 confirms only acceptance.

The executable contract and wake-up policy

I built the local ZED contract below for Python 3.12 and Pydantic 2.13.4. It uses Pydantic validation to reject out-of-range timeouts and expose compute-only arithmetic. It makes 0 AWS calls.

from urllib.parse import urlsplit
from pydantic import BaseModel, ConfigDict, Field, computed_field, field_validator
class AsyncRequest(BaseModel):
    model_config = ConfigDict(strict=True, extra="forbid")
    endpoint_name: str = Field(min_length=1, max_length=63, alias="EndpointName", description="Existing asynchronous endpoint name.")
    input_location: str = Field(max_length=1024, alias="InputLocation", description="S3 URI of the immutable inference payload.")
    inference_id: str = Field(min_length=1, max_length=64, pattern=r"^[A-Za-z0-9-]+$", alias="InferenceId", description="Caller-generated correlation identifier.")
    content_type: str = Field(default="application/json", alias="ContentType", description="Content type accepted by the model container.")
    request_ttl_seconds: int = Field(default=21600, ge=60, le=21600, alias="RequestTTLSeconds", description="Maximum queue residence in seconds.")
    invocation_timeout_seconds: int = Field(default=900, ge=1, le=3600, alias="InvocationTimeoutSeconds", description="Maximum processing duration in seconds.")
    @field_validator("input_location")
    @classmethod
    def check_s3_object(cls, value: str) -> str:
        location = urlsplit(value)
        if location.scheme == "s3" and location.netloc and len(location.path) > 1:
            return value
        raise ValueError("An S3 bucket and object key are required")
class ComputeEstimate(BaseModel):
    model_config = ConfigDict(strict=True, extra="forbid", allow_inf_nan=False)
    hourly_rate_usd: float = Field(gt=0.0, description="Historical example rate, not a current quote.")
    instance_hours: float = Field(ge=0.0, description="Aggregate billable instance-hours across the fleet.")
    @computed_field(description="Compute-only charge before discounts and other services.")
    @property
    def compute_usd(self) -> float:
        return self.hourly_rate_usd * self.instance_hours
# Emit API keyword arguments; the bucket and endpoint are illustrative.
request = AsyncRequest(EndpointName="zed-async", InputLocation="s3://example-input/jobs/job-001.json", InferenceId="job-001")
print(request.model_dump_json(by_alias=True))
# Evaluate assumptions without claiming a measured production result.
for hours in (24.0, 2.0, 2.5, 4.0):
    estimate = ComputeEstimate(hourly_rate_usd=0.23, instance_hours=hours)
    print(estimate.model_dump_json())

Pass request.model_dump(by_alias=True) as keyword arguments to the Boto3 client's invoke_endpoint_async method. Persist its returned output location. This contract validates syntax and numeric limits; it doesn't authorize bucket access or check model compatibility.

Think of the validator as a caliper at admission. It rejects a dimension outside the 3600-second ceiling, but it cannot prove that the submitted work fits inside that dimension.

The following CloudFormation template attaches scaling to an existing, in-service asynchronous endpoint. Supply its variant name and Application Auto Scaling service-linked role ARN. This isn't an endpoint deployment template.

AWSTemplateFormatVersion: '2010-09-09'
Description: ZED scaling for an existing asynchronous endpoint
Parameters:
  EndpointName:
    Type: String
  VariantName:
    Type: String
  ScalingRoleArn:
    Type: String
Resources:
  Target:
    Type: AWS::ApplicationAutoScaling::ScalableTarget
    Properties:
      MinCapacity: 0
      MaxCapacity: 5
      ResourceId:
        Fn::Sub: endpoint/${EndpointName}/variant/${VariantName}
      RoleARN:
        Ref: ScalingRoleArn
      ScalableDimension: sagemaker:variant:DesiredInstanceCount
      ServiceNamespace: sagemaker
  BacklogPolicy:
    Type: AWS::ApplicationAutoScaling::ScalingPolicy
    Properties:
      PolicyName: zed-backlog
      PolicyType: TargetTrackingScaling
      ScalingTargetId:
        Ref: Target
      TargetTrackingScalingPolicyConfiguration:
        TargetValue: 5.0
        CustomizedMetricSpecification:
          MetricName: ApproximateBacklogSizePerInstance
          Namespace: AWS/SageMaker
          Statistic: Average
          Dimensions:
            - Name: EndpointName
              Value:
                Ref: EndpointName
  WakePolicy:
    Type: AWS::ApplicationAutoScaling::ScalingPolicy
    Properties:
      PolicyName: zed-wake
      PolicyType: StepScaling
      ScalingTargetId:
        Ref: Target
      StepScalingPolicyConfiguration:
        AdjustmentType: ChangeInCapacity
        MetricAggregationType: Average
        Cooldown: 300
        StepAdjustments:
          - MetricIntervalLowerBound: 0
            ScalingAdjustment: 1
  WakeAlarm:
    Type: AWS::CloudWatch::Alarm
    Properties:
      MetricName: HasBacklogWithoutCapacity
      Namespace: AWS/SageMaker
      Statistic: Average
      Period: 60
      EvaluationPeriods: 2
      DatapointsToAlarm: 2
      Threshold: 1
      ComparisonOperator: GreaterThanOrEqualToThreshold
      TreatMissingData: missing
      Dimensions:
        - Name: EndpointName
          Value:
            Ref: EndpointName
      AlarmActions:
        - Ref: WakePolicy

The 60-second periods and 2 breaching datapoints follow AWS's example. They don't promise a 120-second cold start. Metric delivery and instance provisioning add delay.

My production infrastructure uses Terraform for orchestration resources. You can express these same scaling resources there, but choose 1 owner for them. Don't let Terraform and CloudFormation manage the same alarm.

Billing evidence and computed results

The public AWS bill example establishes 5.52 USD/day for 24 instance-hours. Everything else below is arithmetic at its historical 0.23 USD/hour rate. I haven't measured ZED's daily runtime.

Scenario Instance-hours/day Compute USD/day Evidence status
Continuously provisioned 24 5.52 AWS public billing example
Concentrated processing 2 0.46 Assumed duty cycle
Processing plus idle allowance 2.5 0.575 Assumed duty cycle
More dispersed activity 4 0.92 Assumed duty cycle
{"type":"bar","data":{"labels":["AWS example: 24 h","Assumption: 2 h","Assumption: 2.5 h","Assumption: 4 h"],"datasets":[{"label":"Compute-only USD/day at historical rate","data":[5.52,0.46,0.575,0.92],"backgroundColor":["#64748b","#3b82f6","#06b6d4","#8b5cf6"]}]} ,"options":{"responsive":true,"plugins":{"title":{"display":true,"text":"Public billing baseline versus calculated duty cycles"}},"scales":{"y":{"beginAtZero":true,"title":{"display":true,"text":"USD/day, compute only"}}}}}

The 2.5-hour scenario is 89.6% below the compute baseline. That's conditional, not demonstrated. Scattered arrivals can repeatedly provision capacity and increase billable time even when total model execution stays unchanged.

I read a billing report like a workshop inventory. Counting only the machine omits the shelves and delivery equipment; an endpoint-only estimate likewise omits supporting services.

ZED's total includes storage and orchestration, with messaging and observability charges added separately. Attribute compute using the REGION-AsyncInf:instanceType usage family from AWS's cost analysis. Compare resource-level instance-hours with actual invoice lines before calling any saving real.

For monitoring, I follow the AWS asynchronous endpoint metrics reference. TimeInBacklog and TotalProcessingTime use milliseconds; ModelLatency uses microseconds. Normalize units first. Otherwise, your dashboard acquires a doctorate in exaggeration.

Start with an alert when ExpiredRequests or NotificationFailures has a sum greater than 0 over a chosen 60-second window. That window is a design choice. Add an application-age alert against your own deadline, because container latency doesn't measure the user's entire wait.

Limits and the next production probe

The local code isn't the complete worker implementation. The template configures only scaling. ZED still needs IAM policies, encrypted storage, notification subscriptions, and tested reconciliation logic.

A load test on an empty road measures the car, not the junction. I want production-shaped arrivals at that junction, including isolated requests after capacity has reached 0.

There are 4 failure cases I would probe: duplicate completion delivery, an ambiguous submission timeout, failed SNS publication, and queue expiration. Record accepted-to-terminal latency, separating cold arrivals from warm ones. Report p50 and p95. Keep failed jobs in the denominator.

InferenceId provides correlation, not a documented exactly-once guarantee. A timeout after submission can leave acceptance uncertain. Don't blindly retry that boundary. Reconcile the ledger and output locations first, and make any replay decision explicit.

Configure both SuccessTopic and ErrorTopic in the endpoint's notification configuration. Give the SageMaker execution role permission to publish, and restrict the completion queue policy to the intended topics. Treat unknown or malformed events as errors rather than success.

The reconciler must inspect overdue nonterminal jobs without depending solely on notifications. Persist known output and failure locations when submission succeeds. Where acceptance remains ambiguous, retain that uncertainty instead of inventing a terminal result.

My existing pipeline experience supports the separation of stage retries from artifacts. It doesn't establish a SageMaker cold-start distribution. That's still missing. The regional price is also unverified; replace the historical rate before budgeting a deployment.

Beyond the zero

More generally, I treat scale-to-zero as a gate, not a guarantee. Measure the waiting outside as carefully as the compute inside, or the apparently empty system still contains unfinished work. For a ZED architecture review, contact me with your arrival profile and billing export, or browse the MLOps and Cloud articles; the invoice won't applaud.


Processing...
Processing...

Please wait

Secure operation