FREE TOOL

SageMaker Inference Cost Calculator

Enter a workload - requests a month, seconds per request, the instance and the serverless memory - and compare what it costs a month on a SageMaker real-time, serverless or asynchronous endpoint and in batch transform jobs, with the request volume where one gets cheaper than another.

  • Your data never leaves your browser: everything is calculated by JavaScript on this page, not on a server.
  • Nothing you enter is uploaded, processed on a server or stored. Check it in your browser's developer tools (Network tab).
  • Once the page has loaded, the tool works without an internet connection.

The workload

Requests are counted as spread evenly over the hours they arrive.

On an instance

Real-time, asynchronous and batch transform run on the instance type you pick.

On serverless

No instance type, no GPUs: 1 to 6 GB of memory, with vCPUs in proportion.

In batch transform

When the requests can wait and be collected in S3 for a job.

See the costs ↓
Designing cost-optimized architectures is a core AWS Solutions Architect Associate domainTry free SAA-C03 practice questions with answers and explanations.SAA-C03 questions →

How each SageMaker inference mode is billed

Real-time endpoints bill every instance by the second while the endpoint is in service, whether it gets requests or not, plus the data the endpoint processes in and out. The calculator keeps one instance all month and adds instances while requests arrive. A request may take up to 60 seconds and 25 MB.

Serverless endpoints bill the time each request runs, by the millisecond, at a price per second that grows with the memory you choose (1 to 6 GB), plus the data processed. Provisioned concurrency bills the memory it keeps warm for as long as it is kept, and the requests it serves at a lower rate. A serverless endpoint has no GPUs, takes requests of up to 4 MB and 60 seconds and at most 200 at once.

Asynchronous endpoints run on instances like a real-time endpoint, at their own price list (mostly the same prices), but queue the requests and can scale to zero instances when the queue is empty. They take requests of up to 1 GB that run up to an hour, and write the results to S3.

Batch transform bills its instances for the duration of each job. There is no endpoint and no data processing charge: the job reads its input from S3 and writes the results there. A record may be up to 100 MB, and an invocation may run up to an hour. AWS publishes no start-up time for a job, so it is an input.

The prices come from the AWS Price List API and are read again on every build of this site (published 2026-09-28). AWS states that the price on its pricing page is charged where the two differ.

Serverless inference prices in US East (N. Virginia)

MemoryOn demand, a second1M requests of 100 msProvisioned, a second keptProvisioned, a second of requests
1024 MB$0.00$2.00$0.00$0.00
2048 MB$0.00$4.00$0.00$0.00
3072 MB$0.0001$6.00$0.00$0.00
4096 MB$0.0001$8.00$0.00$0.00
5120 MB$0.0001$10.00$0.0001$0.00
6144 MB$0.0001$12.00$0.0001$0.00

Frequently asked questions

When is serverless cheaper than a real-time endpoint?

While the requests' total running time costs less than an instance kept all month. Serverless costs the same per request at any volume; a real-time endpoint costs the same per month until it needs a second instance. The break-even above is the monthly volume where the two meet for your request time, memory and instance. If a request costs less on serverless than on a fully busy instance - a fast request, a small memory size, an instance that works on one request at a time - serverless stays cheaper at any volume, up to its limit of 200 requests at once.

Why is the asynchronous endpoint as expensive as real-time?

With requests arriving around the clock, an asynchronous endpoint never empties its queue for long and keeps its instances like a real-time one. It saves only when requests stop for hours - set the hours a day they arrive - or when they are too large or too slow for a real-time endpoint.

Can a real-time endpoint scale to zero?

Only one that hosts inference components, with the endpoint configuration's minimum instance count set to 0. Until an instance has started, which takes several minutes, its invocations fail. The calculator counts a real-time endpoint as always on.

Is what I enter sent anywhere?

No. Everything runs in your browser. The page only counts that the calculator was used and which mode came out cheapest, never the instance type or the numbers.

References

Amazon SageMaker AI pricing
Inference options in Amazon SageMaker AI
Deploy models with Amazon SageMaker Serverless Inference
Autoscale an asynchronous endpoint
Scale an endpoint to zero instances
Batch transform for inference with Amazon SageMaker AI
CreateTransformJob (Amazon SageMaker API Reference)