FREE TOOL

SageMaker Quota Finder

Pick what you are about to run on SageMaker - an endpoint, a training, processing or transform job, a HyperPod cluster, a notebook or a Studio space - and get every quota it counts against, its code, its default and the value to request, with the commands.

  • Your data never leaves your browser: everything is calculated by JavaScript on this page, not on a server.
  • Nothing you enter is uploaded, processed on a server or stored. Check it in your browser's developer tools (Network tab).
  • Once the page has loaded, the tool works without an internet connection.

What you will run

Each use has its own quota per instance type, and its own totals.

Your account's values

Optional. Without them the defaults are compared - an account can have more, or less.

Paste the output of list-service-quotas
aws service-quotas list-service-quotas --service-code sagemaker --region us-east-1 --output json
See the quotas ↓
Service quotas and limits come up in the AWS Solutions Architect Associate examTry free SAA-C03 practice questions with answers and explanations.SAA-C03 questions →

How SageMaker counts quotas

SageMaker gives every instance type a separate quota for each use, per Region: "ml.g5.xlarge for endpoint usage", "ml.g5.xlarge for training job usage", for spot training, warm pools, processing, transform jobs, notebook instances, HyperPod clusters and Studio's JupyterLab and Code Editor apps. Next to them are account-wide totals - instances across all endpoints, across all training jobs - and maxima per endpoint, job or cluster. A plan has to fit all of them, and each counts what already runs in the Region.

The defaults are low. In AWS's quota table almost every instance type starts at 0 for endpoints and on-demand training jobs, CPU instances such as ml.m5.xlarge included, and every GPU type at 0 for every use; the totals across endpoints and jobs start at 4 instances, and HyperPod's at 0. So the first endpoint or training job on a new type fails with ResourceLimitExceeded until you request an increase in the Service Quotas console or with the commands above. A quota that is not in AWS's table for a use means SageMaker does not take that instance type for it - the finder lists only the types each use has a quota for.

The defaults here come from AWS's quota table and are read again on every build of this site (last read on 2026-10-02). Your account may have other values, set by AWS or by an earlier request; paste the output of list-service-quotas to compare against them.

SageMaker's account-wide quotas

The quotas a plan counts against besides the per-instance-type ones, with their defaults in every Region.

QuotaCodeDefault
Longest run time for a processing jobL-EA47BDD3432,000
Longest run time for a training jobL-33A961FD432,000
Maximum number Spot instances allowed per SageMaker HyperPod clusterL-41BFBF4120
Maximum number instances allowed per SageMaker HyperPod clusterL-2CE978FC20
Maximum number of Studio spaces allowed per accountL-8E5333B46
Maximum number of Studio user profiles allowed per accountL-AC46C40F2
Maximum number of instances per endpointL-ED8DEE9B4
Maximum number of instances per processing jobL-A5D6339920
Maximum number of instances per spot training jobL-DB62C86420
Maximum number of instances per training jobL-622CFD7020
Maximum number of running Studio apps allowed per accountL-4E39DDC440
Maximum number of serverless endpointsL-99AD19BF5
Maximum total concurrency that can be allocated across all serverless endpointsL-9630010210
Number of instances across active endpointsL-7A3DF6114
Number of instances across all processing jobsL-F311B08F4
Number of instances across all spot training jobsL-939580824
Number of instances across all training jobsL-00C91CB54
Number of instances across all transform jobsL-60D2A6F04
Total number of Spot instances allowed across SageMaker HyperPod clustersL-97A2C7240
Total number of instances allowed across SageMaker HyperPod clustersL-3308CCC70
Total number of notebook instancesL-04CE2E678

Frequently asked questions

Does an asynchronous endpoint have its own quota?

No. An asynchronous endpoint is created like a real-time one, with an instance type and count, and counts against the same "for endpoint usage" quota and instances across active endpoints. Serverless endpoints have no instance type; they count against the number of serverless endpoints and the concurrency allocated across them.

Why does an endpoint update fail with ResourceLimitExceeded when the endpoint runs fine?

By default SageMaker updates an endpoint blue/green: it provisions the new fleet, shifts the traffic and only then terminates the old fleet. For that time the endpoint runs twice its instances. A rolling deployment replaces a batch of 5-50% of the fleet at a time and needs only that much room.

Do spot training jobs and warm pools use the training job quota?

Managed spot training has its own quotas per instance type and across all spot training jobs. A warm pool keeps a finished job's instances for up to an hour, counted against the instance type's warm pool quota; warm pools do not work with spot instances.

Is what I paste sent anywhere?

No. Everything runs in your browser. The page only counts that the finder was used and for which use, never the instance type, the numbers or the pasted values.

References

Amazon SageMaker AI endpoints and quotas (AWS General Reference)
Supported Regions and Quotas (Amazon SageMaker AI Developer Guide)
Deploy models with Amazon SageMaker Serverless Inference
SageMaker AI Managed Warm Pools
Prerequisites for using SageMaker HyperPod
Blue/Green Deployments (Amazon SageMaker AI Developer Guide)
ListServiceQuotas (Service Quotas API Reference)