How SageMaker counts quotas
SageMaker gives every instance type a separate quota for each use, per Region: "ml.g5.xlarge for endpoint usage", "ml.g5.xlarge for training job usage", for spot training, warm pools, processing, transform jobs, notebook instances, HyperPod clusters and Studio's JupyterLab and Code Editor apps. Next to them are account-wide totals - instances across all endpoints, across all training jobs - and maxima per endpoint, job or cluster. A plan has to fit all of them, and each counts what already runs in the Region.
The defaults are low. In AWS's quota table almost every instance type starts at 0 for endpoints and on-demand training jobs, CPU instances such as ml.m5.xlarge included, and every GPU type at 0 for every use; the totals across endpoints and jobs start at 4 instances, and HyperPod's at 0. So the first endpoint or training job on a new type fails with ResourceLimitExceeded until you request an increase in the Service Quotas console or with the commands above. A quota that is not in AWS's table for a use means SageMaker does not take that instance type for it - the finder lists only the types each use has a quota for.
The defaults here come from AWS's quota table and are read again on every build of this site (last read on 2026-10-02). Your account may have other values, set by AWS or by an earlier request; paste the output of list-service-quotas to compare against them.
SageMaker's account-wide quotas
The quotas a plan counts against besides the per-instance-type ones, with their defaults in every Region.
| Quota | Code | Default |
|---|---|---|
| Longest run time for a processing job | L-EA47BDD3 | 432,000 |
| Longest run time for a training job | L-33A961FD | 432,000 |
| Maximum number Spot instances allowed per SageMaker HyperPod cluster | L-41BFBF41 | 20 |
| Maximum number instances allowed per SageMaker HyperPod cluster | L-2CE978FC | 20 |
| Maximum number of Studio spaces allowed per account | L-8E5333B4 | 6 |
| Maximum number of Studio user profiles allowed per account | L-AC46C40F | 2 |
| Maximum number of instances per endpoint | L-ED8DEE9B | 4 |
| Maximum number of instances per processing job | L-A5D63399 | 20 |
| Maximum number of instances per spot training job | L-DB62C864 | 20 |
| Maximum number of instances per training job | L-622CFD70 | 20 |
| Maximum number of running Studio apps allowed per account | L-4E39DDC4 | 40 |
| Maximum number of serverless endpoints | L-99AD19BF | 5 |
| Maximum total concurrency that can be allocated across all serverless endpoints | L-96300102 | 10 |
| Number of instances across active endpoints | L-7A3DF611 | 4 |
| Number of instances across all processing jobs | L-F311B08F | 4 |
| Number of instances across all spot training jobs | L-93958082 | 4 |
| Number of instances across all training jobs | L-00C91CB5 | 4 |
| Number of instances across all transform jobs | L-60D2A6F0 | 4 |
| Total number of Spot instances allowed across SageMaker HyperPod clusters | L-97A2C724 | 0 |
| Total number of instances allowed across SageMaker HyperPod clusters | L-3308CCC7 | 0 |
| Total number of notebook instances | L-04CE2E67 | 8 |
Frequently asked questions
Does an asynchronous endpoint have its own quota?
No. An asynchronous endpoint is created like a real-time one, with an instance type and count, and counts against the same "for endpoint usage" quota and instances across active endpoints. Serverless endpoints have no instance type; they count against the number of serverless endpoints and the concurrency allocated across them.
Why does an endpoint update fail with ResourceLimitExceeded when the endpoint runs fine?
By default SageMaker updates an endpoint blue/green: it provisions the new fleet, shifts the traffic and only then terminates the old fleet. For that time the endpoint runs twice its instances. A rolling deployment replaces a batch of 5-50% of the fleet at a time and needs only that much room.
Do spot training jobs and warm pools use the training job quota?
Managed spot training has its own quotas per instance type and across all spot training jobs. A warm pool keeps a finished job's instances for up to an hour, counted against the instance type's warm pool quota; warm pools do not work with spot instances.
Is what I paste sent anywhere?
No. Everything runs in your browser. The page only counts that the finder was used and for which use, never the instance type, the numbers or the pasted values.
References
Amazon SageMaker AI endpoints and quotas (AWS General Reference)
Supported Regions and Quotas (Amazon SageMaker AI Developer Guide)
Deploy models with Amazon SageMaker Serverless Inference
SageMaker AI Managed Warm Pools
Prerequisites for using SageMaker HyperPod
Blue/Green Deployments (Amazon SageMaker AI Developer Guide)
ListServiceQuotas (Service Quotas API Reference)