FREE TOOL

SageMaker GPU Instance Calculator

Find the ml.* GPU instances that hold your model on a SageMaker endpoint or training job - weights, KV cache and optimizer states against each instance's GPU memory - with the hourly price in your Region and the quota code to request, or see what fits on an instance.

  • Your data never leaves your browser: everything is calculated by JavaScript on this page, not on a server.
  • Nothing you enter is uploaded, processed on a server or stored. Check it in your browser's developer tools (Network tab).
  • Once the page has loaded, the tool works without an internet connection.

Model

Pick a model, or paste a model's config.json and enter its parameter count.

Paste a config.json

Workload

An endpoint holds the weights and a KV cache for every sequence it serves at once; a training job holds the weights, gradients and optimizer states.

See the instances ↓
Choosing compute for a workload is a core AWS Solutions Architect Associate skillTry free SAA-C03 practice questions with answers and explanations.SAA-C03 questions →

How much GPU memory a model needs

AWS's guide to instance types for large model inference (DJL LMI, the container SageMaker serves large models with) starts from the weights: the number of parameters times the bytes each takes - 4 in fp32, 2 in fp16 or bf16, 1 in fp8 or int8, and half a byte for 4-bit quantized models such as AWQ or GPTQ. A model with 8 billion parameters in bf16 needs 16 GB before it answers anything.

Every sequence it serves then needs a KV cache: for every token, a key and a value per layer, per KV head, per head dimension - 2 × bytes × layers × KV heads × head dimension. AWS's guide multiplies by the hidden size instead, which is right for models where every attention head has its own key and value; Llama 3, Qwen3 and Mistral share 8 KV heads among 32 to 64 attention heads, so the calculator uses the KV heads from the model's config.json - for Llama 3.3 70B, 328 KB per token instead of 2.6 MB. The cache grows with the tokens per sequence (prompt and output) times the sequences served at once.

The engine does not get all of a GPU's memory: LMI's vLLM examples set gpu_memory_utilization to 0.9, leaving the rest to CUDA and activations. A model is split over an instance's GPUs with tensor parallelism, which needs a GPU count that divides the model's attention heads, so a model with 28 heads uses 4 of an instance's 8 GPUs. GPU memory is counted as the EC2 documentation lists it, in GiB the GPU can use - an A10G or L4 sold as 24 GB has 22 GiB.

Training needs far more. With mixed precision and the Adam optimizer, AWS counts about 20 bytes per parameter - the bf16 weight and gradient, two fp32 optimizer states and fp32 copies of the weight and gradient - before the activations, so a 10-billion-parameter model needs at least 200 GB. The calculator spreads these states over all the GPUs of an instance (sharded data parallelism, or FSDP); P instances, which AWS recommends for distributed training, also spread them over as many instances as it takes, up to 20 - the default limit per training job - while a G instance has to hold the job alone. LoRA keeps the base model frozen in bf16 and trains only a small adapter; QLoRA keeps the frozen model in 4 bits.

Why SageMaker answers ResourceLimitExceeded

Every GPU instance type has its own quotas per Region - one for endpoints, one for training jobs, others for notebooks, processing and HyperPod - and AWS sets every GPU instance's endpoint and training quota to 0 by default. The first endpoint on a new instance type fails with an error such as "The account-level service limit 'ml.g5.12xlarge for endpoint usage' is 0 Instances, with current utilization of 0 Instances and a request delta of 1 Instances". The calculator names the quota and its code for the instance it picks; request the increase in the Service Quotas console or with the command it shows, before you deploy.

SageMaker GPU instance types

The GPU instances SageMaker offers, with their US East (N. Virginia) prices per hour on demand. Prices and the instances offered differ by Region; the calculator uses the Region you choose. The prices come from the AWS Price List API and the GPUs from the EC2 documentation, read again on every build of this site; the prices shown were published on 2026-09-28.

InstanceGPUsGPU memoryEndpointTraining
ml.g4dn.xlarge1 × T416 GiB$0.736$0.736
ml.g4dn.2xlarge1 × T416 GiB$0.94$0.94
ml.g4dn.4xlarge1 × T416 GiB$1.51$1.51
ml.g4dn.8xlarge1 × T416 GiB$2.72$2.72
ml.g4dn.12xlarge4 × T464 GiB$4.89$4.89
ml.g4dn.16xlarge1 × T416 GiB$5.44$5.44
ml.g5.xlarge1 × A10G22 GiB$1.41$1.41
ml.g5.2xlarge1 × A10G22 GiB$1.52$1.52
ml.g5.4xlarge1 × A10G22 GiB$2.03$2.03
ml.g5.8xlarge1 × A10G22 GiB$3.06$3.06
ml.g5.12xlarge4 × A10G88 GiB$7.09$7.09
ml.g5.16xlarge1 × A10G22 GiB$5.12$5.12
ml.g5.24xlarge4 × A10G88 GiB$10.18$10.18
ml.g5.48xlarge8 × A10G176 GiB$20.36$20.36
ml.g6.xlarge1 × L422 GiB$1.13$1.13
ml.g6.2xlarge1 × L422 GiB$1.22$1.22
ml.g6.4xlarge1 × L422 GiB$1.65$1.65
ml.g6.8xlarge1 × L422 GiB$2.52$2.52
ml.g6.12xlarge4 × L488 GiB$5.75$5.75
ml.g6.16xlarge1 × L422 GiB$4.25$4.25
ml.g6.24xlarge4 × L488 GiB$8.34$8.34
ml.g6.48xlarge8 × L4176 GiB$16.69$16.69
ml.g6e.xlarge1 × L40S44 GiB$2.61$2.61
ml.g6e.2xlarge1 × L40S44 GiB$2.80$2.80
ml.g6e.4xlarge1 × L40S44 GiB$3.76$3.76
ml.g6e.8xlarge1 × L40S44 GiB$5.66$5.66
ml.g6e.12xlarge4 × L40S176 GiB$13.12$13.12
ml.g6e.16xlarge1 × L40S44 GiB$9.47$9.47
ml.g6e.24xlarge4 × L40S176 GiB$18.83$18.83
ml.g6e.48xlarge8 × L40S352 GiB$37.66$37.66
ml.g7.2xlarge1 × RTX PRO 450032 GiB$3.15$3.15
ml.g7.4xlarge1 × RTX PRO 450032 GiB$3.80$3.80
ml.g7.8xlarge1 × RTX PRO 450032 GiB$5.11$5.11
ml.g7.12xlarge2 × RTX PRO 450064 GiB$8.91$8.91
ml.g7.24xlarge4 × RTX PRO 4500128 GiB$17.82$17.82
ml.g7.48xlarge8 × RTX PRO 4500256 GiB$35.64$35.64
ml.g7e.2xlarge1 × RTX PRO Server 600096 GiB$4.20$4.20
ml.g7e.4xlarge1 × RTX PRO Server 600096 GiB$5.00$5.00
ml.g7e.8xlarge1 × RTX PRO Server 600096 GiB$6.59$6.59
ml.g7e.12xlarge2 × RTX PRO Server 6000192 GiB$10.36$10.36
ml.g7e.24xlarge4 × RTX PRO Server 6000384 GiB$20.72$20.72
ml.g7e.48xlarge8 × RTX PRO Server 6000768 GiB$41.43$41.43
ml.gr6.4xlarge1 × L422 GiB--
ml.gr6.8xlarge1 × L422 GiB--
ml.p3dn.24xlarge8 × V100256 GiB-$35.89
ml.p4d.24xlarge8 × A100320 GiB$25.25$25.25
ml.p4de.24xlarge8 × A100640 GiB$31.56$31.56
ml.p5.4xlarge1 × H10080 GiB--
ml.p5.48xlarge8 × H100640 GiB$63.30$63.30
ml.p5en.48xlarge8 × H2001128 GiB$72.80$72.80
ml.p6-b200.48xlarge8 × B2001432 GiB$131.02$131.02
ml.p6-b300.48xlarge8 × B3002144 GiB$163.78$163.78

Frequently asked questions

Why does my endpoint run out of memory although the model fits?

The weights are only the floor. vLLM reserves the KV cache for the longest sequence and the batch it was configured for (OPTION_MAX_MODEL_LEN, OPTION_MAX_ROLLING_BATCH_SIZE); a model loaded with its full context window - 128K tokens for Llama 3.1 - needs a cache many times its weights. Lower the sequence length, the batch size, or store the cache in fp8.

Why are Inferentia and Trainium instances not listed?

ml.inf2 and ml.trn1 instances run models compiled with the AWS Neuron SDK, which shards and caches them its own way, so GPU memory math does not carry over. Fractional GPU instances (G6f) are left out too.

What about serverless inference?

SageMaker Serverless Inference does not support GPUs, so a model that needs one runs on a real-time or asynchronous endpoint, both billed per instance-hour.

Is what I enter sent anywhere?

No. The calculation runs in your browser. The page only counts that the calculator was used and in which mode, never the model or the numbers.

References

Instance Type Selection (LMI container documentation)
Choosing instance types for large model inference (Amazon SageMaker AI Developer Guide)
Introduction to Model Parallelism (Amazon SageMaker AI Developer Guide)
Specifications for Amazon EC2 accelerated computing instances
Amazon SageMaker AI endpoints and quotas (AWS General Reference)
Amazon SageMaker AI pricing