How much GPU memory a model needs
AWS's guide to instance types for large model inference (DJL LMI, the container SageMaker serves large models with) starts from the weights: the number of parameters times the bytes each takes - 4 in fp32, 2 in fp16 or bf16, 1 in fp8 or int8, and half a byte for 4-bit quantized models such as AWQ or GPTQ. A model with 8 billion parameters in bf16 needs 16 GB before it answers anything.
Every sequence it serves then needs a KV cache: for every token, a key and a value per layer, per KV head, per head dimension - 2 × bytes × layers × KV heads × head dimension. AWS's guide multiplies by the hidden size instead, which is right for models where every attention head has its own key and value; Llama 3, Qwen3 and Mistral share 8 KV heads among 32 to 64 attention heads, so the calculator uses the KV heads from the model's config.json - for Llama 3.3 70B, 328 KB per token instead of 2.6 MB. The cache grows with the tokens per sequence (prompt and output) times the sequences served at once.
The engine does not get all of a GPU's memory: LMI's vLLM examples set gpu_memory_utilization to 0.9, leaving the rest to CUDA and activations. A model is split over an instance's GPUs with tensor parallelism, which needs a GPU count that divides the model's attention heads, so a model with 28 heads uses 4 of an instance's 8 GPUs. GPU memory is counted as the EC2 documentation lists it, in GiB the GPU can use - an A10G or L4 sold as 24 GB has 22 GiB.
Training needs far more. With mixed precision and the Adam optimizer, AWS counts about 20 bytes per parameter - the bf16 weight and gradient, two fp32 optimizer states and fp32 copies of the weight and gradient - before the activations, so a 10-billion-parameter model needs at least 200 GB. The calculator spreads these states over all the GPUs of an instance (sharded data parallelism, or FSDP); P instances, which AWS recommends for distributed training, also spread them over as many instances as it takes, up to 20 - the default limit per training job - while a G instance has to hold the job alone. LoRA keeps the base model frozen in bf16 and trains only a small adapter; QLoRA keeps the frozen model in 4 bits.
Why SageMaker answers ResourceLimitExceeded
Every GPU instance type has its own quotas per Region - one for endpoints, one for training jobs, others for notebooks, processing and HyperPod - and AWS sets every GPU instance's endpoint and training quota to 0 by default. The first endpoint on a new instance type fails with an error such as "The account-level service limit 'ml.g5.12xlarge for endpoint usage' is 0 Instances, with current utilization of 0 Instances and a request delta of 1 Instances". The calculator names the quota and its code for the instance it picks; request the increase in the Service Quotas console or with the command it shows, before you deploy.
SageMaker GPU instance types
The GPU instances SageMaker offers, with their US East (N. Virginia) prices per hour on demand. Prices and the instances offered differ by Region; the calculator uses the Region you choose. The prices come from the AWS Price List API and the GPUs from the EC2 documentation, read again on every build of this site; the prices shown were published on 2026-09-28.
| Instance | GPUs | GPU memory | Endpoint | Training |
|---|---|---|---|---|
| ml.g4dn.xlarge | 1 × T4 | 16 GiB | $0.736 | $0.736 |
| ml.g4dn.2xlarge | 1 × T4 | 16 GiB | $0.94 | $0.94 |
| ml.g4dn.4xlarge | 1 × T4 | 16 GiB | $1.51 | $1.51 |
| ml.g4dn.8xlarge | 1 × T4 | 16 GiB | $2.72 | $2.72 |
| ml.g4dn.12xlarge | 4 × T4 | 64 GiB | $4.89 | $4.89 |
| ml.g4dn.16xlarge | 1 × T4 | 16 GiB | $5.44 | $5.44 |
| ml.g5.xlarge | 1 × A10G | 22 GiB | $1.41 | $1.41 |
| ml.g5.2xlarge | 1 × A10G | 22 GiB | $1.52 | $1.52 |
| ml.g5.4xlarge | 1 × A10G | 22 GiB | $2.03 | $2.03 |
| ml.g5.8xlarge | 1 × A10G | 22 GiB | $3.06 | $3.06 |
| ml.g5.12xlarge | 4 × A10G | 88 GiB | $7.09 | $7.09 |
| ml.g5.16xlarge | 1 × A10G | 22 GiB | $5.12 | $5.12 |
| ml.g5.24xlarge | 4 × A10G | 88 GiB | $10.18 | $10.18 |
| ml.g5.48xlarge | 8 × A10G | 176 GiB | $20.36 | $20.36 |
| ml.g6.xlarge | 1 × L4 | 22 GiB | $1.13 | $1.13 |
| ml.g6.2xlarge | 1 × L4 | 22 GiB | $1.22 | $1.22 |
| ml.g6.4xlarge | 1 × L4 | 22 GiB | $1.65 | $1.65 |
| ml.g6.8xlarge | 1 × L4 | 22 GiB | $2.52 | $2.52 |
| ml.g6.12xlarge | 4 × L4 | 88 GiB | $5.75 | $5.75 |
| ml.g6.16xlarge | 1 × L4 | 22 GiB | $4.25 | $4.25 |
| ml.g6.24xlarge | 4 × L4 | 88 GiB | $8.34 | $8.34 |
| ml.g6.48xlarge | 8 × L4 | 176 GiB | $16.69 | $16.69 |
| ml.g6e.xlarge | 1 × L40S | 44 GiB | $2.61 | $2.61 |
| ml.g6e.2xlarge | 1 × L40S | 44 GiB | $2.80 | $2.80 |
| ml.g6e.4xlarge | 1 × L40S | 44 GiB | $3.76 | $3.76 |
| ml.g6e.8xlarge | 1 × L40S | 44 GiB | $5.66 | $5.66 |
| ml.g6e.12xlarge | 4 × L40S | 176 GiB | $13.12 | $13.12 |
| ml.g6e.16xlarge | 1 × L40S | 44 GiB | $9.47 | $9.47 |
| ml.g6e.24xlarge | 4 × L40S | 176 GiB | $18.83 | $18.83 |
| ml.g6e.48xlarge | 8 × L40S | 352 GiB | $37.66 | $37.66 |
| ml.g7.2xlarge | 1 × RTX PRO 4500 | 32 GiB | $3.15 | $3.15 |
| ml.g7.4xlarge | 1 × RTX PRO 4500 | 32 GiB | $3.80 | $3.80 |
| ml.g7.8xlarge | 1 × RTX PRO 4500 | 32 GiB | $5.11 | $5.11 |
| ml.g7.12xlarge | 2 × RTX PRO 4500 | 64 GiB | $8.91 | $8.91 |
| ml.g7.24xlarge | 4 × RTX PRO 4500 | 128 GiB | $17.82 | $17.82 |
| ml.g7.48xlarge | 8 × RTX PRO 4500 | 256 GiB | $35.64 | $35.64 |
| ml.g7e.2xlarge | 1 × RTX PRO Server 6000 | 96 GiB | $4.20 | $4.20 |
| ml.g7e.4xlarge | 1 × RTX PRO Server 6000 | 96 GiB | $5.00 | $5.00 |
| ml.g7e.8xlarge | 1 × RTX PRO Server 6000 | 96 GiB | $6.59 | $6.59 |
| ml.g7e.12xlarge | 2 × RTX PRO Server 6000 | 192 GiB | $10.36 | $10.36 |
| ml.g7e.24xlarge | 4 × RTX PRO Server 6000 | 384 GiB | $20.72 | $20.72 |
| ml.g7e.48xlarge | 8 × RTX PRO Server 6000 | 768 GiB | $41.43 | $41.43 |
| ml.gr6.4xlarge | 1 × L4 | 22 GiB | - | - |
| ml.gr6.8xlarge | 1 × L4 | 22 GiB | - | - |
| ml.p3dn.24xlarge | 8 × V100 | 256 GiB | - | $35.89 |
| ml.p4d.24xlarge | 8 × A100 | 320 GiB | $25.25 | $25.25 |
| ml.p4de.24xlarge | 8 × A100 | 640 GiB | $31.56 | $31.56 |
| ml.p5.4xlarge | 1 × H100 | 80 GiB | - | - |
| ml.p5.48xlarge | 8 × H100 | 640 GiB | $63.30 | $63.30 |
| ml.p5en.48xlarge | 8 × H200 | 1128 GiB | $72.80 | $72.80 |
| ml.p6-b200.48xlarge | 8 × B200 | 1432 GiB | $131.02 | $131.02 |
| ml.p6-b300.48xlarge | 8 × B300 | 2144 GiB | $163.78 | $163.78 |
Frequently asked questions
Why does my endpoint run out of memory although the model fits?
The weights are only the floor. vLLM reserves the KV cache for the longest sequence and the batch it was configured for (OPTION_MAX_MODEL_LEN, OPTION_MAX_ROLLING_BATCH_SIZE); a model loaded with its full context window - 128K tokens for Llama 3.1 - needs a cache many times its weights. Lower the sequence length, the batch size, or store the cache in fp8.
Why are Inferentia and Trainium instances not listed?
ml.inf2 and ml.trn1 instances run models compiled with the AWS Neuron SDK, which shards and caches them its own way, so GPU memory math does not carry over. Fractional GPU instances (G6f) are left out too.
What about serverless inference?
SageMaker Serverless Inference does not support GPUs, so a model that needs one runs on a real-time or asynchronous endpoint, both billed per instance-hour.
Is what I enter sent anywhere?
No. The calculation runs in your browser. The page only counts that the calculator was used and in which mode, never the model or the numbers.
References
Instance Type Selection (LMI container documentation)
Choosing instance types for large model inference (Amazon SageMaker AI Developer Guide)
Introduction to Model Parallelism (Amazon SageMaker AI Developer Guide)
Specifications for Amazon EC2 accelerated computing instances
Amazon SageMaker AI endpoints and quotas (AWS General Reference)
Amazon SageMaker AI pricing