How a SageMaker endpoint starts, and where it fails
CreateEndpoint and UpdateEndpoint return at once and the endpoint stays Creating or Updating while SageMaker works through the steps below. When one fails, the status becomes Failed and FailureReason says which - the console shows it on the endpoint's page, aws sagemaker describe-endpoint returns it, and the SageMaker Python SDK raises it as UnexpectedStatusException: Error hosting endpoint NAME: Failed. Reason: ....
| Step | What fails there |
|---|---|
| The request | ResourceLimitExceeded - the instance type's endpoint quota, 0 by default for every GPU type - comes back from the call itself, before the endpoint is created |
| Instances | InsufficientInstanceCapacity; in a VPC, too few Availability Zones that offer the type |
| Image and model artifacts | The image or the model.tar.gz cannot be read with the execution role, is in another Region, or the archive cannot be unpacked |
| Container start | CannotStartContainerError: docker run IMAGE serve starts no server |
| Health check | "did not pass the ping health check": the model did not load, or not in time |
An endpoint whose creation failed can only be deleted: fix the cause, delete it and create it again. An update whose instances fail the same way is not completed - check the endpoint's status before you try again. Errors from InvokeEndpoint - ModelError, timeouts, "not found" - come from an endpoint that is InService; the decoder knows those too.
The container contract
Most failures after the instances start come from a container that does not do what SageMaker expects of it:
- SageMaker runs it as
docker run IMAGE serve; theserveargument replaces the Dockerfile'sCMD. Use the exec form ofENTRYPOINT, and nottini. - A web server listens on port 8080 and answers
GET /pingandPOST /invocations. It must accept a connection within 250 ms, answer a ping within 2 seconds and an invocation within 60 seconds. - Pings must get 200 consistently within 8 minutes of the container's start, unless
ContainerStartupHealthCheckTimeoutInSecondsgives it longer. - The model artifacts are unpacked to
/opt/ml/modelbefore the container starts, read-only. The image and the artifacts must be in the same Region as the model, and the execution role must be able to read both. - On GPU instances, the image must not bundle NVIDIA drivers.
Everything the container prints goes to the log group /aws/sagemaker/Endpoints/ENDPOINT_NAME, one stream per variant and instance - the first place to look after a failed health check.
Settings for large models
A large language model takes minutes to download and load, longer than the defaults allow. AWS recommends three production variant settings in the endpoint configuration:
| Setting | When to raise it | Range |
|---|---|---|
VolumeSizeInGB | The model is over 30 GB and the instance has no local disk - slightly above the model's size | 1-512 |
ModelDataDownloadTimeoutInSeconds | The model is over 40 GB | 60-3600 |
ContainerStartupHealthCheckTimeoutInSeconds | The logs show the model loading when the health check gives up | 60-3600 |
A model that does not fit in the instance's GPU memory fails the health check whatever the timeouts. The SageMaker GPU Instance Calculator shows which instances hold a model, and the quota code each needs.
Frequently asked questions
Is the error sent anywhere?
No. It is read in your browser: nothing you paste is uploaded, processed on a server or stored. The page only counts that the decoder was used and which error it recognized, never the text.
The FailureReason only says to check CloudWatch Logs - where are they?
In the log group /aws/sagemaker/Endpoints/ENDPOINT_NAME, in the endpoint's Region. aws logs tail /aws/sagemaker/Endpoints/ENDPOINT_NAME --since 1h prints them. If the group is empty, the container never started or printed nothing - start the image locally with docker run IMAGE serve and the model in /opt/ml/model.
Does a serverless endpoint fail differently?
It runs the same container, with less room: a 10 GB image at most, 1 to 6 GB of memory, 5 GB of disk and no GPU. A container that writes files may fail its health check there when it works on instances - give other write access to those directories in the image.
Why does my request time out after 60 seconds?
A real-time invocation must be answered within 60 seconds and its body may be 25 MB at most. Longer or larger requests belong on an asynchronous endpoint: up to 60 minutes and 1 GB per request, read from S3.
References
Troubleshoot Amazon SageMaker AI model deployments
Custom Inference Code with Hosting Services
SageMaker AI endpoint parameters for large model inference
Give SageMaker AI Hosted Endpoints Access to Resources in Your Amazon VPC
Deploy models with Amazon SageMaker Serverless Inference
InvokeEndpoint
Why does my SageMaker endpoint go into the failed state?
How do I resolve insufficient capacity errors?