The AgentCore limits that cannot be raised
Most AgentCore quotas - agents per account, sessions at once, requests per second - can be raised in Service Quotas. The limits below cannot; an agent that needs more has to be built around them. All apply to the default microVMs compute type (as of 2026-09-30):
| Limit | Value |
|---|---|
| Synchronous request | 15 minutes |
| Streamed response or WebSocket connection | 60 minutes |
| Background (asynchronous) task | 8 hours |
Session lifetime (maxLifetime, settable 60 s - 8 h) | 8 hours |
Idle timeout (idleRuntimeSessionTimeout, settable 60 s - 8 h) | 15 minutes by default |
| Request or response payload | 100 MB |
| Streamed chunk | 10 MB |
| WebSocket frame | 64 KB, 250 frames per second per connection |
| Hardware per session | 2 vCPU, 8 GB |
| Session storage | 1 GB, about 100,000-200,000 files |
| Docker image | 2 GB |
| ZIP package (direct code deployment) | 250 MB, 750 MB unzipped |
| Environment variables | 4 KB on V1; 2.5 KB (container) or 1.5 KB (ZIP) on V2 for now |
V2: healthy /ping after start | 120 seconds |
| Memory: message / event / messages per event | 100 KB / 10 MB / 100 |
| Memory: CreateEvent per actor and session | 5 per second (10 without conversational payloads) |
| Memory: event expiry / strategies per memory | 7-365 days / 6 |
| Code Interpreter: session hardware / disk | 2 vCPU, 8 GB / 10 GB |
| Code Interpreter: synchronous call / asynchronous command | 15 minutes / 8 hours |
| Browser: session hardware / disk | 1 vCPU, 4 GB / 10 GB |
| Browser: automation streams per session | 1 |
When a timeout ends a session, the next request with the same session ID gets a new microVM: nothing from the old one's memory or disk is left. Keep a conversation in AgentCore Memory or S3, not in the process. The AgentCore Error Decoder reads the errors these limits produce.
CPU vs memory billing
AgentCore Runtime bills per second, for CPU and memory separately, and the two are counted differently:
- CPU only while it works. Time spent waiting for the model's response, a tool, an API or a database - I/O wait - is free, as long as no background process keeps running.
- Memory for every second the session exists, at the peak reached so far, from the microVM's start through idle periods until the session ends - which is after the idle timeout, 15 minutes after the last response by default. The minimum is 128 MB.
An agent mostly waits for its model, so CPU is a small part of the bill and memory held through pauses and the idle timeout a large one. The platform version decides how much of it stays: V1 ($0.0895 per vCPU-hour, $0.00945 per GB-hour) keeps memory until the session ends; V2 ($0.1276 and $0.0169) reclaims memory the agent has not used for 120 seconds. AWS's own example - a 10-minute session on V2, waiting 90% of the time on 1 vCPU - costs $0.002127 of CPU and $0.004576 of memory. On V1 a shorter idle timeout (idleRuntimeSessionTimeout) and a smaller memory footprint lower the bill most; a long idle timeout kept for users who pause costs its memory in every session. Browser and Code Interpreter sessions are billed the same way, at V1's rates.
Ways around each limit
- Longer than 15 minutes: stream the response (up to 60 minutes), or answer at once and keep working in the background - the session reports
HealthyBusyon/pingwhile it works, for up to 8 hours. - Longer than 8 hours, more than 2 vCPU or 8 GB, or a GPU: the Instances compute type runs a session on an EC2 instance of the type you choose, in your account, for up to 14 days. It is chosen when the runtime is created and cannot be changed later.
- Payloads over 100 MB and files over 1 GB: put the data in S3 and pass its key; the agent reads it with its execution role.
- Users who pause longer than the idle timeout: raise
idleRuntimeSessionTimeout(up to 8 hours), or let the session end and load the conversation from AgentCore Memory when the next one starts - the cheaper choice when pauses are long. - More than one automation client on a browser: start a Browser session per client.
Frequently asked questions
Is what I enter sent anywhere?
No. The check runs in your browser. The page only counts that the checker was used, with which kind of response and which first limit hit, never the numbers.
Which limits can I raise?
Those marked "adjustable" in AWS's quota tables, through Service Quotas or AWS Support: active sessions (5,000 in US East (N. Virginia) and US West (Oregon), 2,500 elsewhere), 25 new sessions per second, 1,000 data plane requests per second, agents, versions and endpoints per account, and every Gateway quota. The limits on this page marked "cannot be raised" are fixed.
Should I use V1 or V2?
V2 starts every session from a snapshot, so cold starts stay about two seconds whatever the image size (P75 in AWS's tests with images of 200 MB to 2 GB), and it bills the memory the agent actually uses, at a higher rate. It is available in US East (N. Virginia), US East (Ohio), US West (Oregon), Europe (Ireland) and Asia Pacific (Tokyo); CloudFormation and the CDK cannot set it yet, and a V2 runtime takes minutes to create or update. V1 stays the default when platformVersion is left out.
Does the estimate include the model's tokens?
No - only the Runtime session. The model calls are billed by Amazon Bedrock per token; the Bedrock Throttling Calculator shows whether your agent's requests fit the model's quotas.
References
Quotas for Amazon Bedrock AgentCore
AgentCore Runtime microVMs and platform versions
Configure AgentCore lifecycle settings
AgentCore Runtime Instances
Amazon Bedrock AgentCore pricing