How Amazon Bedrock counts a request against the quotas
On the bedrock-runtime endpoint - InvokeModel, Converse and their streaming variants - every model has a quota of tokens per minute (input and output together) and, for some models, of requests per minute, in each Region. A request counts twice:
- When it starts, Bedrock deducts its input tokens plus
max_tokens. If that is more than is left of the minute's quota, the request is throttled with a 429. - While it runs and when it ends, the deduction is adjusted to what the request really used: input tokens + cache write tokens + output tokens × the model's burndown rate. The unused part of
max_tokensgoes back to the quota.
Tokens read from the prompt cache never count. You are billed for the tokens alone, not the burndown: AWS's example is a request to Claude Sonnet 4 with 1,000 input and 100 output tokens, which uses 1,500 tokens of the quota and is billed for 1,100. AWS lists the burndown rates on its "How tokens are counted" page:
- The burndown rate for Anthropic Claude models version 4.8 is 15x for output tokens.
- The burndown rate for Anthropic Claude Opus 5.5, Claude Sonnet 5, Claude Opus 5, and Claude Fable 5.1 is 10x for output tokens.
- For all other Anthropic models version 4.7 and below, the burndown is 5x for output tokens.
- The burndown rate for OpenAI GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna on the bedrock-runtime endpoint is 10x for output tokens.
- For all other models, the burndown rate is 1:1.
Models AWS's page does not name yet (newer Claude models) get the highest Anthropic rate in the calculator, so the result errs on the safe side - enter the rate yourself if AWS names it. AWS's Global cross-Region inference page still has an older, shorter list; the calculator follows the dedicated page.
Which quotas apply
| Quota | Scope | Counts |
|---|---|---|
| On-demand tokens / requests per minute | Per model and Region | Calls with the model ID in the Region |
| Cross-Region tokens / requests per minute | Per model and source Region | Calls with a Geo inference profile (us., eu., ...) |
| Global cross-Region tokens / requests per minute | Per model and source Region | Calls with a Global inference profile (global.) |
| Cross-Model Max Tokens Per Day | Per account and Region, all models | Every model's tokens, estimated from on-demand prices |
| Model tokens per day | Some models, per Region | The model's tokens; doubled for cross-Region calls, dropped once AWS approves a tokens-per-minute increase |
| bedrock-mantle input / output tokens per minute | Per model and Region | The OpenAI- and Anthropic-compatible APIs; separate from bedrock-runtime |
The bedrock-mantle endpoint counts differently: input tokens plus max_tokens are checked against its input quota when a request arrives (the unused part returns at the end), output tokens count against a separate output quota - and when that runs out, generation stops and the response ends early instead of a 429. It has no requests-per-minute quota and no burndown, and AWS publishes its quotas only for some models; for the others throughput depends on the service's capacity.
Concurrency, and why max_tokens matters
Because max_tokens is deducted in full when a request starts, it decides how many requests can run at once, not the tokens you really use. AWS's own example: 3,000 input tokens, 1,000 cache write tokens and 1,000 output tokens with a 5x burndown count 9,000 tokens in the end - but with max_tokens at 32,000 the request first reserves 36,000. Lowered to 1,250, it reserves 5,250, and about seven times as many requests fit in the same quota at the same moment. The calculator shows both numbers: requests per minute (limited by what requests count in the end) and requests at once (limited by what they reserve at the start).
Why you get a 429 at low traffic
- Your account's quota is lower than the default. AWS adjusts an account's defaults by the Region, payment history and fraudulent usage, and new accounts may get reduced quotas. Look up the values applied to your account in Service Quotas, in the Region you call.
- One request reserves more than the quota. A large
max_tokenswith a long input can exceed a small tokens-per-minute quota on its own. - The burndown. An output token of many Claude models uses 5, 10 or 15 tokens of the quota.
- The wrong quota. On-demand, Geo, Global and bedrock-mantle quotas are separate: raising one does not help calls made the other way.
- The daily quota. Cross-Model Max Tokens Per Day is shared by every model and application in the account and Region.
- It is not a quota at all.
ServiceUnavailableException(503) andoverloaded_error(529) mean Bedrock has no capacity at the moment - the Bedrock model finder's "Paste an error" mode explains them.
The Why 429 at low traffic? mode above checks these against your workload and quotas.
Frequently asked questions
How do I request a quota increase?
For bedrock-runtime, request an increase of the model's "Cross-Region InvokeModel tokens per minute" quota in Service Quotas; AWS Support then offers to raise the on-demand tokens per minute and the daily quota too. AWS gives priority to accounts whose traffic already uses up their quota, and grants no increases for Legacy models. For bedrock-mantle, open a support case with the endpoint, Region, model, quota and value.
Does a cross-Region inference profile raise my quota?
It uses a different quota: a Geo or Global profile counts against the model's cross-Region quotas, which are often higher than the on-demand ones. The calculator shows the default of each way the model can be called from your Region.
Where does the data come from?
The default quotas come from the Amazon Bedrock quota table of the AWS General Reference, and the models, their IDs and Regions from AWS's model cards - both read again on every build of this site, the quotas last changed on 2026-09-29. The counting rules and burndown rates are those of the Bedrock User Guide's quota pages, and every build checks they are still there. Some models have no quota in AWS's table; the calculator then asks for your value.
Is what I enter sent anywhere?
No. The calculation runs in your browser. The page only counts that the calculator was used, with the way of calling the model and the first cause the diagnosis finds, never the numbers.
References
How tokens are counted in Amazon Bedrock (Amazon Bedrock User Guide)
Quotas for the bedrock-runtime endpoint (Amazon Bedrock User Guide)
Quotas for the bedrock-mantle endpoint (Amazon Bedrock User Guide)
Quotas for Amazon Bedrock (Amazon Bedrock User Guide)
Amazon Bedrock service quotas (AWS General Reference)
Troubleshooting Amazon Bedrock API error codes (Amazon Bedrock User Guide)
list-service-quotas (AWS CLI Command Reference)