FREE TOOL

Bedrock Throttling Calculator

How many requests per minute and how many at once an Amazon Bedrock model's quotas allow for your workload, what to change, and why you get a 429 ThrottlingException at low traffic.

  • Your data never leaves your browser: everything is calculated by JavaScript on this page, not on a server.
  • Nothing you enter is uploaded, processed on a server or stored. Check it in your browser's developer tools (Network tab).
  • Once the page has loaded, the tool works without an internet connection.

Model

The Region you call Amazon Bedrock in, the model, and the ID your code sends - each way of calling a model has its own quotas.

Request

A typical request, in tokens - CloudWatch's InputTokenCount, CacheWriteInputTokenCount and OutputTokenCount show yours. Tokens read from the prompt cache never count against the quotas.

Traffic

The busiest minute you plan for.

Your quotas

AWS's defaults for us-east-1 fill in by themselves. Your account's can be lower - enter the values Service Quotas shows for it.

See the result ↓
Designing for service quotas and scaling is a core AWS Solutions Architect Associate topicTry free SAA-C03 practice questions with answers and explanations.SAA-C03 questions →

How Amazon Bedrock counts a request against the quotas

On the bedrock-runtime endpoint - InvokeModel, Converse and their streaming variants - every model has a quota of tokens per minute (input and output together) and, for some models, of requests per minute, in each Region. A request counts twice:

  1. When it starts, Bedrock deducts its input tokens plus max_tokens. If that is more than is left of the minute's quota, the request is throttled with a 429.
  2. While it runs and when it ends, the deduction is adjusted to what the request really used: input tokens + cache write tokens + output tokens × the model's burndown rate. The unused part of max_tokens goes back to the quota.

Tokens read from the prompt cache never count. You are billed for the tokens alone, not the burndown: AWS's example is a request to Claude Sonnet 4 with 1,000 input and 100 output tokens, which uses 1,500 tokens of the quota and is billed for 1,100. AWS lists the burndown rates on its "How tokens are counted" page:

  • The burndown rate for Anthropic Claude models version 4.8 is 15x for output tokens.
  • The burndown rate for Anthropic Claude Opus 5.5, Claude Sonnet 5, Claude Opus 5, and Claude Fable 5.1 is 10x for output tokens.
  • For all other Anthropic models version 4.7 and below, the burndown is 5x for output tokens.
  • The burndown rate for OpenAI GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna on the bedrock-runtime endpoint is 10x for output tokens.
  • For all other models, the burndown rate is 1:1.

Models AWS's page does not name yet (newer Claude models) get the highest Anthropic rate in the calculator, so the result errs on the safe side - enter the rate yourself if AWS names it. AWS's Global cross-Region inference page still has an older, shorter list; the calculator follows the dedicated page.

Which quotas apply

QuotaScopeCounts
On-demand tokens / requests per minutePer model and RegionCalls with the model ID in the Region
Cross-Region tokens / requests per minutePer model and source RegionCalls with a Geo inference profile (us., eu., ...)
Global cross-Region tokens / requests per minutePer model and source RegionCalls with a Global inference profile (global.)
Cross-Model Max Tokens Per DayPer account and Region, all modelsEvery model's tokens, estimated from on-demand prices
Model tokens per daySome models, per RegionThe model's tokens; doubled for cross-Region calls, dropped once AWS approves a tokens-per-minute increase
bedrock-mantle input / output tokens per minutePer model and RegionThe OpenAI- and Anthropic-compatible APIs; separate from bedrock-runtime

The bedrock-mantle endpoint counts differently: input tokens plus max_tokens are checked against its input quota when a request arrives (the unused part returns at the end), output tokens count against a separate output quota - and when that runs out, generation stops and the response ends early instead of a 429. It has no requests-per-minute quota and no burndown, and AWS publishes its quotas only for some models; for the others throughput depends on the service's capacity.

Concurrency, and why max_tokens matters

Because max_tokens is deducted in full when a request starts, it decides how many requests can run at once, not the tokens you really use. AWS's own example: 3,000 input tokens, 1,000 cache write tokens and 1,000 output tokens with a 5x burndown count 9,000 tokens in the end - but with max_tokens at 32,000 the request first reserves 36,000. Lowered to 1,250, it reserves 5,250, and about seven times as many requests fit in the same quota at the same moment. The calculator shows both numbers: requests per minute (limited by what requests count in the end) and requests at once (limited by what they reserve at the start).

Why you get a 429 at low traffic

  • Your account's quota is lower than the default. AWS adjusts an account's defaults by the Region, payment history and fraudulent usage, and new accounts may get reduced quotas. Look up the values applied to your account in Service Quotas, in the Region you call.
  • One request reserves more than the quota. A large max_tokens with a long input can exceed a small tokens-per-minute quota on its own.
  • The burndown. An output token of many Claude models uses 5, 10 or 15 tokens of the quota.
  • The wrong quota. On-demand, Geo, Global and bedrock-mantle quotas are separate: raising one does not help calls made the other way.
  • The daily quota. Cross-Model Max Tokens Per Day is shared by every model and application in the account and Region.
  • It is not a quota at all. ServiceUnavailableException (503) and overloaded_error (529) mean Bedrock has no capacity at the moment - the Bedrock model finder's "Paste an error" mode explains them.

The Why 429 at low traffic? mode above checks these against your workload and quotas.

Frequently asked questions

How do I request a quota increase?

For bedrock-runtime, request an increase of the model's "Cross-Region InvokeModel tokens per minute" quota in Service Quotas; AWS Support then offers to raise the on-demand tokens per minute and the daily quota too. AWS gives priority to accounts whose traffic already uses up their quota, and grants no increases for Legacy models. For bedrock-mantle, open a support case with the endpoint, Region, model, quota and value.

Does a cross-Region inference profile raise my quota?

It uses a different quota: a Geo or Global profile counts against the model's cross-Region quotas, which are often higher than the on-demand ones. The calculator shows the default of each way the model can be called from your Region.

Where does the data come from?

The default quotas come from the Amazon Bedrock quota table of the AWS General Reference, and the models, their IDs and Regions from AWS's model cards - both read again on every build of this site, the quotas last changed on 2026-09-29. The counting rules and burndown rates are those of the Bedrock User Guide's quota pages, and every build checks they are still there. Some models have no quota in AWS's table; the calculator then asks for your value.

Is what I enter sent anywhere?

No. The calculation runs in your browser. The page only counts that the calculator was used, with the way of calling the model and the first cause the diagnosis finds, never the numbers.

References

How tokens are counted in Amazon Bedrock (Amazon Bedrock User Guide)
Quotas for the bedrock-runtime endpoint (Amazon Bedrock User Guide)
Quotas for the bedrock-mantle endpoint (Amazon Bedrock User Guide)
Quotas for Amazon Bedrock (Amazon Bedrock User Guide)
Amazon Bedrock service quotas (AWS General Reference)
Troubleshooting Amazon Bedrock API error codes (Amazon Bedrock User Guide)
list-service-quotas (AWS CLI Command Reference)