How Amazon Bedrock bills fine-tuning
A fine-tuned model costs three things. Training is billed once per job, for the number of tokens in the training data times the number of epochs - 5 million tokens trained for 2 epochs bill 10 million. Every custom model is then stored for a monthly fee until you delete it. And a custom model can only be called once it runs somewhere: on Provisioned Throughput, or - for some models - deployed on demand.
The calculator multiplies AWS's prices for the base model in its Region by your plan: training for every job, storage for every model you keep, and serving for every month, 730 hours each, as AWS bills a month. The prices come from the AWS Price List API, read again on every build of this site; the ones shown were published on 2026-09-30.
Provisioned Throughput or on-demand deployment
| Serving | Billed | Can you stop it? |
|---|---|---|
| On-demand deployment | Per input and output token | Yes - you pay only for requests |
| Provisioned Throughput, no commitment | Per model unit per hour | Yes - delete it at any time; billing runs until you do |
| Provisioned Throughput, 1 or 6 months | Per model unit per hour, less for 6 months | Not before the term ends |
A Provisioned Throughput for a custom model is priced like the base model's, although AWS's price list prices some custom models apart - Claude 3 Haiku's model unit costs more customized than as the base model; the calculator uses the custom model's price where the list has one. On-demand deployment exists for Amazon Nova Micro, Lite, Pro and Nova 2 Lite in US East (N. Virginia) and Llama 3.3 70B Instruct in US West (Oregon), for models customized on or after July 16, 2025; Llama 3.3 70B has no Provisioned Throughput at all. For the other models, a single model unit is the floor of what using a fine-tuned model costs, however few requests it serves - the calculator sets it against the base model's tokens for the same traffic.
Frequently asked questions
Why only these models?
These are the text models AWS lists for fine-tuning on Amazon Bedrock, in the Regions it lists - most in exactly one. Image and embedding models such as Nova Canvas are billed per image and are left out.
What about reinforcement fine-tuning and distillation?
Reinforcement fine-tuning (Nova 2 Lite, gpt-oss-20b, Qwen3 32B) is billed per hour of training, which is not known before the job runs, so the calculator does not price it. Distillation fine-tunes the student model with responses the teacher model generates; when Amazon Bedrock adds its own data synthesis, the teacher's calls are billed at its on-demand rates, and the training data can grow to 15,000 prompt-response pairs.
How many model units do I need?
AWS does not publish how many input and output tokens a model unit processes per minute - your AWS account manager has the figure for each model. One model unit is the least you can buy; enter more if your traffic needs it.
Is what I enter sent anywhere?
No. The calculation runs in your browser. The page only counts that the calculator was used, with the way of serving that came out cheapest, never the numbers.
References
Amazon Bedrock pricing
Customize your model to improve its performance for your use case (Amazon Bedrock User Guide)
Customize a model with fine-tuning in Amazon Bedrock (Amazon Bedrock User Guide)
Increase model invocation capacity with Provisioned Throughput (Amazon Bedrock User Guide)
Deploy a custom model for on-demand inference (Amazon Bedrock User Guide)
AWS Price List (AWS Billing User Guide)