Claude API Pricing: How Token Costs Work and How to Keep Your Bill Low

Most founders aren't surprised by the Claude API price list. They are surprised by their own traffic. A prototype costs cents a day, then it goes live, a long system prompt rides along with every request, and the bill is ten times the plan.
This guide shows you how Claude API pricing works and how to keep it under control: official per model prices, the cost saving features Anthropic documents, a worked cost example, five steps to estimate and cap your spend, and the prompts we use to plan a budget.
The short version:
- The Claude API charges per token, and output tokens cost five times as much as input tokens on every current model.
- As of September 2026, prices run from $1 input and $5 output per million tokens on Claude Haiku 4.5 up to $10 and $50 on Claude Fable 5.1.
- Prompt caching turns repeated context into cheap cache reads, which cost 2.5% to 10% of the normal input price depending on the model.
- The Batch API cuts both input and output prices by 50% for work that doesn't need an instant answer, and it stacks with caching.
- Spend limits in the Claude Console let you cap your monthly API bill below your tier's maximum.
What is Claude API pricing?
Claude API pricing is the usage based fee Anthropic charges developers for calling Claude models through the API. You pay per million tokens (MTok), with separate rates for input tokens (what you send), output tokens (what Claude writes back) and cached tokens, and the rate depends on the model you choose. There is no monthly subscription for the API itself: according to the Claude Help Center, usage is paid with prepaid credits you buy in the Claude Console.
A token is a small piece of text: Anthropic's pricing FAQ puts it at about 4 characters or 0.75 words in English. Every request is billed on the tokens you send (system prompt, history, documents, tool definitions), the tokens Claude generates, and extras such as web searches. If you are new to the API itself, start with our guide to the Claude API, then come back here to plan the budget.
The core: what each Claude model costs per token

These are the current models and their standard prices per million tokens, taken from Anthropic's official pricing page and claude.com/pricing, as of September 2026:
| Model | Input | Output | 5 min cache write | Cache hit | Batch input | Batch output |
|---|---|---|---|---|---|---|
| Claude Fable 5.1 | $10 | $50 | $12.50 | $0.25 | $5 | $25 |
| Claude Opus 5.5 | $4 | $20 | $5 | $0.20 | $2 | $10 |
| Claude Sonnet 5.5 | $2 | $10 | $2.50 | $0.20 | $1 | $5 |
| Claude Haiku 4.5 | $1 | $5 | $1.25 | $0.10 | $0.50 | $2.50 |
Prices change, so always check the official pricing page before you commit to a budget.
Three things matter more than the headline numbers:
- Output is the expensive side. On every model, one output token costs five input tokens. Long answers cost far more than long prompts.
- Cache hits are almost free. A cache hit costs 10% of the base input price on most models, 5% on Claude Opus 5.5 and 2.5% on Claude Fable 5.1, per the pricing page.
- Model choice is the biggest lever. Anthropic's models overview positions Fable 5.1 for demanding reasoning and long horizon agent work, Opus 5.5 for long running agentic coding and knowledge work, Sonnet 5.5 as the best mix of speed and intelligence, and Haiku 4.5 as the fastest. Our Claude models guide compares them.
Extras are billed on top: web search costs $10 per 1,000 searches plus the tokens it pulls in, while web fetch has no extra charge beyond tokens. US only inference via the inference_geo setting adds a 1.1x multiplier (pricing page). App subscriptions are billed separately: see Claude pricing.
A worked cost example
Say you run a support assistant on Claude Sonnet 5.5. Each request sends a 5,000 token system prompt plus a 500 token question, and Claude answers in about 400 tokens. Volume: 10,000 requests a month.
Without caching:
- Input: 10,000 × 5,500 tokens = 55M tokens × $2 = $110
- Output: 10,000 × 400 tokens = 4M tokens × $10 = $40
- Total: $150 per month
With prompt caching on the 5,000 token system prompt, assuming the cache expires often enough that 1,000 requests have to write it again and 9,000 read it:
- Cache writes: 1,000 × 5,000 = 5M tokens × $2.50 = $12.50
- Cache reads: 9,000 × 5,000 = 45M tokens × $0.20 = $9
- Uncached input: 10,000 × 500 = 5M tokens × $2 = $10
- Output: $40
- Total: $71.50 per month, about 52% less
Output is now more than half the bill, so the next savings come from shorter answers or from testing Claude Haiku 4.5 on the easy questions. Anthropic's own pricing page example puts 10,000 support tickets on Haiku 4.5 at roughly $37, based on about 3,700 tokens per conversation. Our numbers are illustrative, calculated from official rates.
How to estimate and control Claude API costs in 5 steps

Step 1: Count your real tokens
Don't guess. Anthropic's token counting endpoint returns the input token count of a request before you send it, and it is free to use (it has its own rate limits). Run your real system prompt, a typical user message and your tool definitions through it with the model you plan to use. Anthropic notes that its newer tokenizer produces roughly 30% more tokens for the same text than older ones, so recount whenever you switch models.
Step 2: Pick the cheapest model that passes your test
Run 20 to 50 real examples of the task through two models. If Claude Haiku 4.5 or Claude Sonnet 5.5 gets them right, don't pay Opus or Fable prices for that step. Anthropic's own cost guidance says the same: Haiku for simple tasks, Sonnet for most production workloads, Opus for the most complex reasoning. Route only the hard cases to the expensive model.
Step 3: Cache everything that repeats
System prompts, knowledge documents, tool definitions and long conversation histories are the same on every call. Put them at the start of the prompt and turn on prompt caching. The simplest option per the pricing page is automatic caching with one cache_control field. A 5 minute cache write costs 1.25x the input price and pays off after one cache read; a 1 hour write costs 2x and pays off after two reads. Cached reads also don't count toward your input tokens per minute limit on current models (rate limits), so caching raises your throughput too.
Step 4: Batch everything that can wait
Nightly reports, lead enrichment, tagging and evaluations don't need an answer in two seconds. The Message Batches API charges 50% less on input and output, and Anthropic says most batches finish in less than an hour. Batch and cache discounts stack, according to the pricing page.
Step 5: Cap the bill and watch it weekly
In the Claude Console, go to Settings, then Billing, and set your own spend limit below your tier's cap. As of September 2026, the rate limits page lists monthly caps of $500 on the Start tier, $1,000 on Build and $200,000 on Scale. Workspaces can get their own, lower limits. Then check the Usage page weekly: its charts show tokens, requests and your cache hit rate.
A real case: how Hume cut costs 80% with prompt caching

Hume AI builds emotionally intelligent voice AI. According to Anthropic's customer story, it was founded in 2021 by Alan Cowen, and its flagship product EVI is a voice to voice platform used in healthcare, customer service, tutoring and more. Claude powers EVI conversations through the API.
The facts from the story:
- EVI has handled over 2 million minutes of AI voice conversations across more than 1 million distinct conversations.
- The average conversation lasts about 3 minutes, and many run past 30 minutes.
- 36% of EVI users choose Claude over the other external language models on offer.
- Prompt caching helped Hume reduce costs by 80% and lower latency by 10% or more (Hume customer story).
When prompt caching launched, Anthropic reported a 53% cost reduction for a 10 turn conversation with a long system prompt, and up to 90% for very long cached prompts (Anthropic).
Look at it through the system:
- Count first: long voice conversations mean the same system prompt and a growing history are sent again and again. That is exactly the repeated input caching is built for.
- Right model for the job: EVI offers several models, and many users pick Claude for conversation quality.
- Cache what repeats: the savings came from not paying full price for the same context on every turn.
The lesson: the biggest cost cut usually doesn't come from a cheaper vendor. It comes from not paying full price for context you have already sent.
Three use cases
The examples below are illustrative, not real clients. They show how the same five steps play out in different businesses.
Use case 1: Agency writing client reports
Before: every report runs live through the most expensive model, resending the same 8,000 token brand guide.
After: the brand guide is cached, drafts run on Sonnet 5.5 through the Batch API overnight, and only the final strategy summary goes to Opus 5.5.
Use case 2: SaaS with an in app assistant
Before: long answers, no caching, and costs grow faster than revenue.
After: a cached product knowledge prompt, a cap on answer length, Haiku 4.5 for simple how to questions and Sonnet 5.5 for the rest. A workspace spend limit protects the production budget.
Use case 3: Solo founder enriching leads
Before: a script sends 5,000 lead profiles one at a time through the synchronous API.
After: one batch at half price, the scoring rubric cached, and a spend limit set before the first run. For what to do with the scored leads, see how to get more leads.
How we run this with Claude
Before we ship any API feature, we plan the cost in a Claude chat. Here are the two prompts. Fill in the brackets and paste today's prices from the official pricing page, so Claude calculates with current numbers.
Prompt 1: estimate the monthly bill before you build
Make it yours · 0/3 filled
You are a cost analyst for the Claude API. Here are the current official prices per million tokens: [PASTE PRICE TABLE FROM THE PRICING PAGE]. My feature: [WHAT IT DOES]. Per request, I send about [X] tokens of fixed context (system prompt, documents, tools) and [Y] tokens of variable input, and Claude answers with about [Z] tokens. Expected volume: [REQUESTS PER MONTH]. Calculate the monthly cost for each model in the table, then again with prompt caching on the fixed context and with the Batch API where the task allows it. Show every calculation step and list the assumptions I should verify with the token counting endpoint.Prompt 2: find the savings in an existing workload
Make it yours · 0/3 filled
Here is my Claude API usage for the last [PERIOD]: [MODEL, INPUT TOKENS, OUTPUT TOKENS, CACHE READS, CACHE WRITES, REQUESTS]. Here is my system prompt: [PASTE PROMPT]. Here are the tasks this workload handles: [LIST OF TASKS]. Find the three biggest savings. For each one, tell me whether it comes from model choice, caching, batching or shorter output, estimate the saving with the official prices I pasted above, and tell me the one test I should run before changing anything in production.These prompts give you a budget before you build and a savings list once you are live. Deciding which tasks can move to a cheaper model without hurting quality is where most people get stuck, because that takes testing on your own data.
Where most people get stuck
Reading a price table is easy. Keeping the bill flat while usage grows is not. The usual traps:
- They estimate with words instead of tokens. The real count, with tool definitions and history, is often much higher than the guess.
- They default to the biggest model everywhere. One expensive model on every step, when most steps would pass on a cheaper one.
- They never set a spend limit. The first time they notice runaway usage is when the invoice arrives.
That is exactly the gap the Inner Circle is built for: the playbooks to plan and run your Claude API setup, a new playbook every week, and founders building their own AI systems next to you.