Resources·10 min read·By ·

Claude API Pricing: How Token Costs Work and How to Keep Your Bill Low

A glowing lime meter filling up token by token on a dark desk: every API call adds to the bill

Most founders aren't surprised by the Claude API price list. They are surprised by their own traffic. A prototype costs cents a day, then it goes live, a long system prompt rides along with every request, and the bill is ten times the plan.

This guide shows you how Claude API pricing works and how to keep it under control: official per model prices, the cost saving features Anthropic documents, a worked cost example, five steps to estimate and cap your spend, and the prompts we use to plan a budget.

The short version:

  • The Claude API charges per token, and output tokens cost five times as much as input tokens on every current model.
  • As of September 2026, prices run from $1 input and $5 output per million tokens on Claude Haiku 4.5 up to $10 and $50 on Claude Fable 5.1.
  • Prompt caching turns repeated context into cheap cache reads, which cost 2.5% to 10% of the normal input price depending on the model.
  • The Batch API cuts both input and output prices by 50% for work that doesn't need an instant answer, and it stacks with caching.
  • Spend limits in the Claude Console let you cap your monthly API bill below your tier's maximum.

What is Claude API pricing?

Claude API pricing is the usage based fee Anthropic charges developers for calling Claude models through the API. You pay per million tokens (MTok), with separate rates for input tokens (what you send), output tokens (what Claude writes back) and cached tokens, and the rate depends on the model you choose. There is no monthly subscription for the API itself: according to the Claude Help Center, usage is paid with prepaid credits you buy in the Claude Console.

A token is a small piece of text: Anthropic's pricing FAQ puts it at about 4 characters or 0.75 words in English. Every request is billed on the tokens you send (system prompt, history, documents, tool definitions), the tokens Claude generates, and extras such as web searches. If you are new to the API itself, start with our guide to the Claude API, then come back here to plan the budget.

The core: what each Claude model costs per token

Four price columns rising from small to large, each split into a thin input bar and a tall output bar

These are the current models and their standard prices per million tokens, taken from Anthropic's official pricing page and claude.com/pricing, as of September 2026:

ModelInputOutput5 min cache writeCache hitBatch inputBatch output
Claude Fable 5.1$10$50$12.50$0.25$5$25
Claude Opus 5.5$4$20$5$0.20$2$10
Claude Sonnet 5.5$2$10$2.50$0.20$1$5
Claude Haiku 4.5$1$5$1.25$0.10$0.50$2.50

Prices change, so always check the official pricing page before you commit to a budget.

Three things matter more than the headline numbers:

  • Output is the expensive side. On every model, one output token costs five input tokens. Long answers cost far more than long prompts.
  • Cache hits are almost free. A cache hit costs 10% of the base input price on most models, 5% on Claude Opus 5.5 and 2.5% on Claude Fable 5.1, per the pricing page.
  • Model choice is the biggest lever. Anthropic's models overview positions Fable 5.1 for demanding reasoning and long horizon agent work, Opus 5.5 for long running agentic coding and knowledge work, Sonnet 5.5 as the best mix of speed and intelligence, and Haiku 4.5 as the fastest. Our Claude models guide compares them.

Extras are billed on top: web search costs $10 per 1,000 searches plus the tokens it pulls in, while web fetch has no extra charge beyond tokens. US only inference via the inference_geo setting adds a 1.1x multiplier (pricing page). App subscriptions are billed separately: see Claude pricing.

A worked cost example

Say you run a support assistant on Claude Sonnet 5.5. Each request sends a 5,000 token system prompt plus a 500 token question, and Claude answers in about 400 tokens. Volume: 10,000 requests a month.

Without caching:

  • Input: 10,000 × 5,500 tokens = 55M tokens × $2 = $110
  • Output: 10,000 × 400 tokens = 4M tokens × $10 = $40
  • Total: $150 per month

With prompt caching on the 5,000 token system prompt, assuming the cache expires often enough that 1,000 requests have to write it again and 9,000 read it:

  • Cache writes: 1,000 × 5,000 = 5M tokens × $2.50 = $12.50
  • Cache reads: 9,000 × 5,000 = 45M tokens × $0.20 = $9
  • Uncached input: 10,000 × 500 = 5M tokens × $2 = $10
  • Output: $40
  • Total: $71.50 per month, about 52% less

Output is now more than half the bill, so the next savings come from shorter answers or from testing Claude Haiku 4.5 on the easy questions. Anthropic's own pricing page example puts 10,000 support tickets on Haiku 4.5 at roughly $37, based on about 3,700 tokens per conversation. Our numbers are illustrative, calculated from official rates.

How to estimate and control Claude API costs in 5 steps

Five steps in a row: count, choose, cache, batch, cap

Step 1: Count your real tokens

Don't guess. Anthropic's token counting endpoint returns the input token count of a request before you send it, and it is free to use (it has its own rate limits). Run your real system prompt, a typical user message and your tool definitions through it with the model you plan to use. Anthropic notes that its newer tokenizer produces roughly 30% more tokens for the same text than older ones, so recount whenever you switch models.

Step 2: Pick the cheapest model that passes your test

Run 20 to 50 real examples of the task through two models. If Claude Haiku 4.5 or Claude Sonnet 5.5 gets them right, don't pay Opus or Fable prices for that step. Anthropic's own cost guidance says the same: Haiku for simple tasks, Sonnet for most production workloads, Opus for the most complex reasoning. Route only the hard cases to the expensive model.

Step 3: Cache everything that repeats

System prompts, knowledge documents, tool definitions and long conversation histories are the same on every call. Put them at the start of the prompt and turn on prompt caching. The simplest option per the pricing page is automatic caching with one cache_control field. A 5 minute cache write costs 1.25x the input price and pays off after one cache read; a 1 hour write costs 2x and pays off after two reads. Cached reads also don't count toward your input tokens per minute limit on current models (rate limits), so caching raises your throughput too.

Step 4: Batch everything that can wait

Nightly reports, lead enrichment, tagging and evaluations don't need an answer in two seconds. The Message Batches API charges 50% less on input and output, and Anthropic says most batches finish in less than an hour. Batch and cache discounts stack, according to the pricing page.

Step 5: Cap the bill and watch it weekly

In the Claude Console, go to Settings, then Billing, and set your own spend limit below your tier's cap. As of September 2026, the rate limits page lists monthly caps of $500 on the Start tier, $1,000 on Build and $200,000 on Scale. Workspaces can get their own, lower limits. Then check the Usage page weekly: its charts show tokens, requests and your cache hit rate.

A real case: how Hume cut costs 80% with prompt caching

A voice waveform in lime passing through a glowing cache block and coming out lighter on the other side

Hume AI builds emotionally intelligent voice AI. According to Anthropic's customer story, it was founded in 2021 by Alan Cowen, and its flagship product EVI is a voice to voice platform used in healthcare, customer service, tutoring and more. Claude powers EVI conversations through the API.

The facts from the story:

  • EVI has handled over 2 million minutes of AI voice conversations across more than 1 million distinct conversations.
  • The average conversation lasts about 3 minutes, and many run past 30 minutes.
  • 36% of EVI users choose Claude over the other external language models on offer.
  • Prompt caching helped Hume reduce costs by 80% and lower latency by 10% or more (Hume customer story).

When prompt caching launched, Anthropic reported a 53% cost reduction for a 10 turn conversation with a long system prompt, and up to 90% for very long cached prompts (Anthropic).

Look at it through the system:

  • Count first: long voice conversations mean the same system prompt and a growing history are sent again and again. That is exactly the repeated input caching is built for.
  • Right model for the job: EVI offers several models, and many users pick Claude for conversation quality.
  • Cache what repeats: the savings came from not paying full price for the same context on every turn.

The lesson: the biggest cost cut usually doesn't come from a cheaper vendor. It comes from not paying full price for context you have already sent.

Three use cases

The examples below are illustrative, not real clients. They show how the same five steps play out in different businesses.

Use case 1: Agency writing client reports

Before: every report runs live through the most expensive model, resending the same 8,000 token brand guide.

After: the brand guide is cached, drafts run on Sonnet 5.5 through the Batch API overnight, and only the final strategy summary goes to Opus 5.5.

Use case 2: SaaS with an in app assistant

Before: long answers, no caching, and costs grow faster than revenue.

After: a cached product knowledge prompt, a cap on answer length, Haiku 4.5 for simple how to questions and Sonnet 5.5 for the rest. A workspace spend limit protects the production budget.

Use case 3: Solo founder enriching leads

Before: a script sends 5,000 lead profiles one at a time through the synchronous API.

After: one batch at half price, the scoring rubric cached, and a spend limit set before the first run. For what to do with the scored leads, see how to get more leads.

How we run this with Claude

Before we ship any API feature, we plan the cost in a Claude chat. Here are the two prompts. Fill in the brackets and paste today's prices from the official pricing page, so Claude calculates with current numbers.

Prompt 1: estimate the monthly bill before you build

PROMPT

Make it yours · 0/3 filled

You are a cost analyst for the Claude API. Here are the current official prices per million tokens: [PASTE PRICE TABLE FROM THE PRICING PAGE]. My feature: [WHAT IT DOES]. Per request, I send about [X] tokens of fixed context (system prompt, documents, tools) and [Y] tokens of variable input, and Claude answers with about [Z] tokens. Expected volume: [REQUESTS PER MONTH]. Calculate the monthly cost for each model in the table, then again with prompt caching on the fixed context and with the Batch API where the task allows it. Show every calculation step and list the assumptions I should verify with the token counting endpoint.

Prompt 2: find the savings in an existing workload

PROMPT

Make it yours · 0/3 filled

Here is my Claude API usage for the last [PERIOD]: [MODEL, INPUT TOKENS, OUTPUT TOKENS, CACHE READS, CACHE WRITES, REQUESTS]. Here is my system prompt: [PASTE PROMPT]. Here are the tasks this workload handles: [LIST OF TASKS]. Find the three biggest savings. For each one, tell me whether it comes from model choice, caching, batching or shorter output, estimate the saving with the official prices I pasted above, and tell me the one test I should run before changing anything in production.

These prompts give you a budget before you build and a savings list once you are live. Deciding which tasks can move to a cheaper model without hurting quality is where most people get stuck, because that takes testing on your own data.

Where most people get stuck

Reading a price table is easy. Keeping the bill flat while usage grows is not. The usual traps:

  • They estimate with words instead of tokens. The real count, with tool definitions and history, is often much higher than the guess.
  • They default to the biggest model everywhere. One expensive model on every step, when most steps would pass on a cheaper one.
  • They never set a spend limit. The first time they notice runaway usage is when the invoice arrives.

That is exactly the gap the Inner Circle is built for: the playbooks to plan and run your Claude API setup, a new playbook every week, and founders building their own AI systems next to you.

Frequently asked questions

What is Claude API pricing?

Claude API pricing is the usage based fee Anthropic charges for calling Claude models through the API. You pay per million tokens, with separate rates for input, output and cached tokens, and the rate depends on the model. There is no subscription for the API; usage is paid with prepaid credits in the Claude Console.

How much does the Claude API cost per token?

As of September 2026, Claude Haiku 4.5 costs $1 per million input tokens and $5 per million output tokens, Claude Sonnet 5.5 $2 and $10, Claude Opus 5.5 $4 and $20, and Claude Fable 5.1 $10 and $50. One token is roughly 4 characters or 0.75 words of English. Prices change, so check Anthropic's official pricing page.

How does Claude API billing work?

Claude API usage is billed through prepaid credits that an Admin or Billing user buys in the Claude Console under Settings, Billing, and optional auto reload tops them up below a threshold you set. Credits expire one year after purchase and are non refundable. Failed requests are not charged.

How do I check my Claude API usage?

Open the Usage page in the Claude Console, which shows token and request charts per model, rate limit headroom and your prompt cache hit rate. Your credit balance and spend appear on the Billing page, where you can also set a monthly spend limit below your tier's cap.

How much does the Claude Opus API cost?

As of September 2026, Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens on the Claude API. Cache hits cost $0.20 per million tokens, and the Batch API halves the price to $2 input and $10 output. Check Anthropic's pricing page for current rates.

Knowing it is easy. Running it is the work.

Run the Claude API cost playbook inside the Inner Circle

This breakdown gives you the idea. The Inner Circle gives you the systems to run it on your own business, next to founders who are doing the same.

  • → The full vault: every playbook, prompt pack and system, unlocked
  • → A new copy-paste playbook every week
  • → A community of founders who execute, not just consume
  • → The Money System and the CopyPasteCEO app
Join the Inner Circle →

Not sure yet? Try everything for 7 days for $1 →