Claude API Pricing 2026: Token Costs and a Real Budget
Claude API rates in August 2026 range from USD 1 to 10 per million input tokens and from USD 5 to 50 per million output tokens. Use the official table, a case-based formula, and current Syntalith build prices to decide whether an API integration fits your process.
Claude API billing follows two measurements: tokens the model receives and tokens it produces. A case-based estimate gives a more useful budget than a headline rate.
Syntalith
Claude API is a sensible purchase when a repeatable business process needs model output inside an existing system. The budget has two parts: the provider charge for input and output tokens, plus the work required to make the result reliable in your process.
Anthropic's official rate table, checked on 21 August 2026, lists USD 1 to 10 per million input tokens and USD 5 to 50 per million output tokens. The table and the calculation below are enough to decide whether to test the API for one process.
Current official model rates
Anthropic quotes rates per million tokens (MTok). Input covers the content sent to the model; output covers the answer it generates.
| Model | Input per MTok | Output per MTok | Context window | Typical use |
|---|---|---|---|---|
| Claude Fable 5 | 10 USD | 50 USD | 1M tokens | Difficult reasoning and long analysis |
| Claude Opus 5 | 5 USD | 25 USD | 1M tokens | Complex cases and multi-step work |
| Claude Sonnet 5 | 2 USD | 10 USD | 1M tokens | Routine production work and document handling |
| Claude Haiku 4.5 | 1 USD | 5 USD | 200K tokens | Classification, routing, and short extraction |
These are the current standard rates on the Anthropic pricing page. Anthropic says the introductory Sonnet 5 price of 2 / 10 USD is now standard, so a September increase described in older articles no longer belongs in a current budget. Recheck the source when the model, account, or billing route changes.
The rate alone cannot predict the invoice. A short classification may use a few hundred output tokens; a long report can use thousands. The same process can also use Haiku for routine cases and a stronger model for exceptions.
Calculate one real case
Measure one case type before multiplying anything. Take a representative sample from the last few weeks, record input and output tokens from the API response, and calculate:
Monthly usage in USD =
monthly cases x average input tokens / 1,000,000 x input rate
+ monthly cases x average output tokens / 1,000,000 x output rate
For example, a process handling 2,000 cases each month with 4,000 input tokens and 600 output tokens per case would use:
2,000 x 4,000 / 1,000,000 x input rate
+ 2,000 x 600 / 1,000,000 x output rate
Replace both averages with measurements from your own cases. Include the system instructions, conversation history, attached documents, and each model call made during one case. A multi-step process pays for every call.
The accepted-case cost is more useful than a token forecast on its own:
Cost per accepted case =
provider usage
+ hosting and integrations
+ review time
+ maintenance
That number lets a process owner compare the API with the staff time and error handling already attached to the task.
Three ways to reduce the usage bill
- Assign a model to each task. Use Haiku for classification and routing, Sonnet for routine work, and Opus or Fable for cases whose quality test justifies the extra rate. Record the rule so the choice stays predictable.
- Cache repeated context. Anthropic's prompt caching documentation lists a cache hit at 0.1 times the base input rate. A stable instruction pack, catalogue, or reference set is a good candidate. Measure the hit rate before including the saving in a business case.
- Use batch processing for deferred work. Anthropic's Batch API documentation describes a 50% discount on eligible input and output rates. It suits overnight document processing and enrichment where an immediate answer is unnecessary.
Set a monthly spend limit and write the response to a limit event before launch. The options are a queue, a cheaper model for routine cases, or a human handoff. A limit without an operating response only moves the surprise to a later invoice.
Direct API, team plan, or an integration
The API fits a process that has regular volume, an owner, and a destination for the result. A team subscription fits people who need Claude in their daily work and can keep the work in the application. A custom integration earns its cost when the result must reach CRM, ERP, a shared inbox, a document store, or another controlled system.
If the task runs once or a few times a month, start in the application. If the same person repeats the same steps every week, measure one case and compare the accepted-case cost with the time spent today.
Build cost alongside usage
Token charges are only one line in the decision. Current Syntalith entry prices are:
| Work | Price | Includes |
|---|---|---|
| Implementation specification | from €1,200 net | Process map, technical plan, and a fixed build quote |
| Process automation | from €3,500 net | A repeatable process connected to the required systems |
| Process-running agent or app | from €6,000 net | Multi-step work across systems; larger builds typically reach €6,000 to €35,000 net |
The Syntalith pricing page contains the current offer. A specification is useful when the team needs a fixed decision before commissioning a build. The AI cost-management playbook covers the wider budget, including provider usage and accepted outcomes.
A practical buying sequence
- Name one repeatable process and its monthly volume.
- Measure input and output tokens on representative cases.
- Run a quality test with the people who accept the result.
- Compare accepted-case cost with current staff time, corrections, and delays.
- Set a spend limit, an owner, and a response when the limit is reached.
The Claude API price list remains the source for provider rates. A free process scan can help decide whether the API, a team plan, or a small integration fits the process.
Questions for a buyer
Which Claude model is cheapest? Haiku 4.5 has the lowest current standard rate at 1 USD per million input tokens and 5 USD per million output tokens. Use it after a quality test confirms that the task fits its output.
When does prompt caching help? It helps when a substantial part of the input repeats across many calls. The provider documentation describes cache hits at 0.1 times the base input rate. Count cache writes and misses in the estimate.
Does the API bill include the whole system? The API charge covers model usage. Hosting, integrations, access controls, review, maintenance, and any provider route such as a cloud marketplace belong in the wider budget.
Related articles
Frequently asked questions
- How much does a Claude token cost?
- Anthropic quotes Claude API rates per million input and output tokens. In August 2026 Haiku 4.5 is USD 1 / 5, Sonnet 5 is USD 2 / 10, Opus 5 is USD 5 / 25, and Fable 5 is USD 10 / 50. Check the official price page before signing a budget.
- How do I estimate a Claude API bill?
- Measure input and output tokens on a representative set of real cases. Multiply each average by monthly volume and its rate, then add the cost of the integration, hosting, review, and maintenance.
- Can prompt caching lower Claude API cost?
- Yes. A cache hit is billed at 0.1 times the base input rate. Caching helps when the same instructions, reference material, or examples accompany many requests. Batch processing can halve eligible input and output rates.
Free process scan
Start with a free process scan.
- A 30-minute call with the engineer who would lead the work.
- A review of the processes that cost you the most time and money.
- A written summary of what to automate first and the likely cost range.
The scan chooses one process to assess, and within 2 business days you receive a recommendation, including when a simpler route is the better fit.
€0
30 minutes · written takeaway within 2 business days
Times are shown in your own time zone. We work with clients across time zones.
Describe the process in the form