Skip to content
← Back to blog
OpenAIArticle

OpenAI vs Claude API: model prices and 10,000 tasks costed

Current OpenAI and Anthropic model rates, a shared worked example, and the costs of caching, retries and human review.

Author

Syntalith

Published Updated 3 min read

API cost comes from billed token counts and model rates. The cost of a usable result also depends on attempts, tools and correction effort. Put both figures beside each other before choosing a provider.

We checked the following rates on 30 September 2026 against official pricing. Calculations cover the providers' direct APIs, standard text processing and short context. Amounts are in USD before tax. Cloud reseller offers, regional processing and other modes can have different prices.

Rates per million tokens

A token is a unit into which a model splits text. Input includes instructions, questions, history and supplied material. Output covers billed generation tokens under the model's rules. Tokens do not map to a fixed number of words or pages.

ModelUncached input per 1MOutput per 1M
OpenAI GPT-6 Luna$0.10$0.50
OpenAI GPT-6.1 Sol$2$10
OpenAI GPT-6 Astra$10$50
Claude Haiku 4.5$1$5
Claude Sonnet 5.5$2$10
Claude Opus 5.5$4$20
Claude Fable 5.1$10$50

Sources: OpenAI API pricing and Anthropic API pricing. This table compares rates without ranking quality or speed. Long context, caching, Batch and additional tools require the relevant pricing entries.

One shared example: 10,000 monthly calls

Assume a classification task with a short result: each call uses 2,000 input tokens and 500 billable output tokens. Each of 10,000 tasks needs one attempt. Monthly volume is 20 million input and 5 million output tokens.

Cost = 20 × input rate + 5 × output rate.

ModelInputOutputMonthly total
GPT-6 Luna$2$2.50$4.50
GPT-6.1 Sol$40$50$90
GPT-6 Astra$200$250$450
Claude Haiku 4.5$20$25$45
Claude Sonnet 5.5$40$50$90
Claude Opus 5.5$80$100$180
Claude Fable 5.1$200$250$450

This is a calculation using identical assumed usage. The same document can produce different token counts across providers, and response lengths can differ. Replace assumptions with actual usage counters during a pilot. The totals exclude tools, storage, retries and human review.

What changes the bill

Growing history is easy to overlook. Sending the entire conversation and documents at each step creates repeated input charges according to the provider's billing and cache rules. An agent may take many such steps to handle one case.

Caching can reduce the cost of reused context, subject to provider requirements. Writes, reads and storage duration can have separate rates. Do not apply the cache-read price to all traffic without measuring cache hits.

Batch processing can suit work that does not need an immediate response. Match it to the waiting time the process allows. Calculate paid search, file storage, other tools and application infrastructure separately.

Using a cheaper model before a more expensive one

A workflow can start with a cheaper model and send specified cases to a more capable one. That needs a reliable way to detect an incorrect or incomplete result. A model's statement of confidence is insufficient on its own.

If the first attempt costs $0.004 and the second $0.02, a case sent onward costs $0.024 for both calls. These illustrative rates explain the calculation. Savings depend on the share of escalated cases and the quality of the escalation rule. Check cases that should have been escalated but were incorrectly accepted.

Why review effort can decide the purchase

Illustrative scenario: 20% of 10,000 cases need three minutes of review. That is 10,000 × 20% × 3 / 60 = 100 hours. At an assumed employment cost of €30 per hour, review costs €3,000 monthly. Substitute your own hourly cost and use an explicit exchange rate before combining it with USD charges.

A model that reduces corrections may justify higher token rates. Measure that effect. Record the cost of all attempts, review time and accepted cases:

Cost per accepted case = (API + tools + review + corrections) / accepted cases.

Add build, hosting and maintenance costs to the full workflow budget. Failed attempts still consume money and belong in the numerator.

Choosing a model for deployment

Run the shortlisted models on the same typical cases and exceptions. Define correctness before scoring, record model versions and settings, and measure response time, cost and correction effort. Then check performance on a fresh batch of cases.

Token rates help with an initial estimate. Base the launch decision on the quality and cost of accepted work. For a comparison of the ready-made products, see Claude or ChatGPT for business. For build scope, see OpenAI API implementation.

Syntalith implements OpenAI and Claude solutions. A process audit starts at €600 net; our pricing explains the scope and credit toward a subsequent build. Talk to us about your volume and how you assess results.

OpenAI Select Partner

Syntalith is an OpenAI Select Partner in the OpenAI Partner Network.

We help your team respond to customers faster and find information in company documents. We choose and set up the right AI tools, then teach your team how to use them.

How to introduce ChatGPT and OpenAI at work
Syntalith is a member of Claude Partner Network, Anthropic's partner program.

Denotes membership in Anthropic's partner program for Claude. Not an endorsement of Syntalith's services by Anthropic.

Free process scan

Start with a free process scan.

  • A 30-minute call with the engineer who would lead the work.
  • A review of the processes that cost you the most time and money.
  • A written summary: a possible direction, missing information and the next step.

The scan chooses one process to assess, and within 2 business days you receive a recommendation, including when a simpler route is the better fit.

€0

30 minutes · written takeaway within 2 business days

Book a free process scan (30 min)

Times are shown in your own time zone. We work with clients across time zones.

Describe the process in the form