As of 6 October 2026, a million output tokens costs between $0.50 (GPT-6 Luna) and $50 (GPT-6 Astra, Claude Fable 5.1) across the eighteen offers listed below. That is a 1 to 100 spread on the public rate cards of OpenAI, Anthropic, Google, Mistral AI, DeepSeek and xAI. ActuIA collected these prices from the pricing pages of the vendors and of two clouds, Amazon Bedrock and Google Cloud. They are converted into euros at the European Central Bank reference rate of 5 October, the latest published at the time of the survey, i.e. 1 euro for 1.1204 dollars.
Update. Survey carried out on 6 October 2026 between 00:00 and 01:00 UTC, on the pages linked in each row. A price may change between two surveys; the vendor's page is authoritative.
The 6 October 2026 table, in dollars per million tokens
The amounts are standard prices, for short context and on each vendor's global endpoint. Short context ends at 272,000 input tokens at OpenAI and at 200,000 tokens for Gemini 3.1 Pro and Grok 4.7. At xAI, a request whose prompt reaches that threshold is billed at the higher rate "for all tokens in the request", according to the model page. Google indexes its price on prompt size ("prompts > 200k tokens"), i.e. $4 and $18 beyond that for Gemini 3.1 Pro. OpenAI does not specify on its rate card how the switch at 272,000 tokens applies. The "cached input" column gives the price of a token read back from the prompt cache, excluding write fees. The "batch" column refers to deferred processing interfaces (Batch API).
| Vendor | Model | Input | Cached input | Output | Batch (input / output) | Surveyed | Source |
|---|---|---|---|---|---|---|---|
| OpenAI | GPT-6 Astra (gpt-6-astra) | 10.00 | 1.00 | 50.00 | 5.00 / 25.00 | 6 Oct 2026 | rate card |
| OpenAI | GPT-6.1 Sol (gpt-6.1-sol) | 2.00 | 0.10 | 10.00 | 1.00 / 5.00 | 6 Oct 2026 | rate card |
| OpenAI | GPT-6 Luna (gpt-6-luna) | 0.10 | 0.01 | 0.50 | 0.05 / 0.25 | 6 Oct 2026 | rate card |
| Anthropic | Claude Fable 5.1 | 10.00 | 0.25 | 50.00 | 5.00 / 25.00 | 6 Oct 2026 | rate card |
| Anthropic | Claude Opus 5.5 | 4.00 | 0.20 | 20.00 | 2.00 / 10.00 | 6 Oct 2026 | rate card |
| Anthropic | Claude Sonnet 5.5 | 2.00 | 0.20 | 10.00 | 1.00 / 5.00 | 6 Oct 2026 | rate card |
| Anthropic | Claude Haiku 4.5 | 1.00 | 0.10 | 5.00 | 0.50 / 2.50 | 6 Oct 2026 | rate card |
| Gemini 3.1 Pro Preview | 2.00 | 0.20 | 12.00 | 1.00 / 6.00 | 6 Oct 2026 | rate card | |
| Gemini 3.8 Flash (1) | 0.75 | 0.075 | 3.75 | 0.375 / 1.875 | 6 Oct 2026 | rate card | |
| Gemini 3.5 Flash-Lite | 0.30 | 0.03 | 2.50 | 0.15 / 1.25 | 6 Oct 2026 | rate card | |
| Mistral AI | Mistral Medium 3.5 | 1.50 | 0.15 | 7.50 | 0.75 / 3.75 (2) | 6 Oct 2026 | rate card |
| Mistral AI | Mistral Large 3 | 0.50 | 0.05 | 1.50 | 0.25 / 0.75 (2) | 6 Oct 2026 | rate card |
| Mistral AI | Mistral Small 4 | 0.15 | 0.015 | 0.60 | 0.075 / 0.30 (2) | 6 Oct 2026 | rate card |
| DeepSeek | deepseek-flash (V4.1 Flash), peak hours (3) | 0.30 | 0.006 | 1.20 | n/a | 6 Oct 2026 | rate card |
| DeepSeek | deepseek-flash, off-peak hours | 0.15 | 0.003 | 0.60 | n/a | 6 Oct 2026 | rate card |
| DeepSeek | deepseek-v4-pro, peak hours | 1.32 | 0.044 | 3.96 | n/a | 6 Oct 2026 | rate card |
| DeepSeek | deepseek-v4-pro, off-peak hours | 0.66 | 0.022 | 1.98 | n/a | 6 Oct 2026 | rate card |
| xAI | Grok 4.7 | 2.00 | 0.50 | 6.00 | not offered | 6 Oct 2026 | model page |
(1) Introductory price displayed until 31 December 2026; Google announces $1.50 and $7.50 from 1 January 2027. (2) Amount calculated by ActuIA: Mistral AI applies a 50% discount in batch mode, according to its pricing page and its batch processing documentation. (3) Peak hours: 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday. n/a: not disclosed on the page surveyed. For Grok 4.7, the 21 September launch page announces $2 and $6. The model page displays $0.50 for cached input, then $4, $1 and $12 from 200,000 tokens, and states that the batch API is not supported.
The same prices converted into euros
The conversion applies the reference rate published by the European Central Bank on 5 October 2026. Vendors bill in dollars; the exchange rate difference between the survey date and the invoice date is borne by the customer.
| Model | Input (€) | Cached input (€) | Output (€) |
|---|---|---|---|
| GPT-6 Astra | 8.925 | 0.8925 | 44.63 |
| GPT-6.1 Sol | 1.785 | 0.0893 | 8.93 |
| GPT-6 Luna | 0.089 | 0.0089 | 0.45 |
| Claude Fable 5.1 | 8.925 | 0.2231 | 44.63 |
| Claude Opus 5.5 | 3.570 | 0.1785 | 17.85 |
| Claude Sonnet 5.5 | 1.785 | 0.1785 | 8.93 |
| Claude Haiku 4.5 | 0.893 | 0.0893 | 4.46 |
| Gemini 3.1 Pro Preview | 1.785 | 0.1785 | 10.71 |
| Gemini 3.8 Flash | 0.669 | 0.0669 | 3.35 |
| Gemini 3.5 Flash-Lite | 0.268 | 0.0268 | 2.23 |
| Mistral Medium 3.5 | 1.339 | 0.1339 | 6.69 |
| Mistral Large 3 | 0.446 | 0.0446 | 1.34 |
| Mistral Small 4 | 0.134 | 0.0134 | 0.54 |
| deepseek-flash (peak hours) | 0.268 | 0.0054 | 1.07 |
| deepseek-flash (off-peak hours) | 0.134 | 0.0027 | 0.54 |
| deepseek-v4-pro (peak hours) | 1.178 | 0.0393 | 3.53 |
| deepseek-v4-pro (off-peak hours) | 0.589 | 0.0196 | 1.77 |
| Grok 4.7 | 1.785 | 0.4463 | 5.36 |
Rate card changes since 1 September, date by date
- 1 September: Anthropic launches Claude Fable 5.1 and lowers cache reads to $0.25, i.e. 0.025 times the input price, according to its release notes. Input and output stay at $10 and $50.
- 2 September: Google releases Gemini 3.8 Flash at the introductory price of Gemini 3.7 Flash, $0.75 and $3.75.
- 3 September: OpenAI releases GPT-6 Astra, billed at $10 and $50.
- 10 September: DeepSeek launches V4.1 Flash, called in the API under the name deepseek-flash, and lowers its prices accordingly. The former identifiers deepseek-v4-flash and deepseek-v4-flash-vision-exp are "temporarily" routed to V4.1 Flash and billed at its price; the corresponding models are retired.
- 21 September: xAI releases Grok 4.7 at the price of Grok 4.6.
- 22 September: OpenAI launches GPT-6 Sol at $2 and $10, compared with $4 and $20 for the promotional price of GPT-5.6 Sol, and GPT-6 Luna at $0.10 and $0.50. The same day, Anthropic releases Claude Opus 5.5, at $4 and $20 compared with $5 and $25 for Opus 5, and at $0.20 for cache reads instead of $0.50. ActuIA recalculated a typical agent session with these two rate cards in its 25 September analysis.
- 28 September: Claude Sonnet 5.5 is released at the price of Sonnet 5, $2 and $10, with cache reads at $0.20.
- 29 September: GPT-6.1 Sol keeps GPT-6 Sol's $2 and $10 but halves the cache read price, to $0.10, as ActuIA detailed on 1 October.
The OpenAI changelog dates the three launches to 3, 22 and 29 September. One earlier move remains in force. On 10 August, Anthropic made the introductory price of Claude Sonnet 5, $2 and $10, its standard price, according to its release notes; the increase to $3 and $15 planned for 1 September did not take place.
Three of the September launches are primarily about cache reads. Fable 5.1 and GPT-6.1 Sol change only that line, and Opus 5.5 cuts it by 60% while input and output fall by 20%. According to Anthropic, this is the line that makes up the majority of agentic and coding costs.
Cache and batch: 75% to 98% off re-read input
The discount on a token read back from the cache ranges from 75% for Grok 4.7 ($0.50 versus $2) to 98% for deepseek-flash ($0.006 versus $0.30 at peak hours). It reaches 95% on GPT-6.1 Sol and Claude Opus 5.5, 97.5% on Claude Fable 5.1 and 90% at Google and Mistral AI. The cache only applies to the part of a request that is identical from one call to the next: instructions, tool definitions and reference documents must come before the data specific to each call, otherwise the discount does not apply.
The cache also has an entry cost. At Anthropic, a write costs 1.25 times the input price for a five-minute retention and 2 times for one hour. The vendor calculates that it pays for itself from the first read in the first case, and from the second read in the second. OpenAI also charges for writes, at 1.25 times the short-context input price: $12.50 on GPT-6 Astra, $2.50 on GPT-6.1 Sol and $0.125 on GPT-6 Luna. On Gemini 3.8 Flash, Google adds a storage rental of $0.50 per million tokens per hour, rising to $4.50 on Gemini 3.1 Pro Preview: a cache kept without being read is still paid for.
Batch processing halves the bill at OpenAI, Anthropic, Google and Mistral AI, in exchange for a longer response time. Anthropic specifies that its multipliers stack: batch discount, cache and data residency apply on top of one another. The discount disappears in several configurations surveyed on 6 October: on Amazon Bedrock for Claude Opus 5.5 and Sonnet 5.5, on Mistral AI's regional endpoints and at xAI for Grok 4.7. DeepSeek's page mentions no batch offer.
Four cost items missing from the rate cards
The tokenizer
A price per million tokens does not say how many tokens a given text produces. Anthropic states that Claude 4.7 and later models use a new tokenizer that produces "approximately 30% more tokens for the same text", an increase the vendor says varies with the content. At the same listed price, the bill for the same corpus follows it. With its Kolibri model released on 3 October, Aleph Alpha publishes a compression measurement on a German web corpus: 4.90 for its bilingual tokenizer, 4.35 for GPT-5's and 3.72 for DeepSeek V4's. The same German text thus yields about 30% more tokens at DeepSeek than at Kolibri. ActuIA found no French-specific coefficient on the rate cards surveyed. It can be measured on one's own documents, using the token counting endpoints documented by OpenAI and Anthropic.
Data residency
OpenAI charges 10% more on its regional endpoints, including eu.api.openai.com, for models released from 5 March 2026. Access goes through the sales team and, outside the United States, requires approval of abuse monitoring controls and a data retention amendment, according to its documentation. Anthropic's API offers only two values for inference geography, "global" and "us", the latter billed at 1.1 times, according to its data residency page. No European geography is listed; Bedrock and Google Cloud offer endpoints in Europe, with a 10% premium. Mistral AI applies the same premium on api.eu.mistral.ai, whose data centers are located in the EU and EFTA, and states that its global endpoint "does not commit to a specific inference location".
DeepSeek's peak hours
Since 16 August 2026, DeepSeek has billed its off-peak hours at half the peak-hour price. Peak hours cover 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays. In Paris, that corresponds to 03:00 to 06:00 and 08:00 to 12:00 until 25 October, then 02:00 to 05:00 and 07:00 to 11:00 after the switch to winter time. A European team that calls the API during its working morning therefore pays the full price; weekends and Chinese public holidays are entirely off-peak. The low cache read price stems from an architectural choice that ActuIA described in September.
Reasoning tokens
Reasoning tokens are not visible in the response but are "billed as output tokens", OpenAI writes in its guide; Google labels its line "Output price (including thinking tokens)". Since output is the highest price on every rate card, the reasoning effort setting acts directly on the most expensive line of the bill.
Hosting in Europe: clouds, Mistral's EU endpoint and open weights
Under a location constraint, the reference price is not the one in the main table. The regional rate cards surveyed on 6 October, in dollars per million tokens, add 10% to the vendor's global price.
| Access | Model | Input | Cached input | Output | Batch (input / output) | Surveyed | Source |
|---|---|---|---|---|---|---|---|
| Amazon Bedrock, geographic or regional inference (Paris, Frankfurt) | Claude Opus 5.5 | 4.40 | 0.22 | 22.00 | not offered | 6 Oct 2026 | rate card |
| Amazon Bedrock, same | Claude Sonnet 5.5 | 2.20 | 0.22 | 11.00 | not offered | 6 Oct 2026 | rate card |
| Google Cloud, EU multi-region (5) | Claude Opus 5.5 | 4.40 | 0.22 | 22.00 | 2.20 / 11.00 | 6 Oct 2026 | rate card |
| Google Cloud, EU multi-region | Claude Sonnet 5.5 | 2.20 | 0.22 | 11.00 | 1.10 / 5.50 | 6 Oct 2026 | rate card |
| Google Cloud, non-global endpoint | Gemini 3.8 Flash | 0.825 | 0.0825 | 4.125 | 0.4125 / 2.0625 (6) | 6 Oct 2026 | rate card |
| Mistral, api.eu.mistral.ai endpoint (4) | Mistral Medium 3.5 | 1.65 | 0.165 | 8.25 | not offered | 6 Oct 2026 | docs |
(4) Mistral AI global price multiplied by 1.1, according to the rule published by the vendor. Batch, agents and the files API are not available on its regional endpoints, and function calling is the only supported tool. (5) On the same Google Cloud page, the global endpoint tab displays, as of 6 October, batch prices of $2.50 and $12.50 for Claude Opus 5.5. That is Opus 5's batch price, higher than the European cell, whereas Anthropic bills Opus 5.5 batch at $2 and $10. (6) Amount calculated by ActuIA: Google Cloud states that Gemini models are available in batch mode with a 50% discount, without displaying the amount for this endpoint.
Open-weight models shift the cost from the token to the infrastructure. Mistral Large 3 and Mistral Small 4 are released under the Apache 2.0 license, Mistral Medium 3.5 under a modified MIT license, according to the vendor's model list. Kolibri, released on 3 October by Aleph Alpha, is a mixture-of-experts model with 78 billion parameters, 3 billion of them active, downloadable under the Apache 2.0 license; the launch post does not publish a per-token price. The cost of use then depends on hardware, utilization rate and operations, which no rate card captures.
Two typical tasks, calculated at 6 October prices
First case: 10,000 support tickets, each with 2,000 input tokens, of which 1,500 are shared instructions read from the cache and 500 are specific to the ticket, and 400 response tokens, reasoning included. That is 15 million tokens read from cache, 5 million fresh tokens and 4 million output tokens, excluding cache writes. Second case: summarizing 1,000 documents of 20,000 tokens each, without cache, with 1,000 output tokens per document, i.e. 20 million input tokens and 1 million output tokens.
| Model | 10,000 tickets | 1,000 documents, real time | 1,000 documents, batch |
|---|---|---|---|
| GPT-6.1 Sol | $51.50 (€45.97) | $50.00 | $25.00 |
| GPT-6 Luna | $2.65 (€2.37) | $2.50 | $1.25 |
| Claude Opus 5.5 | $103.00 (€91.93) | $100.00 | $50.00 |
| Claude Sonnet 5.5 | $53.00 (€47.30) | $50.00 | $25.00 |
| Gemini 3.8 Flash | $19.88 (€17.74) | $18.75 | $9.38 |
| Mistral Medium 3.5 | $39.75 (€35.48) | $37.50 | $18.75 |
| deepseek-flash, peak / off-peak hours | $6.39 / $3.20 | $7.20 / $3.60 | n/a |
Without cache, the first case costs $80 on GPT-6.1 Sol and Claude Sonnet 5.5, compared with $51.50 and $53 with it: the cache removes about a third of the bill. Switching to batch halves the second case at the vendors that offer it. These amounts reproduce a rate card on a given volume; they say nothing about each model's success rate on these tasks.
Why an agent's bill does not follow the rate card
Two models at the same price per token do not cost the same per task. Anthropic bills Claude Sonnet 5.5 at the price of Sonnet 5 but claims it "costs up to 30% less per task", because it uses fewer tokens. For Opus 5.5, the vendor says that "at default settings" the model "will cost 40% less than Opus 5 on typical workloads". This internal measurement is not detailed and cannot be verified on a rate card.
The opposite trend is documented by Bain & Company in its Technology Report 2026: the drop in price per token is offset by higher usage, "leaving overall cost per task stubbornly high". The firm estimates that a new agentic workflow "might cost 10 times more than it ultimately should" before prompts, orchestration and model choices are optimized. The case of Uber, which capped Claude Code and Cursor after exhausting its AI budget in four months, illustrates the gap between a stable rate card and consumption that is not. The rate card sets the price of a token; the architecture sets the number of tokens.
Method and survey log
Each price comes from the official page of the vendor or cloud, consulted on 6 October 2026 between 00:00 and 01:00 UTC: OpenAI rate card, Anthropic rate card, Gemini API rate card, Mistral AI rate card, DeepSeek rate card, xAI model documentation, Amazon Bedrock pricing and Google Cloud pricing. Dollar amounts are reported without rounding, except those flagged as calculated in a footnote, at public list price, excluding negotiated discounts, capacity commitments and free tiers. For the Gemini API free tier, Google states that data is used to improve its products, which is not the case on the paid tier. The euro conversion uses the latest ECB reference rate available at the time of the survey.
Three deadlines already appear in the rate cards and in Anthropic's deprecation schedule. The GPT-5.6 Sol promotion remains guaranteed until at least 21 November 2026, and Claude Sonnet 4.5 will be retired from Anthropic's API on 30 November. The price of Gemini 3.8 Flash is due to double on 1 January 2027.