As of September 15, 2026, GPT-6 Astra โ OpenAI's new flagship, launched September 3, 2026 โ costs $10.00 per million input tokens and $50.00 per million output. GPT-5.6 Sol is discounted to $4.00/$20.00 through a promotion running to November 21, GPT-4o is still $2.50/$10.00, and GPT-4o mini is still the cheapest option at $0.15/$0.60. That's the short answer. The longer answer โ what a real task actually costs, and which models are about to disappear โ is more interesting.
Headline per-token rates are the number everyone quotes and the number that matters least. What you actually pay depends on which model you route to, how many reasoning tokens it burns invisibly, whether your input is cached, and whether you run jobs in real time or in batch. Below is the full September 2026 price list, then the math that turns those rates into a monthly bill โ plus the two hard shutdown dates (October 23 and December 11) that will break production traffic still pointed at o1 or o3 if nobody migrates it in time.

Figures from OpenAI's published API pricing as of September 15, 2026. OpenAI revises pricing frequently โ always confirm current rates before committing a budget.
OpenAI API Pricing 2026: The Full Per-Token Breakdown
OpenAI API pricing as of September 15, 2026 covers nine active models across four generations: the new GPT-6 Astra flagship, the GPT-5.6 line (Sol, Terra, Luna), the GPT-5 and GPT-4o family, and the legacy o-series reasoning models. Input rates run from $0.15 to $10.00 per million tokens, output from $0.60 to $50.00, and you pay only for tokens actually processed.
| Model | Input / 1M | Cached input / 1M | Output / 1M | Best for |
|---|---|---|---|---|
| GPT-6 Astra | $10.00 | $1.00 | $50.00 | New Sept 3, 2026 โ computer use, coding, cybersecurity |
| GPT-5.6 Sol (promo) | $4.00 | $0.40 | $20.00 | Frontier reasoning; promo runs to Nov 21, 2026 (list $5.00/$30.00) |
| GPT-5.6 Terra | $2.00 | $0.20 | $12.00 | Balanced reasoning below Sol's cost |
| GPT-5.6 Luna | $0.20 | $0.02 | $1.20 | High-volume, latency-sensitive tasks |
| GPT-5 (legacy) | $1.25 | $0.125 | $10.00 | Deprecated; API shutdown Dec 11, 2026 |
| GPT-4o | $2.50 | $1.25 | $10.00 | General-purpose default; still API-available |
| GPT-4o mini | $0.15 | $0.075 | $0.60 | Cheapest general-purpose option |
| o3 | $2.00 | $0.50 | $8.00 | Deep reasoning; API shutdown Dec 11, 2026 |
| o1 (legacy) | $15.00 | $7.50 | $60.00 | Deprecated; API shutdown Oct 23, 2026 |
Rates are per 1M tokens, USD, standard tier, as published on OpenAI's official API pricing page and OpenAI's GPT-6 Astra announcement, accessed September 15, 2026. OpenAI revises pricing frequently โ o-series rates in particular have fallen sharply since launch, while GPT-6 Astra launched well above every existing tier. Always confirm against the live pricing page before committing a budget.
GPT-6 Astra: OpenAI's New $10/$50 Flagship, Launched September 3, 2026
OpenAI launched GPT-6 Astra on September 3, 2026, priced at $10.00 per million input tokens and $50.00 per million output โ a new premium tier sitting above every existing model, including GPT-5.6 Sol. Cached input runs $1.00 per million (a 90% discount), Fast mode doubles the standard rate for lower latency, and Batch/Flex processing halves it to $5.00/$25.00. According to CNBC's September 3 report, Astra rolled out first to organizations in OpenAI's cybersecurity access program before expanding to the general API and Azure/AWS Bedrock.
Astra is not a Sol replacement โ both remain live in the API at their own price points. This likely means OpenAI is now running a genuine multi-tier premium strategy rather than a single ever-cheaper flagship: Astra for the highest-stakes agentic and coding work, Sol for general frontier reasoning, and Terra/Luna for everyday and high-volume traffic. Whether that pricing gap holds is unproven โ Sol itself was list-priced at $5.00/$30.00 in July and cut to $4.00/$20.00 by August 21, so a similar promotional discount on Astra within its first few months would not be surprising.
GPT-5.6 Sol, Terra, and Luna: Why OpenAI's Flagship Got More Expensive, Then Cheaper Again
OpenAI replaced its GPT-5 lineup with three GPT-5.6 tiers on July 29, 2026: Sol (flagship), Terra (mid-tier), and Luna (budget). The headline change at launch was that Sol cost more than GPT-5 did โ $5.00 versus $1.25 per million input tokens โ reversing several straight generations of flagship-model price cuts. According to OpenAI's own GPT-5.6 announcement, the reason was infrastructure: Sol is the first OpenAI model to run in production on dedicated Cerebras wafer-scale chips, delivering up to 750 tokens per second versus roughly 70 tokens per second on a typical Nvidia H100 cluster running the same weights.
On August 21, 2026, OpenAI cut Sol's API price by 20% on input and 33% on output โ to $4.00 and $20.00 per million tokens โ with that promotional rate confirmed to run at least through November 21, 2026. Terra and Luna, cut 20% and 80% respectively on July 30, sit at $2.00 and $0.20 input and have not moved since. Most existing production traffic can migrate to one of those two tiers without much dollar movement, leaving Sol (and now Astra) for the narrower slice of latency-critical, highest-capability workloads.
What OpenAI API Pricing Looks Like Per Real Task
A token is roughly 0.75 words, so 1,000 tokens is about 750 words. Most real requests are small โ a few hundred tokens in, a few hundred out. The table below prices six common workloads on the model you'd actually pick, assuming standard (uncached) input.
Route a chat message to the right agent
GPT-5.6 Luna ยท 300 in / 20 out
Classify a support ticket
GPT-4o mini ยท 500 in / 50 out
Summarize a 10-page PDF
GPT-4o ยท 8K in / 600 out
Draft a marketing email
GPT-4o ยท 800 in / 500 out
Answer a hard math/logic question
o3 ยท 1K in / 12K out*
Analyze a 200-page legal contract
GPT-5.6 Sol (promo) ยท 150K in / 4K out
Run a multi-step coding/agentic task
GPT-6 Astra ยท 40K in / 6K out
*The o3 example assumes ~12K hidden reasoning tokens billed at the output rate โ the single most underestimated line item in any reasoning-model budget.
The lesson: individual calls are cheap, but volume compounds fast. A product doing 5 million GPT-4o summaries a month at $0.026 each is spending $130,000 โ and almost all of that is avoidable with the right routing and caching, covered below. For how this rolls up across the industry, see our AI Spending dashboard.
Why the o3 and o1 Reasoning Models Cost More โ And Are Being Retired
Reasoning models think before they answer, and that thinking is made of tokens you pay for. A GPT-4o answer might be 400 output tokens. The same prompt on o3 can generate 8,000-20,000 internal reasoning tokens plus the visible answer โ all billed at the $8.00-per-million output rate. That's why o3, at a lower headline price than the older o1, can still cost 10-15x more than GPT-4o for the same question. Both are on the way out: OpenAI's deprecations page lists o1, o1-pro, o3-mini, and o4-mini for API shutdown on October 23, 2026, with full o3 and o3-pro following on December 11, 2026 alongside the original GPT-5 snapshot family. GPT-5.6 Terra is OpenAI's recommended migration path off all of them for most reasoning-heavy production traffic.
Hidden reasoning tokens
5K-20K tokens per answer, billed at output rate, invisible until the bill arrives
Longer outputs
Reasoning answers run 3-5x longer than a comparable GPT-4o response
Retry sensitivity
A failed or truncated reasoning chain still bills for every token generated
Context reuse
Without caching, long reasoning prompts re-bill full input on every call
The practical rule: don't send a task to o3 or Sol unless GPT-4o mini, GPT-4o, or Terra demonstrably fails it. Reasoning and flagship-tier models are a scalpel, not a default. For roughly 80% of production calls, GPT-4o mini or Luna is both cheaper and fast enough.
OpenAI API Pricing vs Anthropic and Google in 2026
OpenAI isn't priced in a vacuum. Anthropic's Claude pricing and Google's Gemini API pricing set the competitive floor, and GPT-6 Astra now costs more per input token than every flagship competitor listed below. Here's how the top and budget tiers line up as of September 2026.
| Model | Provider | Input / 1M | Output / 1M |
|---|---|---|---|
| GPT-6 Astra | OpenAI | $10.00 | $50.00 |
| GPT-5.6 Sol (promo) | OpenAI | $4.00 | $20.00 |
| GPT-4o | OpenAI | $2.50 | $10.00 |
| GPT-4o mini | OpenAI | $0.15 | $0.60 |
| Claude Opus 4.5 | Anthropic | $5.00 | $25.00 |
| Claude Sonnet 5 | Anthropic | $2.00 | $10.00 |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 |
| Gemini 3.1 Pro | $2.00 | $12.00 | |
| Gemini 3.8 Flash | $0.75 | $3.75 |
Sources: OpenAI, Anthropic, and Google published API pricing pages, accessed September 15, 2026.
The takeaway: GPT-6 Astra costs 5x Gemini 3.1 Pro on input and 2x Claude Opus 4.5, the widest gap OpenAI's top tier has opened over either competitor's flagship this year. GPT-4o mini remains the cheapest credible general-purpose model at $0.15 input, ahead of Gemini 3.8 Flash's $0.75. Both Anthropic and Google cut flagship prices earlier in 2026 while OpenAI moved the opposite direction with Astra โ a genuine divergence in strategy, not just a one-off. For the full competitive picture, see our OpenAI vs Anthropic enterprise breakdown.
Four Levers That Cut an OpenAI API Bill 50-70%
Do this
- โ Prompt caching: up to 90% off on GPT-6 Astra, GPT-5.6, and GPT-5
- โ Batch API: 50% off for non-urgent jobs (24h window; half-price on Astra too)
- โ Route simple tasks to GPT-4o mini or Luna (17-50x cheaper input than Astra)
- โ Trim system prompts โ they bill on every single call
Stop doing this
- โ Defaulting every call to Astra, Sol, o3, or o1
- โ Sending uncached 20K-token system prompts each request
- โ Paying real-time rates for overnight batch work
- โ Still routing production traffic to o1 or o3-mini past October 23
Stacked together, these are not marginal. A team caching a 15K-token system prompt across 2 million calls, batching its analytics jobs, and downshifting classification to GPT-4o mini routinely takes a $40,000 monthly bill under $14,000 โ same product, same quality. The single highest-ROI change is usually prompt caching, because most production apps resend the same instructions on every request.
What the headline pricing misses
The clean story โ "OpenAI keeps getting cheaper" โ doesn't survive contact with GPT-6 Astra. Astra costs 2x GPT-5.6 Sol's promotional rate and 8x GPT-5.6 Luna, a second consecutive flagship price increase after Sol broke a multi-year run of cuts back in July. Anyone budgeting off last year's per-token math and routing production traffic straight to the newest flagship without checking current rates first could see a bill spike, not a discount.
A second caveat: the discount math above assumes standard list pricing. Enterprise customers frequently negotiate committed-use or volume contracts off-list, and those terms aren't published, so a large team's actual per-token cost can differ meaningfully from every number in this article. This likely means the gap between GPT-6 Astra and its competitors is smaller in practice for high-volume enterprise buyers than the public price list suggests, though there's no public data to size that gap precisely. A third: Sol's $4.00/$20.00 rate is a promotion confirmed only through November 21, 2026 โ anyone building a 2027 budget off today's promotional price is guessing, not calculating. Cross-model tracking sites such as Artificial Analysis are a reasonable way to sanity-check list prices against real-world benchmarking data before committing a budget.
The per-token number on the pricing page is not your bill.
Routing, caching, and batching decide whether GPT-4o costs you $14K or $40K a month โ and that gap is entirely an engineering choice.
See how these pricing moves translate into revenue in our OpenAI Revenue 2026 breakdown, or how GPT-5.6 Sol's Cerebras deployment works in our Cerebras pricing and speed breakdown. Track AI model economics on the AI Valuations and AI Spending dashboards at Value Add VC.
Latest from the Pulse
Get VC data most people never see
โ 100% free
Weekly benchmarks, valuations, and fund data. Join 5,000+ investors. No spam.