AI & TechnologyJune 21, 2026ยท13 min readยทยทLast updated: 2026-09-15

OpenAI API Pricing 2026: GPT-6 Astra, GPT-4o, o3, and GPT-5.6 Cost Per Token Breakdown

Every OpenAI model, what it costs per million tokens as of September 2026, what a real task actually costs to run, and the four levers that cut an API bill by more than half.

TC
Trace Cohen
Founder, Value Add Holdings LLC ยท 3x founder (BrandYourself, Launch.it, SPOT) ยท 65+ investments ยท Based in Boca Raton, FL
65+Investments3xFounder$200M+Funds Tracked

Quick Answer

$10 per million input tokens is what GPT-6 Astra costs via the OpenAI API as of its September 3, 2026 launch, more than double GPT-5.6 Sol's discounted $4 promotional rate. That Sol promo runs through November 21, and o1 shuts down October 23.

As of September 15, 2026, GPT-6 Astra โ€” OpenAI's new flagship, launched September 3, 2026 โ€” costs $10.00 per million input tokens and $50.00 per million output. GPT-5.6 Sol is discounted to $4.00/$20.00 through a promotion running to November 21, GPT-4o is still $2.50/$10.00, and GPT-4o mini is still the cheapest option at $0.15/$0.60. That's the short answer. The longer answer โ€” what a real task actually costs, and which models are about to disappear โ€” is more interesting.

Headline per-token rates are the number everyone quotes and the number that matters least. What you actually pay depends on which model you route to, how many reasoning tokens it burns invisibly, whether your input is cached, and whether you run jobs in real time or in batch. Below is the full September 2026 price list, then the math that turns those rates into a monthly bill โ€” plus the two hard shutdown dates (October 23 and December 11) that will break production traffic still pointed at o1 or o3 if nobody migrates it in time.

OpenAI API Pricing 2026: GPT-6 Astra, GPT-4o, o3, and GPT-5.6 Cost Per Token Breakdown
$10.00/M
new Sept 3, 2026
GPT-6 Astra input price
$4.00/M
-20% since Aug 21
GPT-5.6 Sol (promo)
$0.15/M
unchanged
Cheapest model (GPT-4o mini)
Oct 23, 2026
o3/o3-pro follow Dec 11
o1 API shutdown

Figures from OpenAI's published API pricing as of September 15, 2026. OpenAI revises pricing frequently โ€” always confirm current rates before committing a budget.

OpenAI API Pricing 2026: The Full Per-Token Breakdown

OpenAI API pricing as of September 15, 2026 covers nine active models across four generations: the new GPT-6 Astra flagship, the GPT-5.6 line (Sol, Terra, Luna), the GPT-5 and GPT-4o family, and the legacy o-series reasoning models. Input rates run from $0.15 to $10.00 per million tokens, output from $0.60 to $50.00, and you pay only for tokens actually processed.

ModelInput / 1MCached input / 1MOutput / 1MBest for
GPT-6 Astra$10.00$1.00$50.00New Sept 3, 2026 โ€” computer use, coding, cybersecurity
GPT-5.6 Sol (promo)$4.00$0.40$20.00Frontier reasoning; promo runs to Nov 21, 2026 (list $5.00/$30.00)
GPT-5.6 Terra$2.00$0.20$12.00Balanced reasoning below Sol's cost
GPT-5.6 Luna$0.20$0.02$1.20High-volume, latency-sensitive tasks
GPT-5 (legacy)$1.25$0.125$10.00Deprecated; API shutdown Dec 11, 2026
GPT-4o$2.50$1.25$10.00General-purpose default; still API-available
GPT-4o mini$0.15$0.075$0.60Cheapest general-purpose option
o3$2.00$0.50$8.00Deep reasoning; API shutdown Dec 11, 2026
o1 (legacy)$15.00$7.50$60.00Deprecated; API shutdown Oct 23, 2026

Rates are per 1M tokens, USD, standard tier, as published on OpenAI's official API pricing page and OpenAI's GPT-6 Astra announcement, accessed September 15, 2026. OpenAI revises pricing frequently โ€” o-series rates in particular have fallen sharply since launch, while GPT-6 Astra launched well above every existing tier. Always confirm against the live pricing page before committing a budget.

GPT-6 Astra: OpenAI's New $10/$50 Flagship, Launched September 3, 2026

OpenAI launched GPT-6 Astra on September 3, 2026, priced at $10.00 per million input tokens and $50.00 per million output โ€” a new premium tier sitting above every existing model, including GPT-5.6 Sol. Cached input runs $1.00 per million (a 90% discount), Fast mode doubles the standard rate for lower latency, and Batch/Flex processing halves it to $5.00/$25.00. According to CNBC's September 3 report, Astra rolled out first to organizations in OpenAI's cybersecurity access program before expanding to the general API and Azure/AWS Bedrock.

Astra is not a Sol replacement โ€” both remain live in the API at their own price points. This likely means OpenAI is now running a genuine multi-tier premium strategy rather than a single ever-cheaper flagship: Astra for the highest-stakes agentic and coding work, Sol for general frontier reasoning, and Terra/Luna for everyday and high-volume traffic. Whether that pricing gap holds is unproven โ€” Sol itself was list-priced at $5.00/$30.00 in July and cut to $4.00/$20.00 by August 21, so a similar promotional discount on Astra within its first few months would not be surprising.

GPT-5.6 Sol, Terra, and Luna: Why OpenAI's Flagship Got More Expensive, Then Cheaper Again

OpenAI replaced its GPT-5 lineup with three GPT-5.6 tiers on July 29, 2026: Sol (flagship), Terra (mid-tier), and Luna (budget). The headline change at launch was that Sol cost more than GPT-5 did โ€” $5.00 versus $1.25 per million input tokens โ€” reversing several straight generations of flagship-model price cuts. According to OpenAI's own GPT-5.6 announcement, the reason was infrastructure: Sol is the first OpenAI model to run in production on dedicated Cerebras wafer-scale chips, delivering up to 750 tokens per second versus roughly 70 tokens per second on a typical Nvidia H100 cluster running the same weights.

On August 21, 2026, OpenAI cut Sol's API price by 20% on input and 33% on output โ€” to $4.00 and $20.00 per million tokens โ€” with that promotional rate confirmed to run at least through November 21, 2026. Terra and Luna, cut 20% and 80% respectively on July 30, sit at $2.00 and $0.20 input and have not moved since. Most existing production traffic can migrate to one of those two tiers without much dollar movement, leaving Sol (and now Astra) for the narrower slice of latency-critical, highest-capability workloads.

What OpenAI API Pricing Looks Like Per Real Task

A token is roughly 0.75 words, so 1,000 tokens is about 750 words. Most real requests are small โ€” a few hundred tokens in, a few hundred out. The table below prices six common workloads on the model you'd actually pick, assuming standard (uncached) input.

Route a chat message to the right agent

GPT-5.6 Luna ยท 300 in / 20 out

$0.00008

Classify a support ticket

GPT-4o mini ยท 500 in / 50 out

$0.00011

Summarize a 10-page PDF

GPT-4o ยท 8K in / 600 out

$0.026

Draft a marketing email

GPT-4o ยท 800 in / 500 out

$0.007

Answer a hard math/logic question

o3 ยท 1K in / 12K out*

$0.098

Analyze a 200-page legal contract

GPT-5.6 Sol (promo) ยท 150K in / 4K out

$0.68

Run a multi-step coding/agentic task

GPT-6 Astra ยท 40K in / 6K out

$0.70

*The o3 example assumes ~12K hidden reasoning tokens billed at the output rate โ€” the single most underestimated line item in any reasoning-model budget.

The lesson: individual calls are cheap, but volume compounds fast. A product doing 5 million GPT-4o summaries a month at $0.026 each is spending $130,000 โ€” and almost all of that is avoidable with the right routing and caching, covered below. For how this rolls up across the industry, see our AI Spending dashboard.

Why the o3 and o1 Reasoning Models Cost More โ€” And Are Being Retired

Reasoning models think before they answer, and that thinking is made of tokens you pay for. A GPT-4o answer might be 400 output tokens. The same prompt on o3 can generate 8,000-20,000 internal reasoning tokens plus the visible answer โ€” all billed at the $8.00-per-million output rate. That's why o3, at a lower headline price than the older o1, can still cost 10-15x more than GPT-4o for the same question. Both are on the way out: OpenAI's deprecations page lists o1, o1-pro, o3-mini, and o4-mini for API shutdown on October 23, 2026, with full o3 and o3-pro following on December 11, 2026 alongside the original GPT-5 snapshot family. GPT-5.6 Terra is OpenAI's recommended migration path off all of them for most reasoning-heavy production traffic.

Hidden reasoning tokens

5K-20K tokens per answer, billed at output rate, invisible until the bill arrives

Longer outputs

Reasoning answers run 3-5x longer than a comparable GPT-4o response

Retry sensitivity

A failed or truncated reasoning chain still bills for every token generated

Context reuse

Without caching, long reasoning prompts re-bill full input on every call

The practical rule: don't send a task to o3 or Sol unless GPT-4o mini, GPT-4o, or Terra demonstrably fails it. Reasoning and flagship-tier models are a scalpel, not a default. For roughly 80% of production calls, GPT-4o mini or Luna is both cheaper and fast enough.

OpenAI API Pricing vs Anthropic and Google in 2026

OpenAI isn't priced in a vacuum. Anthropic's Claude pricing and Google's Gemini API pricing set the competitive floor, and GPT-6 Astra now costs more per input token than every flagship competitor listed below. Here's how the top and budget tiers line up as of September 2026.

ModelProviderInput / 1MOutput / 1M
GPT-6 AstraOpenAI$10.00$50.00
GPT-5.6 Sol (promo)OpenAI$4.00$20.00
GPT-4oOpenAI$2.50$10.00
GPT-4o miniOpenAI$0.15$0.60
Claude Opus 4.5Anthropic$5.00$25.00
Claude Sonnet 5Anthropic$2.00$10.00
Claude Haiku 4.5Anthropic$1.00$5.00
Gemini 3.1 ProGoogle$2.00$12.00
Gemini 3.8 FlashGoogle$0.75$3.75

Sources: OpenAI, Anthropic, and Google published API pricing pages, accessed September 15, 2026.

The takeaway: GPT-6 Astra costs 5x Gemini 3.1 Pro on input and 2x Claude Opus 4.5, the widest gap OpenAI's top tier has opened over either competitor's flagship this year. GPT-4o mini remains the cheapest credible general-purpose model at $0.15 input, ahead of Gemini 3.8 Flash's $0.75. Both Anthropic and Google cut flagship prices earlier in 2026 while OpenAI moved the opposite direction with Astra โ€” a genuine divergence in strategy, not just a one-off. For the full competitive picture, see our OpenAI vs Anthropic enterprise breakdown.

Four Levers That Cut an OpenAI API Bill 50-70%

Do this

  • โœ“ Prompt caching: up to 90% off on GPT-6 Astra, GPT-5.6, and GPT-5
  • โœ“ Batch API: 50% off for non-urgent jobs (24h window; half-price on Astra too)
  • โœ“ Route simple tasks to GPT-4o mini or Luna (17-50x cheaper input than Astra)
  • โœ“ Trim system prompts โ€” they bill on every single call

Stop doing this

  • โœ• Defaulting every call to Astra, Sol, o3, or o1
  • โœ• Sending uncached 20K-token system prompts each request
  • โœ• Paying real-time rates for overnight batch work
  • โœ• Still routing production traffic to o1 or o3-mini past October 23

Stacked together, these are not marginal. A team caching a 15K-token system prompt across 2 million calls, batching its analytics jobs, and downshifting classification to GPT-4o mini routinely takes a $40,000 monthly bill under $14,000 โ€” same product, same quality. The single highest-ROI change is usually prompt caching, because most production apps resend the same instructions on every request.

What the headline pricing misses

The clean story โ€” "OpenAI keeps getting cheaper" โ€” doesn't survive contact with GPT-6 Astra. Astra costs 2x GPT-5.6 Sol's promotional rate and 8x GPT-5.6 Luna, a second consecutive flagship price increase after Sol broke a multi-year run of cuts back in July. Anyone budgeting off last year's per-token math and routing production traffic straight to the newest flagship without checking current rates first could see a bill spike, not a discount.

A second caveat: the discount math above assumes standard list pricing. Enterprise customers frequently negotiate committed-use or volume contracts off-list, and those terms aren't published, so a large team's actual per-token cost can differ meaningfully from every number in this article. This likely means the gap between GPT-6 Astra and its competitors is smaller in practice for high-volume enterprise buyers than the public price list suggests, though there's no public data to size that gap precisely. A third: Sol's $4.00/$20.00 rate is a promotion confirmed only through November 21, 2026 โ€” anyone building a 2027 budget off today's promotional price is guessing, not calculating. Cross-model tracking sites such as Artificial Analysis are a reasonable way to sanity-check list prices against real-world benchmarking data before committing a budget.

The per-token number on the pricing page is not your bill.

Routing, caching, and batching decide whether GPT-4o costs you $14K or $40K a month โ€” and that gap is entirely an engineering choice.

See how these pricing moves translate into revenue in our OpenAI Revenue 2026 breakdown, or how GPT-5.6 Sol's Cerebras deployment works in our Cerebras pricing and speed breakdown. Track AI model economics on the AI Valuations and AI Spending dashboards at Value Add VC.

Get VC data most people never see

โ€” 100% free

Weekly benchmarks, valuations, and fund data. Join 5,000+ investors. No spam.

Frequently Asked Questions

How much does the OpenAI API cost in 2026?

As of September 15, 2026, OpenAI API pricing spans $0.15 per million input tokens for GPT-4o mini up to $10.00 per million for the new GPT-6 Astra flagship, which launched September 3, 2026 at $10.00 input and $50.00 output per million tokens. GPT-5.6 Sol is discounted to $4.00 input and $20.00 output through a promotion running to November 21, 2026, while GPT-4o remains $2.50 input and $10.00 output. You only pay for tokens actually processed, billed per request.

How much does GPT-6 Astra cost compared to GPT-5.6 Sol?

GPT-6 Astra costs $10.00 per million input tokens and $50.00 per million output tokens, more than double Sol's promotional $4.00/$20.00 rate (list price $5.00/$30.00). OpenAI positioned Astra as its most capable model yet for computer use, coding, and cybersecurity work, priced as a distinct premium tier above Sol rather than a direct replacement โ€” Sol and Terra both remain live in the API.

Why is the o3 reasoning model more expensive than GPT-4o?

The o3 reasoning model costs about $2.00 input and $8.00 output per million tokens, but the real cost driver is hidden reasoning tokens. A single o3 answer can burn 5,000 to 20,000 internal reasoning tokens billed at the output rate, so a task that costs $0.01 on GPT-4o can cost $0.15 or more on o3 despite similar headline pricing. OpenAI has scheduled o3 and o3-pro for API shutdown on December 11, 2026.

How can I reduce my OpenAI API bill?

The four biggest levers are prompt caching (up to 90% off repeated input on GPT-6 Astra, GPT-5.6, and GPT-5 models), the Batch API (50% off for non-urgent jobs), routing simple tasks to GPT-4o mini or the low-cost GPT-5.6 Luna tier instead of GPT-4o, Sol, or Astra, and trimming system prompts. Combined, these routinely cut a production API bill by 50-70% without changing output quality.

Which OpenAI models are being shut down in 2026?

OpenAI is retiring o1, o1-pro, o3-mini, o4-mini, and both deep-research models from the API on October 23, 2026, followed by full o3, o3-pro, and the original GPT-5, GPT-5 mini/nano/pro snapshots on December 11, 2026. Teams still on any of these should migrate to GPT-5.6 Terra or GPT-6 Astra before those dates to avoid a hard API error.

Is GPT-4o mini still the cheapest option in 2026?

Yes, narrowly. GPT-4o mini's $0.15 per million input tokens undercuts the GPT-5.6 Luna tier's $0.20 by 25%, though Luna's $1.20 output price is exactly double GPT-4o mini's $0.60. For input-heavy workloads GPT-4o mini stays cheapest; for balanced input/output mixes the two are close enough that task quality should decide.

Companies & investors in this article

Explore 45+ free VC tools, dashboards, and recommended startup software.

Get VC data most people never see

Weekly benchmarks & analysis. Join 5,000+ investors.