OpenAI API Pricing (Sept 2026): GPT-6 Astra, Sol & Luna Costs

OpenAI’s flagship GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens, while the budget GPT-6 Luna costs $0.10 and $0.50. Between them sits GPT-6 Sol at $2 and $10. That 100x spread is why model choice, not traffic, usually decides your bill.

Prices have moved three times since July, so older guides are already wrong. This article covers the current rates, the multipliers that change them, and how to estimate your own costs.

cloudzero.com

How much does the OpenAI API cost right now?

Short answer: it depends on the model. OpenAI’s official pricing page lists GPT-6 Astra at $10 input and $50 output per million tokens at the Standard tier, GPT-6 Sol at $2 and $10, and GPT-6 Luna at $0.10 and $0.50.

GPT-6 model (Standard, short context)InputCached inputOutput
gpt-6-astra$10.00$1.00$50.00
gpt-6-sol$2.00$0.20$10.00
gpt-6-luna$0.10$0.01$0.50

OpenAI released Sol and Luna on September 22, 2026, and that was 19 days after Astra. Per 1,000 tokens, Sol costs $0.002 for input and $0.01 for output.

image 14 1

zylo.com

How does OpenAI API billing work?

Short answer: you pay per token, with separate rates for what you send (input) and what the model writes back (output). The API is pay-as-you-go with no subscription or seat fees, and ChatGPT Plus or Business subscriptions include no API credit. Billing runs on prepaid credits that draw down as you use the API.

A token is roughly three-quarters of an English word, so a million tokens is about 750,000 words. Four line items appear on most invoices:

  • Input tokens: your prompt, system instructions, conversation history and attached documents.
  • Output tokens: the model’s reply. This includes reasoning tokens on reasoning models, which is why short answers can still cost more than expected.
  • Cached input: a repeated prompt prefix billed at a steep discount. On GPT-6, cached input is 10% of the normal input rate.
  • Cache writes: cache writes are billed at 1.25x the uncached input rate on GPT-5.6 Sol, and GPT-6 Sol’s listed cache-write rate ($2.50) follows the same ratio.

Whichever endpoint you use, the price is the same. The Responses, Chat Completions, Realtime, Batch and Assistants APIs are not priced separately.

What happened to GPT-5.6 and older models?

Short answer: GPT-5.6 is still sold, but GPT-6 Sol and Luna now undercut it.

The GPT-5.6 family has three tiers, and the ladder changed on July 30, 2026. Sol is $4/$20 (promotional through at least November 21, 2026), Terra is $2/$12, and Luna is $0.20/$1.20. The July 30 cut is what brought Luna and Terra down to those rates.

The GPT-6 launch didn’t replace the whole lineup. OpenAI shipped no GPT-6 Terra, so the middle tier of GPT-5.6 has no successor. That leaves Terra at $2/$12, the same input price as GPT-6 Sol with a higher output price, which makes it hard to justify for new projects.

Older models remain available too. GPT-5.5 stays at $5/$30. GPT-5.4 lists at $2.50 input, $0.25 cached input and $15 output, with a 1.05M-token context window. GPT-6 Sol is therefore both cheaper and newer than GPT-5.4.

One caution: cheaper doesn’t mean better on every task. On two of the three coding and computer-use charts OpenAI published, GPT-5.6 Sol’s best score beat GPT-6 Sol’s. Test your own workload before migrating.

What do long context, Batch, Flex and Fast mode do to the price?

Short answer: long prompts cost more, and the processing tier can halve or double your rate.

  • Long context: On GPT-5.6 Sol, prompts over 272K input tokens are priced at 2x input and 1.5x output for the full request. On GPT-6, Astra’s long-context rate is $20 input and $75 output, and Sol’s is $4 and $15. The 272K breakpoint appears in third-party tracking of Astra’s pricing. Confirm it on each model page.
  • Batch: asynchronous jobs get 50% off. GPT-6 Sol falls to $1 input and $5 output.
  • Flex: Flex matches Batch pricing for most models while staying on the synchronous API.
  • Fast mode: This is the low-latency tier at 2x Standard, renamed from Priority processing on July 30, 2026.
  • Data residency: Regional processing endpoints carry a 10% uplift for eligible models.

How do you calculate your cost per API call?

Short answer: multiply tokens by the per-million rate, then divide by one million.

Cost = (input tokens × input rate + output tokens × output rate) ÷ 1,000,000

Say a support bot sends 2,000 input tokens and gets 500 back, across 100,000 requests a month (200M input, 50M output):

ModelPer requestMonthly cost
GPT-6 Astra$0.045$4,500
GPT-6 Sol$0.009$900
GPT-6 Luna$0.00045$45

Cache 1,500 of Sol’s 2,000 input tokens and the per-request cost drops to about $0.0063, or roughly $630 a month. Run the same job asynchronously on Batch and Sol falls to about $450. These figures are my own arithmetic from the published rates. They ignore cache-write charges and tool fees.

Use the token-counting guide in OpenAI’s docs to measure real prompts. Third-party calculators such as BenchLM’s also let you enter monthly volume and cache share, though OpenAI’s documentation is the authority on rates.

What do embeddings, Whisper and tools cost?

Embeddings are cheap. text-embedding-3-small costs $0.02 per million tokens and text-embedding-3-large costs $0.13. Batch halves both, and only input is billed, since no text is generated. Indexing 10,000 documents of 500 tokens each costs about ten cents on the small model.

Transcription is billed per minute. OpenAI’s page lists gpt-transcribe at $0.0045 per minute, gpt-4o-transcribe at $0.006, and gpt-4o-mini-transcribe at $0.003. The older whisper-1 was billed at $0.006 per minute, but that figure comes from a third-party guide, so check before relying on it.

Tools add separate charges. Web search costs $10 per 1,000 calls plus content tokens at model rates, file search calls cost $2.50 per 1,000, and hosted containers start at $0.03 per 20-minute session.

Fine-tuning is closing. OpenAI is winding down its fine-tuning platform, which is no longer open to new users, and existing fine-tuned models stay available until their base models are deprecated. Budget for prompt engineering and caching instead.

Is Azure OpenAI cheaper than the OpenAI API?

Short answer: per-token rates are broadly the same, but Azure adds deployment options with different economics.

Azure offers pay-as-you-go pricing at rates generally comparable to the direct API, plus provisioned throughput units (PTUs) that reserve capacity. Microsoft says Standard deployment for Astra, Sol and Luna spans 28 Global regions plus US and EU Data Zones, and Data Zone Provisioned Throughput carries a 10% premium over Global.

The real trade-off is commitment. PTUs bill per reserved hour whether or not you use them, so they are a commitment rather than a discount. Choose Azure for enterprise governance, private networking or existing Microsoft agreements. Choose the direct API when you want the newest models and features first, since Azure matches OpenAI’s token rates but not its timing.

image 15 1

How does OpenAI pricing compare with Claude and Gemini?

Short answer: OpenAI leads on the cheap end, and the three vendors are close in the middle.

TierOpenAIAnthropicGoogle
PremiumGPT-6 Astra $10/$50Fable 5.1 $10/$50Gemini 3.1 Pro $2/$12
MidGPT-6 Sol $2/$10Opus 5.5 $4/$20; Sonnet 5 $2/$10Gemini 3.8 Flash $0.75/$3.75*
BudgetGPT-6 Luna $0.10/$0.50Haiku 4.5 $1/$5Gemini 2.5 Flash-Lite $0.10/$0.40

Three caveats matter more than the headline numbers:

  1. Gemini 3.8 Flash’s $0.75/$3.75 rate is introductory through December 31, 2026, and rises to $1.50/$7.50 on January 1, 2027.
  2. Claude models from 4.7 onward use a tokenizer that produces roughly 30% more tokens for the same text, so per-million comparisons flatter Anthropic slightly.
  3. Gemini output pricing includes thinking tokens, so reasoning-heavy prompts bill at the output rate even when the visible answer is short.

Compare cost per completed task on your own prompts, not the rate card alone.

How can you reduce your OpenAI API bill?

  1. Route by difficulty. Send classification, extraction and summaries to Luna, everyday agent work to Sol, and only the tasks Sol measurably fails to Astra. Luna’s input price is 1/100th of Astra’s.
  2. Structure prompts for caching. Put static instructions and examples first, and variable user content last.
  3. Batch anything that can wait. Nightly enrichment, evals and bulk classification qualify for the 50% discount.
  4. Watch long context. One request crossing 272K tokens reprices the whole call, so trim or chunk documents.
  5. Cap output length and set reasoning effort no higher than the task needs.
  6. Set spend limits in the dashboard and track cost per feature, not just the monthly total.
  7. Re-check prices quarterly. The July 30 cuts and September launches show how fast rates change. Sol’s GPT-5.6 rate is promotional and only guaranteed through November 21, 2026.
image 16 1

Frequently asked questions

Is the OpenAI API free? (OpenAI API pricing free)

Not on an ongoing basis that I could confirm on OpenAI’s pricing page. The API is prepaid and pay-as-you-go, and a ChatGPT subscription doesn’t include API credit. For low-cost testing, GPT-6 Luna at $0.10 per million input tokens makes experiments cost cents.

What is GPT API pricing?

GPT API pricing is per million tokens, split into input and output. GPT-6 Luna costs $0.10/$0.50, GPT-6 Sol $2/$10 and GPT-6 Astra $10/$50, with discounts for cached input and Batch and surcharges for long context and Fast mode.

What is ChatGPT API pricing?

There is no separate ChatGPT API. You pay the token rate of whichever model you call. The chat-latest alias, which tracks the current ChatGPT model, lists at $5 input, $0.50 cached input and $30 output per million tokens.

What are OpenAI’s pricing plans?

For developers, the API is pay-as-you-go with usage tiers that raise rate limits as you spend. Consumer and business ChatGPT plans are billed separately and don’t include API usage. Fast mode, Flex and Batch are processing options rather than plans.

Is there an OpenAI API pricing calculator?

OpenAI’s pricing docs link an image-input cost calculator rather than a full-bill one. To estimate a bill, multiply input and output tokens by the per-million rates and divide by 1,000,000, or use a third-party calculator such as BenchLM’s and check the rates against OpenAI’s page.

What is the OpenAI API pricing for all models?

The main current rates per million input/output tokens are:

  • GPT-6 Astra: $10/$50
  • GPT-6 Sol: $2/$10
  • GPT-6 Luna: $0.10/$0.50
  • GPT-5.6 Sol: $4/$20 (promotional)
  • GPT-5.6 Terra: $2/$12
  • GPT-5.6 Luna: $0.20/$1.20
  • GPT-5.5: $5/$30
  • GPT-5.4: $2.50/$15

Embeddings run $0.02–$0.13 per million tokens.

What is GPT-5.4 API pricing?

GPT-5.4 costs $2.50 per million input tokens, $0.25 cached and $15 output. That is half the price of GPT-5.5, but GPT-6 Sol is cheaper still at $2/$10.

What is Azure OpenAI pricing?

Azure OpenAI offers Standard pay-as-you-go pricing at rates generally comparable to OpenAI’s direct API, plus Provisioned Throughput Units for reserved capacity. Data Zone provisioned deployments cost about 10% more than Global. Check the Azure pricing page for current regional rates.

Author Bio: Hamid Ali is a technology and AI writer covering OpenAI API pricing, generative AI models, cloud computing, and emerging AI tools. He focuses on clear, research-driven insights that help readers understand AI costs, model pricing, and practical technology trends.

Author Name: Hamid Ali
Email: johanharwen314@gmail.com

Also Read:

Comcast Data Breach Lawsuit Settlement
PlayStation Video Library Licensing Agreements
How to Fix Slow Startup on Windows 10 and 11

Leave a Reply

Your email address will not be published. Required fields are marked *