Total unique visitors
Browse by category Chatbots Image Generation Video Generation Audio & Voice Coding Writing Productivity Research AI Agents Free Tier Table
Home page Ask Cat on AI

Ask CatAI Tool SummaryOpenRouter

GPT-6 Astra Lands on OpenRouter: Five Providers for One Model, and the Priciest Costs 4x the Cheapest (Checked Sept 2026)

🐾 Quick facts
  • Free tier:There is
  • Cheapest paid plan:US$0/mo and up
  • Free quota:More than 25 free models, 4 free providers, daily limit of 50 …
  • Last checked:2026-09-21

Article last updated:2026-09-05

OpenAI’s new flagship, GPT-6 Astra, went live on OpenRouter on 2026-09-04.

The obvious headline is that you no longer need a ChatGPT subscription to call it. The more useful headline is this: the same GPT-6 Astra is served by five separate provider endpoints, and the most expensive one charges four times what the cheapest one charges for output.

Here is every endpoint, what it costs, and what the cheap tier gives up.

1. The basics

Read directly from the OpenRouter model page openai/gpt-6-astra on 2026-09-05:

  • Listed: 2026-09-04
  • Context window: 1,050,000 tokens
  • Max completion: 128,000 tokens
  • Supported: function calling (tools, tool_choice) and structured outputs via JSON schema in response_format
  • Web search: US$10 per 1,000 calls

The headline price at the top of the page is US$10 per million input tokens and US$50 per million output tokens, with cache reads at US$1.00/M and cache writes at US$12.50/M.

2. Five endpoints, a 4x spread

Scroll down the same page and the provider list reads (per million tokens):

Provider endpointInputOutputCache read
OpenAI FlexUS$5.00US$25.00US$0.50
AzureUS$10.00US$50.00US$1.00
OpenAIUS$10.00US$50.00US$1.00
Azure (US)US$11.00US$55.00US$1.10
OpenAI FastUS$20.00US$100.00US$2.00

Top row to bottom row: 4x on input, 4x on output.

In plain terms, one run on the priciest endpoint buys you four runs on the cheapest one.

3. Cross-checked against OpenAI’s own price list

One source is not enough, so here is OpenAI’s official API pricing page (developers.openai.com) for the same model:

Official tierInputCached inputCache writesOutput
StandardUS$10.00US$1.00US$12.50US$50.00
BatchUS$5.00US$0.50US$6.25US$25.00
Fast ModeUS$20.00US$2.00US$25.00US$100.00

Cross-checking the two:

  • OpenRouter’s OpenAI and Azure endpoints match the official Standard price exactly.
  • OpenRouter’s OpenAI Fast matches official Fast Mode exactly.
  • OpenRouter’s OpenAI Flex (US$5 / US$25) matches the official Batch numbers — but the official price list does not list a Flex tier for GPT-6 Astra at all.

We are flagging that last one honestly: the naming does not line up across the two sources, so treat it as unverified. Before you route production traffic there, send one small request and confirm the latency and billing match what you expect.

One more line from the official page worth knowing: all three tiers have a separate long-context price, roughly double the short-context rate. That is the same mechanic we covered in GPT-6 Astra double-bills past 272K — a million-token window is not an invitation to paste your whole repo in.

4. What the cheap tier trades away

The three tiers differ mostly on one axis: how long you are willing to wait.

  • Fast (US$20 / US$100): priority queue, quick responses. Worth it when a human is staring at the screen.
  • Standard (US$10 / US$50): the everyday default.
  • Flex / Batch (US$5 / US$25): let it take its time. Right for bulk cleanup, offline analysis, and overnight automation — nobody is waiting, so nobody should pay for speed.

The rule is simple: someone waiting on screen means the expensive tier; nobody waiting means the cheap one. That single decision cuts most bills roughly in half.

5. The part people miss: the default is not the cheapest

Plenty of people assume OpenRouter automatically picks the lowest price for them. It does not.

OpenRouter’s own documentation describes the default as price-weighted load balancing: selection is weighted by the inverse square of price, with providers that had recent outages deprioritised. The example in the docs is a US$1/M provider receiving 9x the traffic of a US$3/M provider.

So the cheap endpoint gets most of the traffic — not all of it. Your request can still land on an expensive one.

To actually pin the price down, you have to say so in the request. We cover exactly how in How to pin OpenRouter to the cheap provider.

6. Three-line summary

  1. GPT-6 Astra hit OpenRouter on 2026-09-04 with a 1.05M-token context and 128K max output.
  2. Five endpoints for one model, output from US$25 to US$100 per million tokens — a 4x spread; the premium buys speed.
  3. Cheap is not automatic — the default is price-weighted routing, not lowest-price routing.

Official links

More on this site

All figures were read directly from the official pages on 2026-09-05. Prices change; check the official page before you commit.

Let's take a look at these

More verified articles on this tool

Go to the official website

Affiliate Links Notice