90% off cached reads
Shared system prompts and prefixes can be cached automatically. Reads cost $1.00 instead of $10.00. Cache writes cost $12.50, so reuse the prefix across many requests.
Power has a price. Know yours.
The short answer: GPT-6 Astra costs $10.00 per 1M input tokens and $50.00 per 1M output tokens on the standard API tier. That's a 2.5x price jump over GPT-5.6 Sol ($4 / $20). Cached reads cost $1.00, cache writes cost $12.50, and Batch API cuts everything in half.
VERIFIED Official pricing page cross-checked with five independent sources
All prices are per 1M tokens and billed at the chosen model's input/output rates.
| Item | Price per 1M tokens |
|---|---|
| Input | $10.00 |
| Cached input (read) | $1.00 |
| Cache writes | $12.50 |
| Output | $50.00 |
Pricing tables hide the real question: what does a typical request cost?
| Scenario | Input tokens | Output tokens | Cost |
|---|---|---|---|
| Short chat turn | 5,000 | 800 | $0.09 |
| Code review | 50,000 | 5,000 | $0.75 |
| Long document summary | 200,000 | 2,000 | $2.10 |
Both input and output prices are 2.5x higher than GPT-5.6 Sol's current $4 / $20 promotional price.
| Model | Input / 1M | Output / 1M |
|---|---|---|
| GPT-5.6 Sol | $4.00 | $20.00 |
| GPT-6 Astra | $10.00 | $50.00 |
On a blended workload (cache reads + fresh input + output), the effective per-run cost rises roughly 76% - from about $3.08 to $7.70 per 1M tokens. Whether that's worth it depends on what you're building. Review the GPT-6 Astra benchmark results to determine whether the reported capability gains justify the higher price; if your app is price-sensitive and doesn't need the extra capability, Sol remains cheaper.
Use the right billing mode for the workload instead of paying the realtime premium everywhere.
Shared system prompts and prefixes can be cached automatically. Reads cost $1.00 instead of $10.00. Cache writes cost $12.50, so reuse the prefix across many requests.
For bulk generation, backfills, and evaluation runs, Batch API bills at half price. The tradeoff is a 24-hour turnaround instead of instant responses.
Keep requests under 272K input tokens when cost matters. Chunk long inputs or cache aggressively before reaching for the full 1.05M window.
Batch rates: input $5.00 / cached input $0.50 / cache writes $6.25 / output $25.00 per 1M tokens.
Estimate a single request with fresh input, cached reads, cache writes, and output.
Need a capability check? See GPT-6 Astra benchmark results or compare GPT-6 Astra vs Fable 5.1 pricing.
Third-party aggregators may offer Astra alongside other models under one API key. Compare current provider terms before choosing a billing route.
We do not publish an unverified TokenRa price comparison here. For official rates, see the API pricing page. Pricing can change; verify figures before launch.
Short answers for the decisions that affect your bill.
$10.00 per 1M input tokens and $50.00 per 1M output tokens on the standard API tier. Cached reads are $1.00 and cache writes are $12.50.
No. It is 2.5x more expensive across both input and output ($10 / $50 versus $4 / $20).
Requests exceeding 272K input tokens bill at 2x the standard rate. This applies to the long end of Astra's 1.05M-token context window.
Free-tier availability is not confirmed in the sources used for this page. Check the official pricing documentation for the current answer.