AI API Pricing Compared 2026: What Every Major Model Costs per Token

One table, five labs, a 1,000x price spread - and the discounts that matter more than the list prices.

Saganote
Saganote ·
5 Min Read

TL;DR: Frontier AI API pricing currently ranges from about $1 to $50 per million tokens, depending on the model and whether you're paying for input or output. Claude Fable 5.1 remains the most expensive mainstream frontier option at $10/$50, while GPT-5.6 Luna and other efficiency-focused models push costs much lower. Prices verified September 2, 2026.

AI API pricing spans a much wider range than subscription plans suggest. Verified September 2, 2026, frontier models range from low-cost production models to $50-per-million-token output at the top end. The biggest changes are model generations: Anthropic now has Fable 5.1, Opus 5, and Sonnet 5, OpenAI has GPT-5.6 Sol, Terra, and Luna, and xAI has moved to Grok 4.6. Choosing the right model tier can matter more than choosing the provider.

ModelProviderInput ($/MTok)Output ($/MTok)
Claude Fable 5.1Anthropic$10.00$50.00
GPT-5.6 SolOpenAI$5.00$30.00
Claude Opus 5Anthropic$5.00$25.00
GPT-5.6 TerraOpenAI$2.00$12.00
Gemini 3.1 ProGoogle$2.00*$12.00*
Claude Sonnet 5Anthropic$2.00$10.00
Sonar ProPerplexity$3.00$15.00
Gemini 3.5 FlashGoogle$1.50$9.00
Sonar Deep ResearchPerplexity$2.00$8.00
GPT-5.6 LunaOpenAI$1.00$6.00
Grok 4.6xAI$2.00$6.00
Claude Haiku 4.5Anthropic$1.00$5.00
Grok 4.3xAI$1.25$2.50
SonarPerplexity$1.00$1.00

The Workhorse Tier: $6 to $15 Output

Most production applications have more economical choices. GPT-5.6 Terra costs $2/$12, Claude Sonnet 5 costs $2/$10, Gemini 3.5 Flash costs $1.50/$9, and GPT-5.6 Luna costs $1/$6. These models provide a much lower cost base for applications that do not need the maximum capability of frontier-tier models.

The Budget Tier: $1 to $6 Output

The lowest-cost models can dramatically reduce production bills. Perplexity Sonar costs $1/$1, Claude Haiku 4.5 costs $1/$5, GPT-5.6 Luna costs $1/$6, and Grok 4.3 costs $1.25/$2.50 per million input/output tokens. For high-volume applications, these differences can matter more than small benchmark gaps between models.

The Workhorse Tier: $6 to $15 Is Where Production Apps Live

Most real products run here. Terra, Gemini 3.1 Pro, and Sonnet 5 cluster between $10 and $15 output. Luna at $1/$6 and Gemini 3.5 Flash at $1.50/$9 undercut them with near-flagship benchmark scores. Watch one date: Sonnet 5's $2/$10 is introductory and rises to $3/$15 on September 1, a 50% jump for anyone who builds on the intro rate and forgets.

The Budget Tier: Cheaper Than You Think

Haiku 4.5 and Grok 4.3 handle classification, extraction, and routing at $5 and $2.50 output. Sonar Small runs $0.20 flat with live web search included. GPT-5 nano's $0.05 input rate makes bulk embedding-adjacent work nearly free. Teams overpay most often here - routing simple tasks to frontier models out of habit, not need.

The Discounts That Beat the List Prices

Three mechanisms cut these rates harder than switching labs. Batch processing halves OpenAI and Anthropic prices for anything asynchronous. Prompt caching drops repeat-context reads to a tenth or less - Anthropic charges $1 per million to re-read Fable 5 context that costs $10 fresh. And xAI hands out up to $175 a month in free credits through its data-sharing program, with the privacy trade that implies.

Hidden meters run the other way. OpenAI and Anthropic bill web search separately (Anthropic at $10 per 1,000 searches), while Perplexity bakes search into every Sonar token. Add-ons, not token rates, decide many real-world bills.

Which API Should You Build On

AI API pricing rewards a two-model strategy: one budget model for volume, one frontier model for the hard 5% of requests. Cost-sensitive products start with Haiku, Luna, or Flash and escalate selectively. Search products should price Sonar first because the bundled search fee usually wins. Teams locked into one lab for compliance reasons should at least benchmark Grok's $2/$6 to know what leverage they're leaving unused.

AI API pricing changes faster than any page can promise, so treat the verification date as part of the data. Subscription-side numbers live in our full AI pricing guide, with per-provider detail in the Claude, ChatGPT, Gemini, and Grok pricing pages.

Frequently Asked Questions

What is the cheapest AI API in 2026?
GPT-5 nano at $0.05 per million input tokens for bulk work, and Perplexity's Sonar Small at $0.20/$0.20 for search-grounded tasks.
What is the most expensive AI model API?
Anthropic's Fable 5 at $10 per million input tokens and $50 per million output - double the price of GPT-5.6 Sol.
How do batch discounts work?
OpenAI and Anthropic cut token prices 50% for asynchronous jobs submitted in batches, typically returned within 24 hours.
Is the Grok API really cheaper than GPT and Claude?
Yes, at the frontier tier - Grok 4.5 costs $2/$6 per million tokens against Sol's $5/$30 and Fable 5's $10/$50. xAI also offers up to $175 a month in credits for data sharing.
What is prompt caching?
A discount for re-sent context. Cached input costs a tenth or less of fresh input - $1 versus $10 per million on Fable 5.

Share this
Saganote

About Author

Saganote

Saganote is an independent technology publication covering artificial intelligence, cybersecurity, startups, software, consumer technology, and innovation. Our editorial team researches, writes, and reviews original news, analysis, and explainers to provide accurate, timely, and well-sourced coverage of the technology industry.