AI Models & LLMs

Best LLM APIs

An LLM API lets your app send text to a language model and get a reply back, paying per token instead of a monthly seat. This page ranks the main providers developers use in September 2026 on the quality of their best models, price, developer features, reliability and data terms.

All prices are list prices per million tokens (input / output) from each provider's pricing page on 23 September 2026. Quality scores lean on the Artificial Analysis Intelligence Index, an independent benchmark average.

Quick answer

The Anthropic Claude API is the best LLM API for quality, because it serves Claude Opus 5.5, the top-ranked model, at $4 / $20 per million tokens. The OpenAI API has the widest price range, from GPT-6 Luna at $0.10 / $0.50 to GPT-6 Astra. Google's Gemini API has the best free tier. OpenRouter is best for testing many models with one key, and DeepSeek is cheapest.

Top picks at a glance

Scoreboard

Scores are out of 10. The overall score is the weighted average of the criteria below.

#ToolOverallModel qualityPricingDeveloper featuresReliability & reachData & compliancePrice fromBest for
1Anthropic Claude API
Anthropic
8.910.07.59.08.59.0$1 / $5 per 1M tokens (Haiku 4.5)Best models for coding agents and knowledge work
2OpenAI API
OpenAI
8.99.08.59.58.58.5$0.10 / $0.50 per 1M tokens (GPT-6 Luna)Widest model range, from $0.10 budget tier to GPT-6 Astra
3Google Gemini API
Google
8.58.09.38.58.58.0Free tier; paid from $0.10 / $0.40 per 1M tokens
Free tier
Best free tier and cheap multimodal models
4OpenRouter
OpenRouter
8.59.58.08.58.07.0Model price + 5.5% fee on card top-ups
Free tier
One API key for hundreds of models
5SpaceXAI Grok API
SpaceXAI
7.98.28.87.57.56.5$1 / $2 per 1M tokens (grok-build-0.1)Cheap near-frontier coding model
6Mistral API
Mistral AI
7.76.59.07.57.59.0$0.10 / $0.10 per 1M tokens (Ministral 3B)European provider with cheap mid-size models
7Meta Model API
Meta
7.78.39.06.56.56.5$0.10 / $0.20 per 1M tokens (Contributor tier)Fast, cheap Muse Spark reasoning
8DeepSeek API
DeepSeek
7.67.210.07.07.05.0$0.15 / $0.60 per 1M tokens (off-peak)Lowest prices for bulk work

Expert reviews

#1 · Best models for coding agents and knowledge work

Anthropic Claude API

by Anthropic · Usage-based · $1 / $5 per 1M tokens (Haiku 4.5)
8.9/10

The Claude API gives you the best model on the market: Claude Opus 5.5, #1 on the Artificial Analysis Intelligence Index at 58. The lineup is simple. Haiku 4.5 costs $1 / $5 per million tokens, Sonnet 5 costs $2 / $10, Opus 5.5 costs $4 / $20 and Fable 5.1 costs $10 / $50. Every current model except Haiku has a 1M-token context and 128K output.

The developer features are strong: prompt caching (reads cost as little as 2.5% of the input price on Fable), a 50% Batch discount, an effort setting to trade quality for cost, and the same model IDs on AWS, Google Cloud and Microsoft Foundry. A US-only inference option costs 10% extra.

The weak spot is the budget end. Haiku 4.5 is almost a year old, so there is no cheap Claude to rival GPT-6 Luna or Gemini 2.5 Flash-Lite. Pick it if quality on coding and agent tasks matters most. Skip it if you need the lowest cost per token at huge volume.

Score breakdown

Model quality10.0
Pricing7.5
Developer features9.0
Reliability & reach8.5
Data & compliance9.0

Key facts

Pricing
$1 / $5 per 1M tokens (Haiku 4.5) (Sonnet 5 $2 / $10; Opus 5.5 $4 / $20; Fable 5.1 $10 / $50. Batch 50% off. Cache reads 2.5–10% of input price. US-only inference at 1.1x.)
Free option
No
Platforms
API, AWS Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS
Flagship
Claude Opus 5.5 (AA index 58, #1)
Context
1M tokens on Opus 5.5, Fable 5.1 and Sonnet 5
Max output
128K sync; up to 300K on Batch (beta)
Clouds
Claude API, Bedrock, Google Cloud, Microsoft Foundry

What we like

  • Home of the top-ranked model, Claude Opus 5.5
  • Same models on AWS, Google Cloud and Microsoft Foundry
  • Deep caching and Batch discounts
  • Clear model lineup with long retirement notice

Watch out for

  • No modern budget model; Haiku 4.5 dates from October 2025
  • Text and image in, text out only (no native audio or image output)
  • Opus models use many output tokens
#2 · Widest model range, from $0.10 budget tier to GPT-6 Astra

OpenAI API

by OpenAI · Usage-based · $0.10 / $0.50 per 1M tokens (GPT-6 Luna)
8.9/10

The OpenAI API has the best spread of price points. GPT-6 Luna costs $0.10 / $0.50 per million tokens, GPT-6 Sol costs $2 / $10, and GPT-6 Astra costs $10 / $50. All three share a context of about 1M tokens. On the Artificial Analysis index, Astra scores 53 and Sol 48, behind Claude Opus 5.5 but ahead of everything from Google.

It also has the broadest platform: Batch at half price, 90% off cached input, a Fast mode for Astra, and widely used official SDKs. Watch the long-context surcharge. Prompts over the short-context limit roughly double the input price.

Pick it if you want one vendor that covers everything from cheap classification to frontier reasoning. Skip it if you only need the single best model for coding agents, where Claude Opus 5.5 leads at less than half Astra's price.

Score breakdown

Model quality9.0
Pricing8.5
Developer features9.5
Reliability & reach8.5
Data & compliance8.5

Key facts

Pricing
$0.10 / $0.50 per 1M tokens (GPT-6 Luna) (GPT-6 Sol $2 / $10; GPT-6 Astra $10 / $50. Long-context rates higher. Batch 50% off; cached input 90% off.)
Free option
No
Platforms
API, Python SDK, Node SDK
Flagship
GPT-6 Astra (AA index 53)
Value model
GPT-6 Sol, $2 / $10 (AA index 48)
Context
About 1.05M tokens on GPT-6 models
Launched
GPT-6 Astra 3 Sep; Sol and Luna 22 Sep 2026

What we like

  • Three GPT-6 tiers cover every budget
  • GPT-6 Luna is one of the cheapest capable models
  • Mature SDKs, Batch and caching
  • Very large developer ecosystem

Watch out for

  • GPT-6 Astra is the priciest mainstream model
  • Long-context surcharge on all GPT-6 tiers
  • Sol and Luna are brand new, with little independent testing
#3 · Best free tier and cheap multimodal models

Google Gemini API

by Google · Freemium · Free tier; paid from $0.10 / $0.40 per 1M tokens
8.5/10

The Gemini API has the most generous free tier of any major provider, and its paid Flash models are cheap and fast. Gemini 3.8 Flash costs $0.75 in and $3.75 out per million tokens, runs at about 283 tokens per second in Artificial Analysis tests and takes text, images, audio and video. Gemini 2.5 Flash-Lite is the cheapest option at $0.10 / $0.40.

The big gap is the top end. Gemini 3.5 Pro still has not shipped, so Google's best model scores 41 on the Artificial Analysis index, well behind Claude and GPT-6. Plan for price rises too: 3.8 Flash doubles to $1.50 / $7.50 on 1 January 2027.

Pick it if you want to prototype for free, need cheap multimodal input or use Google Cloud. Skip it if you need frontier reasoning or coding quality.

Score breakdown

Model quality8.0
Pricing9.3
Developer features8.5
Reliability & reach8.5
Data & compliance8.0

Key facts

Pricing
Free tier; paid from $0.10 / $0.40 per 1M tokens (Gemini 3.8 Flash $0.75 / $3.75 until 31 Dec 2026, then $1.50 / $7.50. Gemini 3.1 Pro Preview $2 / $12. Batch and Flex 50% off.)
Free option
Yes
Platforms
API, Google AI Studio, Vertex AI
Best model
Gemini 3.8 Flash (AA index 41)
Free tier
Yes, on most Flash models
Search grounding
5,000 free requests/month on Gemini 3.x
Context
1M tokens

What we like

  • Free tier on most models
  • Cheap, fast Flash models with 1M context
  • Native image, audio and video input
  • 5,000 free Google Search grounding calls a month

Watch out for

  • No frontier-class model while Gemini 3.5 Pro is delayed
  • Gemini 3.8 Flash price doubles on 1 January 2027
  • Check free-tier data-use terms before sending private data
#4 · One API key for hundreds of models

OpenRouter

by OpenRouter · Usage-based · Model price + 5.5% fee on card top-ups
8.5/10

OpenRouter is not a model maker. It is a router: one API key and one OpenAI-style endpoint for 455 models from Anthropic, OpenAI, Google, Meta, SpaceXAI (formerly xAI), DeepSeek, Z.ai, Moonshot and dozens of hosts. It usually passes through the vendor's list price and charges a 5.5% fee when you buy credits by card.

That makes it the easiest way to compare models on your own prompts, switch when a new one launches, or fall back to another provider during an outage. New models appear fast: Claude Opus 5.5, GPT-6 Sol and Grok 4.7 were all listed within a day of launch. It is also an easy way to reach open models such as GLM-5.3 and Kimi K3 without running servers.

Pick it if you want flexibility, quick testing or access to open models without running servers. Skip it if you need a direct contract, strict data-residency terms or the lowest possible cost at very high volume.

Score breakdown

Model quality9.5
Pricing8.0
Developer features8.5
Reliability & reach8.0
Data & compliance7.0

Key facts

Pricing
Model price + 5.5% fee on card top-ups (Free models: 50 requests/day, or 1,000/day after buying $10 of credit. Bring-your-own-key free up to $25,000/month, then 5%.)
Free option
Yes
Platforms
API, OpenAI-compatible SDKs
Models listed
455 (API count, 23 Sep 2026)
Credit fee
5.5% ($0.80 minimum) by card; 5% crypto
Free models
Yes, rate-limited
Failed requests
Not billed

What we like

  • Hundreds of models behind one key
  • New models listed within a day of launch
  • Automatic fallback across providers
  • Free, rate-limited models for testing

Watch out for

  • 5.5% fee on credit purchases
  • An extra company in your data path
  • Features can lag the vendor's own API
#5 · Cheap near-frontier coding model

SpaceXAI Grok API

by SpaceXAI · Usage-based · $1 / $2 per 1M tokens (grok-build-0.1)
7.9/10

SpaceXAI, the company formerly called xAI, sells Grok models through its API at low prices. The new Grok 4.7 costs $2 in and $6 out per million tokens for prompts under 200K tokens, and scores 46 on the Artificial Analysis index, close to GPT-6 Sol at a lower output price. SpaceXAI reports 71.0% on DeepSWE v1.1.

Speed is a plus: Artificial Analysis measured Grok 4.7 at about 188 tokens per second. The trade-offs are token use and context. It can use many reasoning tokens on complex tasks, which eats into the low price, and it has no Batch API. Its 500K context is half of what Claude, OpenAI and Google offer, although the older Grok 4.3 has 1M tokens at $1.25 / $2.50.

Pick it if you want a low-cost model for long coding sessions. Skip it if you need 1M-token context on the flagship, batch discounts or a long enterprise track record.

Score breakdown

Model quality8.2
Pricing8.8
Developer features7.5
Reliability & reach7.5
Data & compliance6.5

Key facts

Pricing
$1 / $2 per 1M tokens (grok-build-0.1) (Grok 4.7 $2 / $6 under 200K tokens, $4 / $12 above. Fast variant 2x price for 2x speed.)
Free option
No
Platforms
API, Cursor, Cloud marketplaces
Flagship
Grok 4.7, released 21 Sep 2026 (AA index 46)
Context
500K (Grok 4.5–4.7); 1M on Grok 4.3
Cached input
$0.50 per 1M tokens (Grok 4.7)

What we like

  • Near-frontier Grok 4.7 at $2 / $6
  • Cheaper older models with 1M context
  • Day-one support in Cursor
  • Fast output, about 188 tokens per second (Artificial Analysis)

Watch out for

  • No Batch API for Grok 4.7
  • 500K context on current flagships
  • Heavy reasoning-token use raises real costs
#6 · European provider with cheap mid-size models

Mistral API

by Mistral AI · Usage-based · $0.10 / $0.10 per 1M tokens (Ministral 3B)
7.7/10

Mistral is the main European option. Its API prices are low across the range: Ministral 3 models from $0.10 per million tokens, Mistral Small 4 at $0.15 / $0.60, and Mistral Large 3 at $0.50 / $1.50. Cached input is 90% off and Batch is half price. Several models, including Mistral Small 4, are also released as open weights, so you can move them in-house later.

The weakness is quality at the top. Mistral Medium 3.5 scores 14 on the Artificial Analysis index, far behind US and Chinese flagships, so Mistral is not the choice for hard reasoning or agentic coding.

Pick it if you want a European vendor, cheap models for routine tasks, or an easy path from API to self-hosting. Skip it if you need frontier-level answers.

Score breakdown

Model quality6.5
Pricing9.0
Developer features7.5
Reliability & reach7.5
Data & compliance9.0

Key facts

Pricing
$0.10 / $0.10 per 1M tokens (Ministral 3B) (Mistral Small 4 $0.15 / $0.60; Mistral Large 3 $0.50 / $1.50; Mistral Medium 3.5 $1.50 / $7.50. Cached input 90% off; Batch half price.)
Free option
No
Platforms
API, Mistral Vibe (formerly Le Chat), Self-hosted
Headquarters
Paris, France
Cheapest model
Ministral 3 3B, $0.10 / $0.10
Open weights
Mistral Small 4 (Apache 2.0) and others
Embeddings
Mistral Embed $0.10; Codestral Embed $0.15

What we like

  • European company, useful for EU procurement
  • Low prices across the lineup
  • Many models also available as open weights
  • Cheap embedding models

Watch out for

  • Well behind the frontier on reasoning and coding
  • Smaller ecosystem than OpenAI, Anthropic or Google
#7 · Fast, cheap Muse Spark reasoning

Meta Model API

by Meta · Usage-based · $0.10 / $0.20 per 1M tokens (Contributor tier)
7.7/10

Meta's Model API sells its closed Muse Spark models. Muse Spark 1.3 scores 48 on the Artificial Analysis index, level with GPT-6 Sol, and runs at about 213 tokens per second, one of the fastest results for a model this capable. The standard price is $1.25 in and $4.25 out per million tokens.

Meta also offers a "Contributor" tier at just $0.10 / $0.20. The catch: Meta can use that traffic to improve its products. That is fine for public data or testing, not for customer or company secrets.

The platform is young. It has fewer SDKs, integrations and enterprise controls than OpenAI, Anthropic or Google. Pick it if you need fast, cheap, near-frontier answers and can choose the right tier for your data. Skip it if you need mature enterprise tooling or strict data guarantees.

Score breakdown

Model quality8.3
Pricing9.0
Developer features6.5
Reliability & reach6.5
Data & compliance6.5

Key facts

Pricing
$0.10 / $0.20 per 1M tokens (Contributor tier) (Standard tier $1.25 / $4.25 for Muse Spark 1.3; cached input $0.15. Contributor traffic may be used by Meta to improve its products.)
Free option
No
Platforms
API, Muse Code
Flagship
Muse Spark 1.3, released 2 Sep 2026 (AA index 48)
Speed
About 213 tokens/s (Artificial Analysis)
Context
1,048,576 tokens

What we like

  • Near-frontier quality at a low price
  • Very fast output
  • Ultra-cheap Contributor tier for non-sensitive work

Watch out for

  • Contributor tier lets Meta use your traffic
  • Young platform with limited tooling
  • Fewer published coding results than rivals
#8 · Lowest prices for bulk work

DeepSeek API

by DeepSeek · Usage-based · $0.15 / $0.60 per 1M tokens (off-peak)
7.6/10

DeepSeek's API is the cheapest serious option. DeepSeek-Flash, which runs DeepSeek V4.1 Flash, costs $0.30 in and $1.20 out per million tokens at peak and half that off-peak. Cached input can cost as little as $0.003. It supports a 1M-token context, up to 384K tokens of output, images, JSON output and tool calls, and it uses an OpenAI-compatible format.

Quality is solid, not frontier: V4.1 Flash scores 39 on the Artificial Analysis index. The bigger worry for many firms is compliance, because the service is run from China. Because the weights are MIT-licensed, you can get the same model from US hosts or run it yourself.

Pick it if you run huge volumes of routine work such as extraction, tagging or first drafts. Skip it if your data rules forbid China-based processors; use the open weights through another host instead.

Score breakdown

Model quality7.2
Pricing10.0
Developer features7.0
Reliability & reach7.0
Data & compliance5.0

Key facts

Pricing
$0.15 / $0.60 per 1M tokens (off-peak) (DeepSeek-Flash peak $0.30 / $1.20; cache hits from $0.003. V4-Pro $0.66–$1.32 in, $1.98–$3.96 out. Off-peak is half price.)
Free option
No
Platforms
API, OpenAI-compatible SDKs
Models
DeepSeek-Flash (V4.1 Flash) and DeepSeek-V4-Pro
Context
1M tokens; up to 384K output
Peak hours (UTC)
01:00–04:00 and 06:00–10:00 on weekdays
Concurrency
2,500 requests (Flash), 500 (Pro)

What we like

  • Lowest prices of any major API
  • Off-peak and cache discounts
  • 1M context and 384K output
  • Same model available as MIT-licensed open weights

Watch out for

  • China-based service may fail compliance checks
  • Well below frontier quality
  • Peak-hour prices double

How we scored these tools

Each tool is scored 0–10 on the criteria below, using public evidence: independent benchmarks, vendor documentation and pricing pages, aggregate user ratings and reputable reviews. The overall score is the weighted average. Nobody pays to be listed. Read our full methodology.

CriterionWeightWhat we look at
Model quality30%How good the provider's best available models are, based on the Artificial Analysis Intelligence Index and published benchmarks.
Pricing25%List prices across the lineup, plus caching, batch and off-peak discounts.
Developer features20%Context length, tool use, structured output, batch APIs, caching, SDKs and multimodal support.
Reliability & reach15%Track record, speed, rate limits and availability on major clouds.
Data & compliance10%Data handling options, regional processing and suitability for regulated industries.

Price comparison: flagship and budget models

Provider Top model (price per 1M in / out) Cheapest model AA index of top model
Anthropic Opus 5.5: $4 / $20 Haiku 4.5: $1 / $5 58
OpenAI GPT-6 Astra: $10 / $50 GPT-6 Luna: $0.10 / $0.50 53
Meta Muse Spark 1.3: $1.25 / $4.25 Contributor tier: $0.10 / $0.20 48
SpaceXAI Grok 4.7: $2 / $6 grok-build-0.1: $1 / $2 46
Google Gemini 3.8 Flash: $0.75 / $3.75 2.5 Flash-Lite: $0.10 / $0.40 41
DeepSeek V4-Pro: $1.32 / $3.96 (peak) Flash: $0.15 / $0.60 (off-peak) 39 (Flash)
Mistral Medium 3.5: $1.50 / $7.50 Ministral 3 3B: $0.10 / $0.10 14 (Medium 3.5)

OpenRouter resells most of these at list price plus a 5.5% card top-up fee.

How to cut your API bill

  • Cache your prompts. Put fixed instructions and documents first. Cached reads are 90% off at OpenAI and Mistral, 95% off on Claude Opus 5.5 and 97.5% off on Claude Fable 5.1.
  • Batch anything that can wait. Anthropic, OpenAI, Google and Mistral all charge half price for batch jobs.
  • Route by difficulty. Send easy tasks to a cheap model (GPT-6 Luna, Gemini Flash, DeepSeek) and only hard ones to a flagship.
  • Lower the effort setting. Reasoning models bill their hidden thinking as output. Lower effort often gives the same answer for far fewer tokens.
  • Watch context surcharges. OpenAI charges more above its short-context limit, Grok above 200K tokens.

How to choose

  1. Define the task and quality bar. Write 20–50 test prompts with good answers.
  2. Test three or four models through OpenRouter or free tiers. Measure accuracy, speed and cost per task, not just price per token.
  3. Check data terms. Regulated industries may need a provider on their existing cloud (Claude on AWS, Google or Microsoft), US-only processing, or a self-hosted open model.
  4. Plan for change. Models update monthly. Keep your code provider-agnostic, for example with an OpenAI-compatible client, so you can switch quickly.

For the models themselves, see best AI models. For self-hosting, see best open-source LLMs.

Expert tips
  1. Prototype on the Gemini API free tier or OpenRouter's free models, then move to paid tiers once you know which model wins on your prompts.
  2. Structure prompts as fixed instructions first, changing user input last. That single change can cut input costs by up to 90% with caching.
  3. Schedule DeepSeek batch work outside 01:00–04:00 and 06:00–10:00 UTC on weekdays to get the half-price off-peak rate.
  4. Keep sensitive data off Meta's Contributor tier and any free tier whose terms allow training on inputs.
  5. If you build on Gemini 3.8 Flash, budget now for the jump to $1.50 / $7.50 per million tokens on 1 January 2027.

Jargon explained

API
A way for one program to talk to another. An LLM API lets your app send a prompt to a model over the internet and get the answer back.
Price per million tokens
How APIs bill. A token is about three-quarters of a word; you pay separately for input you send and output the model writes.
Prompt caching
The provider stores the start of a prompt you send often. Re-sending it later costs a fraction of the normal price.
Batch API
A way to send many requests that finish within hours instead of seconds, usually at half price.
Rate limit
The maximum number of requests or tokens you can send per minute or day on your account.

Frequently asked questions

What is the best LLM API?

For quality, the Anthropic Claude API, because it serves Claude Opus 5.5, the top-ranked model as of 23 September 2026. For range and ecosystem, the OpenAI API. For free usage, the Google Gemini API.

What is the cheapest LLM API?

DeepSeek's API: DeepSeek-Flash costs $0.15 / $0.60 per million tokens off-peak, with cached input from $0.003. GPT-6 Luna ($0.10 / $0.50) and Gemini 2.5 Flash-Lite ($0.10 / $0.40) are close and run by US companies.

Is there a free LLM API?

Yes. Google's Gemini API has a free tier on most Flash models. OpenRouter offers free, rate-limited models (50 requests a day, or 1,000 after buying $10 of credit). Mistral's free Vibe plan (formerly Le Chat) includes $10 a month of API credits.

Should I use OpenRouter or go direct?

Use OpenRouter to test and compare models quickly, or to get automatic fallbacks. Go direct once you settle on a provider at high volume, to avoid the 5.5% credit fee and get vendor-specific features and contracts.

Can I use Claude through AWS or Google Cloud?

Yes. Claude models are available on Amazon Bedrock, Google Cloud and Microsoft Foundry, as well as Anthropic's own API. This lets you use existing cloud billing and data agreements.

How much does it cost to run a chatbot on an LLM API?

It depends on traffic and model. As a rough example, one million short chats using 1,000 input and 500 output tokens each would cost about $7,000 on GPT-6 Sol at list price before caching ($2,000 input + $5,000 output), or about $350 on GPT-6 Luna. Always measure on your own traffic.

Sources

Every fact on this page comes from public information. Vendor figures are labelled as vendor claims.

  1. Claude models overview and pricing (Anthropic)
  2. Claude Opus product page (Anthropic)
  3. OpenAI API pricing (OpenAI)
  4. OpenAI releases GPT-6 Sol and Luna (MarkTechPost)
  5. Gemini API pricing (Google)
  6. OpenRouter FAQ (OpenRouter)
  7. OpenRouter models (OpenRouter)
  8. Grok models and pricing (SpaceXAI)
  9. Introducing Grok 4.7 (SpaceXAI)
  10. Mistral API pricing (Mistral AI)
  11. Muse Spark 1.3 API pricing (OpenRouter)
  12. DeepSeek API models and pricing (DeepSeek)
  13. LLM Leaderboard (Artificial Analysis)
  14. xAI is becoming SpaceXAI (The Verge)
  15. Vibe gets to work (Mistral AI)