AI Models & LLMs

Best AI Models

Also known as: AI model

September 2026 was the busiest month for new AI models in a year. OpenAI shipped three GPT-6 models, Anthropic shipped Claude Fable 5.1 and then Claude Opus 5.5, Meta shipped Muse Spark 1.3, and SpaceXAI (formerly xAI) shipped Grok 4.7. Open-weight labs in China (Z.ai, Alibaba, Moonshot, DeepSeek) are now only a few points behind the closed leaders.

We ranked the 11 best models you can actually use today. We scored each one on independent tests (mainly the Artificial Analysis Intelligence Index and the Arena human-vote leaderboard), coding and agent results, API price, speed and context window, and how easy it is to get access. Prices are per million tokens (about 750,000 words of input) and are correct as of 23 September 2026. See how we rank for our method.

Quick answer

Claude Opus 5.5 is the best AI model right now. It tops the Artificial Analysis Intelligence Index (58) and costs $4/$20 per million tokens, less than half the price of GPT-6 Astra or Claude Fable 5.1. Muse Spark 1.3 is our runner-up because it matches GPT-6 Sol on the index for less money and runs faster. Pick GPT-6 Sol for strong results at $2/$10, Gemini 3.8 Flash for speed and low cost, and GLM-5.3 or DeepSeek V4 if you need open weights you can run yourself.

Top picks at a glance

Scoreboard

Scores are out of 10. The overall score is the weighted average of the criteria below.

#ToolOverallIntelligence & reasoningCoding & agentsPrice-performanceSpeed & contextAvailability & opennessPrice fromBest for
1Claude Opus 5.5
Anthropic
9.09.89.68.08.08.0$4 / $20 per 1M tokensThe strongest all-round model for coding, agents and knowledge work
2Muse Spark 1.3
Meta
8.38.67.89.29.07.0$1.25 / $4.25 per 1M tokens
Free tier
Fast, cheap frontier-class answers and multimodal agents
3GPT-6 Astra
OpenAI
8.39.39.26.07.57.0$10 / $50 per 1M tokensHard reasoning, research and computer use inside the OpenAI ecosystem
4GPT-6 Sol
OpenAI
8.28.68.49.07.07.0$2 / $10 per 1M tokensCoding and agent work on a budget
5Claude Fable 5.1
Anthropic
8.29.39.35.56.57.5$10 / $50 per 1M tokensVery long agent runs and the hardest reasoning tasks
6GLM-5.3
Z.ai (Zhipu AI)
8.18.08.09.06.59.0$1.40 / $4.40 per 1M tokensThe smartest open-weight model for coding and security work
7Gemini 3.8 Flash
Google
8.17.47.59.59.58.5$0.75 / $3.75 per 1M tokens
Free tier
High-volume apps that need speed and low cost
8Grok 4.7
SpaceXAI (formerly xAI)
8.08.27.88.57.57.5$2 / $6 per 1M tokensLong-running agent tasks at a mid-range price
9Qwen3.8 Max
Alibaba
7.98.08.08.85.58.0$2 / $6 per 1M tokensMultimodal coding and document work with an open-weight fallback
10DeepSeek V4
DeepSeek
7.86.87.29.87.59.5$0.15 / $0.60 per 1M tokens (V4.1 Flash, off-peak)The cheapest capable model, with MIT-licensed weights
11Kimi K3
Moonshot AI
7.67.88.07.05.58.5$3 / $15 per 1M tokensOpen-weight agents that browse, code and use a computer

Expert reviews

#1 · The strongest all-round model for coding, agents and knowledge work

Claude Opus 5.5

by Anthropic · Usage-based · $4 / $20 per 1M tokens
9.0/10

Opus 5.5 is the model to beat. One day after launch it sits at the top of the Artificial Analysis Intelligence Index with 58 points at max effort, five points clear of GPT-6 Astra and Claude Fable 5.1. Even at high effort it scores 54, which still beats every non-Anthropic model.

The price is the surprise. At $4 input and $20 output per million tokens it costs 20% less per token than Opus 5 did, and Anthropic says it costs about 40% less to run on typical work because it uses fewer tokens. Fable 5.1 and GPT-6 Astra both cost $10/$50, so Opus 5.5 gives you more intelligence for well under half the price.

Anthropic says it performs at the level of Fable 5.1 on most work and matches Fable and Mythos in biology and cybersecurity. Its own coding numbers back that up: 66.4% on Terminal-Bench 4.0 versus 55.8% for Fable 5.1. It also has a 1M-token context window and 128K output.

The catch: it is brand new, so independent coding tests and Arena votes will take a few weeks to settle. For now, it is our default pick for developers and heavy users of the Claude app.

Score breakdown

Intelligence & reasoning9.8
Coding & agents9.6
Price-performance8.0
Speed & context8.0
Availability & openness8.0

Key facts

Pricing
$4 / $20 per 1M tokens (Cache reads $0.20/M; fast mode $8/$40. Also in the Claude apps.)
Free option
No
Platforms
Web, iOS, Android, Desktop, API, AWS, Google Cloud, Microsoft Azure
Released
22 September 2026
AA Intelligence Index
58 (max effort), #1
Context / max output
1M / 128K tokens
Terminal-Bench 4.0
66.4% (Anthropic-reported)
Humanity's Last Exam
67.7% with tools (Anthropic-reported)

What we like

  • Highest score on the Artificial Analysis Intelligence Index (58)
  • Costs $4/$20, far less than Astra or Fable 5.1 at $10/$50
  • 1M-token context and 128K output
  • Available on the Claude API, AWS, Google Cloud and Azure from day one

Watch out for

  • Too new for Arena human-vote rankings
  • Most benchmark numbers so far come from Anthropic itself
  • No free API tier
#2 · Fast, cheap frontier-class answers and multimodal agents

Muse Spark 1.3

by Meta · Usage-based · $1.25 / $4.25 per 1M tokens
8.3/10

Muse Spark 1.3 is the model that put Meta back in the frontier race. It scores 48 on the Artificial Analysis Intelligence Index, the same as GPT-6 Sol, and it is far faster: about 213 tokens per second at max effort, roughly twice the speed of GPT-6 Sol. On the Arena human-vote board it scores 1493, close to Claude Opus 5.

Pricing is low. The standard API tier costs $1.25 input and $4.25 output per million tokens. A Contributor tier drops that to $0.10/$0.20. Launch coverage reports that Meta can use Contributor traffic to train its products, so read the terms and keep private data off it. The Meta Model API is compatible with the OpenAI SDK and has built-in web search with citations.

Meta reports 75.4% on DeepSWE 1.1 and 98.5% on a long-context retrieval test inside its 1M-token window. Those are Meta's numbers, and independent coding results still trail the Claude and GPT-6 leaders.

Muse Spark is closed, unlike Meta's older Llama models. Bloomberg reported that Meta is rolling it out to its free Meta AI assistant and its social apps; developers get it through Muse Code and the Meta Model API. For developers who want speed and low cost with near-frontier smarts, it is a strong choice. On our weights its price and speed lift it to second place overall, even though GPT-6 Astra and Claude Fable 5.1 score higher on intelligence.

Score breakdown

Intelligence & reasoning8.6
Coding & agents7.8
Price-performance9.2
Speed & context9.0
Availability & openness7.0

Key facts

Pricing
$1.25 / $4.25 per 1M tokens (Contributor tier $0.10/$0.20; launch coverage says Meta can use that traffic for training. Meta AI chat is free.)
Free option
Yes
Platforms
Meta AI, Muse Code, API
Released
2 September 2026
AA Intelligence Index
48 (max effort), 45 (xhigh)
Output speed
About 213 tokens/s at max effort (AA)
Arena text score
1493 (13 Sept 2026)
Context window
1M tokens

What we like

  • Ties GPT-6 Sol on the Intelligence Index (48)
  • About 213 tokens per second, much faster than rivals at this level
  • $1.25/$4.25 standard pricing
  • 1M-token context and built-in web search on the API

Watch out for

  • Closed weights, unlike older Llama models
  • Cheapest tier reportedly lets Meta use your data
  • Coding claims are mostly Meta's own so far
#3 · Hard reasoning, research and computer use inside the OpenAI ecosystem

GPT-6 Astra

by OpenAI · Usage-based · $10 / $50 per 1M tokens
8.3/10

GPT-6 Astra is OpenAI's top model and the second-smartest model you can use today. It scores 53 on the Artificial Analysis Intelligence Index at max effort, tied with Claude Fable 5.1, and 52 at xhigh. It handles text and images, reads up to 1.05M tokens, and writes up to 128K tokens in one reply. Through the API it works with web search, code interpreter, computer use and MCP tools.

The weak spot is cost. At $10 input and $50 output per million tokens it costs 2.5 times as much as Claude Opus 5.5, which now scores higher. Long prompts get pricier still: past 272K input tokens, input costs double and output costs 1.5 times as much.

Access is also narrower than the name suggests. In regular ChatGPT it appears as GPT-6 Pro and is limited to the Pro plans (from $100; the $200 tier has been closed to new sign-ups since 10 September 2026) plus Business and Enterprise. Plus users only get it inside ChatGPT Work and Codex, and Free and Go users do not get it. OpenAI rolled it out gradually, so some eligible accounts waited.

Pick Astra if your team already builds on OpenAI tools and needs the top tier. Otherwise GPT-6 Sol gives most of the quality for a fifth of the price.

Score breakdown

Intelligence & reasoning9.3
Coding & agents9.2
Price-performance6.0
Speed & context7.5
Availability & openness7.0

Key facts

Pricing
$10 / $50 per 1M tokens (Cached input $1/M. Prompts over 272K tokens cost 2x input and 1.5x output. In ChatGPT Pro, Business and Enterprise as GPT-6 Pro.)
Free option
No
Platforms
Web, iOS, Android, Desktop, API, AWS, Microsoft Azure
Released
3 September 2026
AA Intelligence Index
53 (max effort)
Context / max output
1.05M / 128K tokens
Knowledge cutoff
30 April 2026
Reasoning effort levels
low, medium, high, xhigh, max

What we like

  • Ties for the second-highest Intelligence Index score (53)
  • 1.05M-token context and 128K output
  • Full OpenAI tool stack: web search, code interpreter, computer use, MCP
  • Five effort levels let you trade cost for quality

Watch out for

  • $10/$50 is 2.5x the price of Claude Opus 5.5, which scores higher
  • Long-context surcharge above 272K tokens
  • Not available to ChatGPT Free, Go, or Plus users in regular chat
#4 · Coding and agent work on a budget

GPT-6 Sol

by OpenAI · Usage-based · $2 / $10 per 1M tokens
8.2/10

GPT-6 Sol is the best deal from OpenAI. It sits one step below Astra, uses training methods developed for Astra, and costs $2 input and $10 output per million tokens, a fifth of Astra's price and about half of GPT-5.6 Sol's.

On the Artificial Analysis Intelligence Index it scores 48 at max effort, level with Meta's Muse Spark 1.3 and one point above GPT-5.6 Sol. At lower effort levels the score falls quickly (44 at xhigh, 43 at high), so use max effort for hard problems. OpenAI says Sol at xhigh effort scored 33.2% on AutomationBench against 26.9% for Claude Opus 5, at a much lower cost per task.

The context window is 1.05M tokens, though input tops out at 922K, and output can reach 128K. Through the Responses API it gets web search, file search, code interpreter and computer use.

The odd part is access. At launch Sol was in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users, but not in the regular ChatGPT chat window. For developers that doesn't matter. For API coding agents where Opus 5.5 is too expensive, Sol is our pick. Its sibling GPT-6 Luna ($0.10/$0.50) handles simple, high-volume jobs.

Score breakdown

Intelligence & reasoning8.6
Coding & agents8.4
Price-performance9.0
Speed & context7.0
Availability & openness7.0

Key facts

Pricing
$2 / $10 per 1M tokens (Cached input $0.20/M. Input costs 2x above 272K tokens. In ChatGPT Work and Codex on Plus and higher.)
Free option
No
Platforms
Web, Desktop, API
Released
22 September 2026
AA Intelligence Index
48 (max effort), 44 (xhigh)
Context / max output
1.05M (922K input) / 128K tokens
Knowledge cutoff
20 April 2026
Price change
About 50% cheaper than GPT-5.6 Sol

What we like

  • Scores 48 on the Intelligence Index for $2/$10
  • Effort levels from none to max let you control cost
  • 1.05M-token context and 128K output
  • Full OpenAI tool support in the Responses API

Watch out for

  • Quality drops noticeably below max effort
  • Not in the regular ChatGPT chat window at launch
  • Input price doubles past 272K tokens
#5 · Very long agent runs and the hardest reasoning tasks

Claude Fable 5.1

by Anthropic · Usage-based · $10 / $50 per 1M tokens
8.2/10

Fable 5.1 is Anthropic's largest model and, until Opus 5.5 arrived three weeks later, was the top model on most leaderboards. It still scores 53 on the Artificial Analysis Intelligence Index, level with GPT-6 Astra. On the Arena human-vote board, Fable models hold the top spot (Fable 5 high at 1506) and Fable 5.1 max sits at 1498.

Anthropic reports 60.9% on Humanity's Last Exam without tools and 77.9% on OSWorld 2.0 (partial credit), both clear gains over Fable 5 and Opus 5. Anthropic still recommends it for the most demanding reasoning and long-horizon agent work, or when Opus 5.5 at high effort falls short on your own tests.

Fable 5.1 costs the same $10/$50 as Fable 5, but cache reads fell 75% to $0.25 per million, and Anthropic estimates typical workloads cost about 25% less. It is also slower than Opus 5.5 (about 61 to 66 tokens per second at top effort).

The honest take: Opus 5.5 now scores higher for 40% of the price, so most people should switch. Keep Fable 5.1 for jobs where you have proof it does better. Its twin, Claude Mythos 5.1, is the same model with looser safeguards, open only to vetted US security and life-science groups.

Score breakdown

Intelligence & reasoning9.3
Coding & agents9.3
Price-performance5.5
Speed & context6.5
Availability & openness7.5

Key facts

Pricing
$10 / $50 per 1M tokens (Cache reads $0.25/M (75% cut vs Fable 5). In paid Claude plans.)
Free option
No
Platforms
Web, iOS, Android, Desktop, API, AWS, Google Cloud, Microsoft Azure
Released
1 September 2026
AA Intelligence Index
53 (max and xhigh effort)
Arena text score
1498 (Fable 5.1 max, 13 Sept 2026)
Terminal-Bench 4.0
55.8% (Anthropic-reported)
Context / max output
1M / 128K tokens

What we like

  • Tied second on the Intelligence Index (53) and near the top of Arena
  • Strong on long, multi-hour agent tasks
  • Cache reads now $0.25/M, which cuts costs for repeated context
  • Available on every major cloud

Watch out for

  • Opus 5.5 scores higher at 40% of the price
  • Slower than Opus 5.5
  • $10/$50 list price is among the highest here
#6 · The smartest open-weight model for coding and security work

GLM-5.3

by Z.ai (Zhipu AI) · Open source · $1.40 / $4.40 per 1M tokens
8.1/10

GLM-5.3 is the second-highest open-weight model on the Artificial Analysis Intelligence Index, scoring 45 at max effort (Xiaomi's MiMo-V2.6-Pro scores 46), one point behind Grok 4.7 and level with Qwen3.8 Max. You can download the 753-billion-parameter weights from Hugging Face or pay Z.ai $1.40 input and $4.40 output per million tokens.

Z.ai built GLM-5.3 by extending post-training on the GLM-5.2 base. Its main gains are coding and cybersecurity. Z.ai reports 66.9% on DeepSWE 1.1 and 84.5% on CyberGym, a vulnerability-finding test. Those security skills are why Z.ai held back the weights for about two weeks for a safety review.

The model has a 1M-token context window and three reasoning levels (low, high, max). Speed is middling at about 61 tokens per second on Artificial Analysis. On Arena it scores 1483, level with GPT-5.6 Sol.

The weights use a custom GLM-5.3 license, so read it before shipping a product. Self-hosting a 753B model also needs a multi-GPU server. For most teams, the cheap API is the easy route, and running it yourself is the option when data can't leave your servers.

Score breakdown

Intelligence & reasoning8.0
Coding & agents8.0
Price-performance9.0
Speed & context6.5
Availability & openness9.0

Key facts

Pricing
$1.40 / $4.40 per 1M tokens (Cached input $0.26/M. Weights free to download on Hugging Face.)
Free option
No
Platforms
API, Web, Self-hosted
Released
August 2026 (weights after a two-week safety review)
AA Intelligence Index
45 (max), second among open-weight models
Parameters
753B
Context window
1M tokens
Arena text score
1483 (13 Sept 2026)

What we like

  • Among the highest-scoring open-weight models on the Intelligence Index (45)
  • Downloadable weights for self-hosting
  • $1.40/$4.40 API price
  • Strong vendor-reported coding and security scores

Watch out for

  • Custom license, not MIT or Apache
  • 753B parameters need serious hardware to self-host
  • About 61 tokens per second, slower than US leaders
#7 · High-volume apps that need speed and low cost

Gemini 3.8 Flash

by Google · Freemium · $0.75 / $3.75 per 1M tokens
8.1/10

Google's best model you can use today is not a Pro model. Gemini 3.5 Pro was promised for June and has slipped repeatedly. Google now says it is testing with partners, with no public date. That leaves Gemini 3.8 Flash as the flagship.

It is a very good flagship for a Flash model. It scores 41 on the Artificial Analysis Intelligence Index, well below the leaders, but on the Arena human-vote board it scores 1493, level with Muse Spark 1.3 and Claude Opus 5. People like its answers more than its test scores suggest.

Speed and price are where it wins. It outputs about 283 tokens per second, the fastest model in this list. Until 31 December 2026 it costs $0.75 input and $3.75 output per million tokens. From 1 January 2027 that doubles to $1.50/$7.50, so budget for the rise. The Gemini API also has a free tier.

It reads text, images, audio, video and PDFs across a 1M-token window, but output caps at about 65K tokens, half of what Claude and GPT-6 allow. Use it for chat apps, document processing and fast agents. Pick Opus 5.5 or Astra for the hardest reasoning.

Score breakdown

Intelligence & reasoning7.4
Coding & agents7.5
Price-performance9.5
Speed & context9.5
Availability & openness8.5

Key facts

Pricing
$0.75 / $3.75 per 1M tokens (Intro price until 31 Dec 2026, then $1.50/$7.50. Free API tier available.)
Free option
Yes
Platforms
API, Google AI Studio, Vertex AI, Antigravity, Android Studio
Generally available
2 September 2026
AA Intelligence Index
41 (high)
Output speed
About 283 tokens/s (AA)
Context / max output
1,048,576 / 65,536 tokens
Arena text score
1493 (13 Sept 2026)

What we like

  • Fastest model here at about 283 tokens per second
  • Free API tier and $0.75/$3.75 intro pricing
  • Takes text, image, audio, video and PDF input
  • Arena score of 1493 is close to the top models

Watch out for

  • Intelligence Index of 41 trails the leaders by 17 points
  • Price doubles on 1 January 2027
  • 65K output cap is half of Claude's and GPT-6's
#8 · Long-running agent tasks at a mid-range price

Grok 4.7

by SpaceXAI (formerly xAI) · Usage-based · $2 / $6 per 1M tokens
8.0/10

Grok 4.7 is a solid step up from Grok 4.6 at the same price. SpaceXAI built it on a larger base model and trained it longer on tasks that take hours, not single replies. It is also meant to check its own work more often before answering.

It scores 46 on the Artificial Analysis Intelligence Index, level with the top open-weight model (MiMo-V2.6-Pro) and two points below GPT-6 Sol. The biggest gain is in agentic coding: SpaceXAI reports 37.6% on Terminal-Bench 4.0, up from 20.3% for Grok 4.6, though independent testing by Artificial Analysis found 26%, per The Decoder. That is still far behind Claude Opus 5.5's 66.4% on the same test.

The price is $2 input and $6 output per million tokens for prompts under 200K tokens, doubling above that. That undercuts GPT-6 Sol on output. The context window is 500K tokens, half of most rivals. Speed is a strength: Artificial Analysis measured about 188 tokens per second.

The model launched on the SpaceXAI API, in Cursor, in Grok Build and in the Grok app on the same day. A Fast variant runs about twice as quickly at about twice the price. Pick Grok 4.7 if you already use Cursor or the Grok ecosystem. Otherwise Sol or Opus 5.5 is the stronger buy.

Score breakdown

Intelligence & reasoning8.2
Coding & agents7.8
Price-performance8.5
Speed & context7.5
Availability & openness7.5

Key facts

Pricing
$2 / $6 per 1M tokens (Cached input $0.50/M. Above 200K prompt tokens: $4/$12.)
Free option
No
Platforms
API, Grok app, Cursor, GitHub Copilot, Grok Build
Released
21 September 2026
AA Intelligence Index
46 (xhigh and high)
Context window
500K tokens
Terminal-Bench 4.0
37.6% (SpaceXAI-reported, up from 20.3%)
Knowledge cutoff
May 2026

What we like

  • Scores 46 on the Intelligence Index for $2/$6
  • Big jump in agentic coding over Grok 4.6
  • About 188 tokens per second (Artificial Analysis)
  • No output length limit on the API

Watch out for

  • 500K context is half of most rivals
  • Prices double for prompts of 200K tokens or more
  • Most benchmark numbers come from SpaceXAI itself
#9 · Multimodal coding and document work with an open-weight fallback

Qwen3.8 Max

by Alibaba · Usage-based · $2 / $6 per 1M tokens
7.9/10

Qwen3.8 Max is Alibaba's largest model: 2.4 trillion parameters, of which about 95 billion are active for each token. That design, called mixture of experts, keeps it cheaper to run than its size suggests. The 0902 snapshot scores 45 on the Artificial Analysis Intelligence Index, level with GLM-5.3.

The API costs $2 input and $6 output per million tokens, and it accepts text, images and video. Alibaba reports 67.7 on SWE-bench Pro and 86.1 on OSWorld-Verified, and says the September update more than doubled its TerminalBench 3.0 score, from 11.3 to 29.0.

It is also the first Max-class Qwen you can download. But the open release (Qwen3.8-2.4T-A95B) is text-only, always runs in thinking mode, and supports 262K tokens natively rather than the API's 1M. It uses a custom license, not the Apache 2.0 license of smaller Qwen models.

The weak point is speed: about 39 tokens per second on Artificial Analysis, among the slowest here. Choose it if you want a strong multimodal API at a low price and like having downloadable weights as a backup.

Score breakdown

Intelligence & reasoning8.0
Coding & agents8.0
Price-performance8.8
Speed & context5.5
Availability & openness8.0

Key facts

Pricing
$2 / $6 per 1M tokens (Cached input $0.25/M. Text-only base weights (Qwen3.8-2.4T-A95B) downloadable under a custom license.)
Free option
No
Platforms
API, Alibaba Cloud Model Studio, Self-hosted
Released
3 August 2026; 0902 update on 2 September
AA Intelligence Index
45 (0902 snapshot)
Parameters
2.4T total, about 95B active
Context window
1M tokens (API); 262K native in open weights
SWE-bench Pro
67.7 (Alibaba-reported)

What we like

  • Scores 45 on the Intelligence Index at $2/$6
  • Takes text, image and video input
  • Base weights are downloadable
  • Big coding gains in the September snapshot

Watch out for

  • Slow, at about 39 tokens per second
  • Open weights lack vision and the full 1M context
  • Custom license on the large model
#10 · The cheapest capable model, with MIT-licensed weights

DeepSeek V4

by DeepSeek · Open source · $0.15 / $0.60 per 1M tokens (V4.1 Flash, off-peak)
7.8/10

DeepSeek V4 is the family to pick when price matters most. It comes in two main versions on the API. V4.1 Flash, released on 10 September, costs $0.15 input and $0.60 output per million tokens off-peak. V4 Pro 0813 costs $0.66/$1.98 off-peak. Prices double during peak hours in China's working day (01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays).

Both use the MIT license, the most permissive in this list. You can run them, change them and sell products built on them without a revenue cap.

On the Artificial Analysis Intelligence Index, the new V4.1 Flash scores 39 and runs at about 219 tokens per second. Oddly, the bigger V4 Pro 0813 scores lower at 36. DeepSeek reports 80.6% on SWE-bench Verified and 67.9 on Terminal-Bench 2.0 for V4 Pro, but independent tests so far put it well behind the US leaders and behind GLM-5.3 and Kimi K3.

Both support a 1M-token context and very long outputs of up to 384K tokens. For bulk jobs like tagging, summarizing or first-draft code, nothing here beats it on cost. For hard reasoning, spend more on a model higher up this list.

Score breakdown

Intelligence & reasoning6.8
Coding & agents7.2
Price-performance9.8
Speed & context7.5
Availability & openness9.5

Key facts

Pricing
$0.15 / $0.60 per 1M tokens (V4.1 Flash, off-peak) (V4 Pro $0.66/$1.98 off-peak. Peak-hour prices are double. Weights free under MIT.)
Free option
No
Platforms
API, Web, iOS, Android, Self-hosted
Latest releases
V4 Pro 0813 (13 Aug 2026), V4.1 Flash (10 Sept 2026)
AA Intelligence Index
39 (V4.1 Flash), 36 (V4 Pro 0813)
License
MIT
Context / max output
1M / 384K tokens
Output speed
About 219 tokens/s for V4.1 Flash (AA)

What we like

  • Lowest prices here: $0.15/$0.60 off-peak for V4.1 Flash
  • MIT license with no usage limits
  • 1M context and up to 384K output
  • V4.1 Flash is fast at about 219 tokens per second

Watch out for

  • Intelligence Index of 36 to 39 trails the leaders by 20 points
  • Peak-hour prices are double
  • Vendor coding claims are not matched by independent tests yet
#11 · Open-weight agents that browse, code and use a computer

Kimi K3

by Moonshot AI · Open source · $3 / $15 per 1M tokens
7.6/10

Kimi K3 is the largest open-weight model released so far, with 2.8 trillion parameters and 104 billion active per token. It scores 44 on the Artificial Analysis Intelligence Index at max effort, one point behind GLM-5.3 and Qwen3.8 Max. On Arena it scores 1485, a little above GLM-5.3.

Moonshot's own numbers point to agent work as its strength: 91.2 on BrowseComp (web research), 84.8 on OSWorld-Verified (using a computer) and 88.3 on Terminal-Bench 2.1. It also handles images well. Treat these as vendor claims until more independent results arrive.

The drawback is price. At $3 input and $15 output per million tokens, Kimi K3 costs more than GPT-6 Sol ($2/$10), which scores higher. Cache hits drop input to $0.30, which helps agents that resend the same context. Speed is slow, at about 37 tokens per second.

The weights are downloadable under the Kimi K3 License. It is fine for most uses, but companies that sell the model as a hosted service and earn over $20 million a year need a separate agreement. Use it through the Kimi app or self-host it when you need an open model with strong agent skills.

Score breakdown

Intelligence & reasoning7.8
Coding & agents8.0
Price-performance7.0
Speed & context5.5
Availability & openness8.5

Key facts

Pricing
$3 / $15 per 1M tokens (Cache hits $0.30/M. Weights free under the Kimi K3 License; large hosting providers need a separate deal.)
Free option
No
Platforms
Web, iOS, Android, API, Self-hosted
Released
July 2026 (API 16 July, weights 26 July)
AA Intelligence Index
44 (max)
Parameters
2.8T total, 104B active
Context window
1,048,576 tokens
Arena text score
1485 (13 Sept 2026)

What we like

  • Scores 44 on the Intelligence Index with open weights
  • Strong vendor-reported browsing and computer-use scores
  • 1M-token context
  • Cheap cache hits at $0.30/M

Watch out for

  • $3/$15 costs more than GPT-6 Sol, which scores higher
  • Slow, at about 37 tokens per second
  • License limits large hosting providers

How we scored these tools

Each tool is scored 0–10 on the criteria below, using public evidence: independent benchmarks, vendor documentation and pricing pages, aggregate user ratings and reputable reviews. The overall score is the weighted average. Nobody pays to be listed. Read our full methodology.

CriterionWeightWhat we look at
Intelligence & reasoning35%Artificial Analysis Intelligence Index, Humanity's Last Exam and similar hard reasoning tests, plus Arena human votes.
Coding & agents25%Results on Terminal-Bench, DeepSWE, OSWorld and other tests where the model writes code or runs multi-step tasks with tools.
Price-performance15%API price per million tokens compared with what the model scores. Cheaper for the same quality scores higher.
Speed & context10%Output speed in tokens per second (Artificial Analysis) and how much text the model can read at once.
Availability & openness15%Whether you can use it in a consumer app, on which plans, on which clouds, and whether the weights are downloadable.

The frontier at a glance

Intelligence Index scores are from Artificial Analysis on 23 September 2026 (highest effort level). Prices are standard API list prices per million input/output tokens.

Model Maker AA Index Price (in/out) Context Open weights
Claude Opus 5.5 Anthropic 58 $4 / $20 1M No
GPT-6 Astra OpenAI 53 $10 / $50 1.05M No
Claude Fable 5.1 Anthropic 53 $10 / $50 1M No
GPT-6 Sol OpenAI 48 $2 / $10 1.05M No
Muse Spark 1.3 Meta 48 $1.25 / $4.25 1M No
Grok 4.7 SpaceXAI 46 $2 / $6 500K No
GLM-5.3 Z.ai 45 $1.40 / $4.40 1M Yes (custom)
Qwen3.8 Max Alibaba 45 $2 / $6 1M Partly (custom)
Kimi K3 Moonshot 44 $3 / $15 1M Yes (custom)
Gemini 3.8 Flash Google 41 $0.75 / $3.75 1M No
DeepSeek V4.1 Flash DeepSeek 39 $0.15 / $0.60 1M Yes (MIT)

The gap between the top closed model and the best open one (MiMo-V2.6-Pro, 46) is now 12 points. Anthropic holds three of the top four spots on the index once you count Opus 5 (51).

What changed in September 2026

  • Prices fell sharply. Claude Opus 5.5 cut per-token prices 20% versus Opus 5. GPT-6 Sol and Luna cost about half of the GPT-5.6 models they replace. Grok 4.7 kept Grok 4.6's price while getting smarter.
  • Three tiers are now normal. OpenAI sells Astra ($10/$50), Sol ($2/$10) and Luna ($0.10/$0.50). Anthropic sells Fable, Opus and Sonnet. Pick the tier that matches the task, not the brand's top model by default.
  • Some models are held back. Claude Mythos 5.1 is the same model as Fable 5.1 with looser safeguards, and only vetted US security and life-science groups can use it. Gemini 3.5 Pro, announced for June, is still only with test partners.
  • Open weights caught up on price, not yet on quality. GLM-5.3, Qwen3.8 Max, Kimi K3 and DeepSeek V4 all ship downloadable weights, but the best of them scores 45 against Opus 5.5's 58.

How to choose

  1. Hardest problems, cost no object: Claude Opus 5.5 at max effort. Try GPT-6 Astra if you're tied to OpenAI.
  2. Coding agents on a budget: GPT-6 Sol or Claude Opus 5.5 at medium effort (Opus 5.5 still scores 51 there). See our best AI for coding list for full tools like Cursor and Claude Code.
  3. High-volume apps: Gemini 3.8 Flash for speed, Muse Spark 1.3 for a smarter mid-price option, DeepSeek V4.1 Flash for the lowest bill.
  4. Data must stay on your servers: GLM-5.3 or Kimi K3 if you have big GPUs; DeepSeek V4 if you need the MIT license. See best open-source LLMs and best local LLMs.
  5. Just want a chat app: You don't pick raw models there. Compare apps in best AI chatbots or our ChatGPT vs Claude vs Gemini face-off.

How to read the benchmarks

We lean on two independent sources. The Artificial Analysis Intelligence Index runs the same set of reasoning, knowledge, math and coding tests on every model and combines them into one score. The Arena leaderboard ranks models by millions of blind human votes on which answer people prefer. The two can disagree: Gemini 3.8 Flash scores only 41 on the index but ranks near the top on Arena, because people like its answers.

Vendor numbers (Terminal-Bench, DeepSWE, OSWorld, Humanity's Last Exam) are useful but each lab picks its own settings, tools and test versions. We label them as vendor-reported and give them less weight. Brand-new models such as Opus 5.5 and GPT-6 Sol have not yet collected enough Arena votes to appear there.

Finally, effort level matters. Most models now let you choose how long they think. GPT-6 Sol scores 48 at max effort but 43 at high. Always compare models at the effort level you will actually pay for.

What a million tokens actually costs you

A token is roughly three-quarters of a word. A typical coding agent task might read 200K tokens and write 20K. Using list prices and ignoring caching:

  • Claude Opus 5.5: about $0.80 input + $0.40 output = $1.20
  • GPT-6 Astra: about $2.00 + $1.00 = $3.00 (the 272K surcharge doesn't apply yet)
  • GPT-6 Sol: about $0.40 + $0.20 = $0.60
  • DeepSeek V4.1 Flash (off-peak): about $0.03 + $0.01 = $0.04

Real bills differ because models use different numbers of tokens for the same job, and prompt caching can cut input costs by 75% to 97%. Anthropic, for example, says Opus 5.5 uses fewer tokens than Opus 5 on typical work. Run your own test on 20 to 50 real tasks before you commit. For more on APIs, see best LLM APIs.

Expert tips
  1. Before paying for GPT-6 Astra or Claude Fable 5.1, run your task on Claude Opus 5.5 at high effort first. It scores 54 there on the Intelligence Index, above both at max, for well under half the price.
  2. Set the effort level on purpose. GPT-6 Sol drops from 48 to 43 between max and high effort, while Opus 5.5 stays at 51 even at medium. Test two levels and keep the cheapest one that passes.
  3. Turn on prompt caching for agents and chatbots that resend the same system prompt or codebase. Cache reads cost $0.20/M on Opus 5.5 and $0.25/M on Fable 5.1, versus $4 and $10 for normal input.
  4. If you use Gemini 3.8 Flash, budget for the price doubling to $1.50/$7.50 on 1 January 2027. If you use DeepSeek, schedule batch jobs outside peak hours (01:00–04:00 and 06:00–10:00 UTC on weekdays) to pay half.
  5. Keep private or customer data off Meta's Muse Spark Contributor tier. Launch coverage reports that the low $0.10/$0.20 price comes with Meta being allowed to use that traffic; check the terms first.

Jargon explained

Token
A chunk of text, about three-quarters of a word. AI models read and write in tokens, and APIs charge per million of them.
Context window
How much text a model can look at in one go, including your prompt, files and its own reply. 1M tokens is several long books.
Artificial Analysis Intelligence Index
A single score from an independent company that runs the same reasoning, math, knowledge and coding tests on every model, so you can compare them fairly.
Open weights
The model's trained numbers are published, so you can download it and run it on your own hardware. The license still sets rules on how you use it.
Effort level
A setting that controls how long a model thinks before answering. Higher effort gives better answers on hard problems but costs more and takes longer.
Mixture of experts
A design where only part of a huge model switches on for each token. It lets a 2.8-trillion-parameter model run at the cost of a much smaller one.

Frequently asked questions

What is the best AI model right now?

Claude Opus 5.5, released on 22 September 2026. It scores 58 on the Artificial Analysis Intelligence Index, five points ahead of GPT-6 Astra and Claude Fable 5.1, and costs $4/$20 per million tokens against their $10/$50.

What are Claude Fable and Claude Mythos?

They are Anthropic's largest models. Anthropic says Fable 5.1 and Mythos 5.1 are the same model with different safeguards. Fable 5.1 is generally available in the Claude apps and API at $10/$50. Mythos 5.1 has looser limits on cybersecurity and biology topics and is only open to vetted US organizations through Anthropic's Cyber Verification and Life Sciences Verification programs. Most people cannot access Mythos.

Is Gemini 3.5 Pro available?

Not publicly, as of 23 September 2026. Google announced it in May for a June launch, then delayed it several times. Google now says it is testing with partners and will release it when it is ready. The best Gemini model you can use today is Gemini 3.8 Flash.

What is the difference between GPT-6 Astra, Sol and Luna?

They are three price tiers. Astra ($10/$50) is the smartest and appears as GPT-6 Pro in ChatGPT's paid Pro, Business and Enterprise plans. Sol ($2/$10) is for coding and agents at a fifth of the price. Luna ($0.10/$0.50) handles simple, high-volume jobs like summaries and short answers.

What is the best open-source AI model?

Xiaomi's MiMo-V2.6-Pro scores highest among open-weight models on the Artificial Analysis Intelligence Index (46), with GLM-5.3 from Z.ai (45) and Kimi K3 (44) close behind. See our best open-source LLMs ranking. If you need the most permissive license, DeepSeek V4 uses MIT. Note that most of these use custom licenses, so check the terms. Our best open-source LLMs ranking goes deeper.

Which AI model is the cheapest?

Among capable models, DeepSeek V4.1 Flash at $0.15 input and $0.60 output per million tokens off-peak, and GPT-6 Luna at $0.10/$0.50. Meta's Muse Spark 1.3 Contributor tier ($0.10/$0.20) is cheaper still, but launch coverage says Meta can use that data. Gemini 3.8 Flash has a free API tier.

Which model has the biggest context window?

Most leaders now read about 1 million tokens (roughly 550,000 to 750,000 words). GPT-6 Astra and Sol list 1.05M. Grok 4.7 is the exception at 500K. For output, Claude and GPT-6 allow up to 128K tokens and DeepSeek up to 384K, while Gemini 3.8 Flash caps at about 65K.

Sources

Every fact on this page comes from public information. Vendor figures are labelled as vendor claims.

  1. Claude Opus 5.5 (Anthropic)
  2. Introducing Claude Fable 5.1 and Claude Mythos 5.1 (Anthropic)
  3. Models overview (Anthropic)
  4. Anthropic releases Opus 5.5 with lower prices and Fable-level performance (TechCrunch)
  5. LLM Leaderboard: Intelligence Index (Artificial Analysis)
  6. Text Arena leaderboard (Arena (LMArena))
  7. GPT-6 Astra model page (OpenAI)
  8. GPT-6 Sol model page (OpenAI)
  9. OpenAI cuts GPT-6 prices in half with Sol and Luna (The Next Web)
  10. How to use GPT-6 Astra when it rolls out to you (Engadget)
  11. Gemini 3.5: frontier intelligence with action (Google)
  12. Gemini API release notes (Google)
  13. Gemini 3.8 Flash model page (Google)
  14. Gemini Developer API pricing (Google)
  15. Gemini 3.5 Pro delays due to coding performance (9to5Google)
  16. Muse Spark on Meta Model API (Meta)
  17. Introducing Muse Spark 1.3 (Meta)
  18. xAI API release notes (xAI)
  19. SpaceXAI Releases Grok 4.7 (MarkTechPost)
  20. GLM-5.3 model card (Z.ai / Hugging Face)
  21. Z.ai pricing (Z.ai)
  22. Qwen3.8-2.4T-A95B model card (Alibaba Qwen / Hugging Face)
  23. Qwen3.8-Max: Features, Benchmarks, and Pricing (DataCamp)
  24. Kimi K3 model card (Moonshot AI / Hugging Face)
  25. Kimi API pricing (Moonshot AI)
  26. DeepSeek API models and pricing (DeepSeek)
  27. DeepSeek-V4-Pro-0813 model card (DeepSeek / Hugging Face)
  28. About ChatGPT Pro tiers (OpenAI Help Center)
  29. xAI launches Grok 4.7 at bargain prices, but benchmarks reveal a wide gap to Claude and GPT-6 (The Decoder)
  30. xAI is becoming SpaceXAI (The Verge)