Qwen3.8 Max logo

Proprietary model

Released August 2026

Qwen3.8 Max Intelligence, Performance & Price Analysis

Model summary

Intelligence

56
Artificial Analysis Intelligence Index
4 out of 4 units for Intelligence.

Speed

61.5
Output tokens per second
2 out of 4 units for Speed.

Price

Input
$2.00
per 1M tokens
Output
$6.00
per 1M tokens
2 out of 4 units for Price.

Cache Hit Price

$0.25
USD per 1M tokens
2 out of 4 units for Cache Hit Price.

Verbosity

150M
Output tokens from Intelligence Index
4 out of 4 units for Verbosity.

Qwen3.8 Max is amongst the leading models in intelligence and reasonably priced when comparing to other models of similar price. It's also slower than average and very verbose. The model supports text, image, and video input, outputs text, and has a 1M tokens context window.

Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, placing it well above average among comparable models (median: 32). When evaluating the Intelligence Index, it generated 150M tokens, which is very verbose in comparison to the median of 66M.

Pricing for Qwen3.8 Max is $2.00 per 1M input tokens (somewhat expensive, median: $1.75) and $6.00 per 1M output tokens (moderately priced, median: $10.00). In total, it cost $1749.16 to evaluate Qwen3.8 Max on the Intelligence Index.

At 62 tokens per second, Qwen3.8 Max is slower than average (71).

ReasoningYes

This page shows the reasoning version of this model.

A non-reasoning variant may also exist.

Input modality

Supports: text, image, and video

Output modality

Supports: text

Context window1M
~1500 A4 pages of size 12 Arial font

Metrics are compared against models of the same class:

  • Non-reasoning models → compared only with other non-reasoning models
  • Reasoning models → compared across both reasoning and non-reasoning
  • Open weights models → compared only with other open weights models of the same size class:
    • Tiny: ≤4B parameters
    • Small: 4B–40B parameters
    • Medium: 40B–150B parameters
    • Large: >150B parameters
  • Proprietary models → compared across proprietary and open weights models of the same price range, using a blended 3:1 input/output price ratio:
    • <$0.15 per 1M tokens
    • $0.15–$1 per 1M tokens
    • >$1 per 1M tokens

Highlights

Artificial Analysis Intelligence Index · Higher is better
Claude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelMuse Spark 1.2 (xhigh)Logo of Muse Spark 1.2 (xhigh)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelDeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model61605957565454515050443824
Output tokens per second · Higher is better
Gemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelDeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelClaude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning model213176145143111104686462595439
Weighted average cost (USD) per Intelligence Index task · Lower is better
DeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelMuse Spark 1.2 (xhigh)Logo of Muse Spark 1.2 (xhigh)Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning model$0.03$0.08$0.14$0.36$0.38$0.40$0.56$0.57$0.86$1.14$1.23$2.34$3.15

Intelligence

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR
Claude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelGPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelMuse Spark 1.2 (xhigh)Logo of Muse Spark 1.2 (xhigh)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelDeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model6160605959585756565656555554545454535352515050443824
Reasoning models are indicated by a lightbulb icon

Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Artificial Analysis Intelligence Index by Open Weights / Proprietary

Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR
Claude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelGPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelMuse Spark 1.2 (xhigh)Logo of Muse Spark 1.2 (xhigh)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelDeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model6160605959585756565656555554545454535352515050443824
Reasoning models are indicated by a lightbulb icon

Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Indicates whether the model weights are available. Models are labelled as 'Commercial Use Restricted' if the weights are available but commercial use is limited (typically requires obtaining a paid license).

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better

Agentic real-world work tasks, (Elo-500)/2000

Claude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelGPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelMuse Spark 1.2 (xhigh)Logo of Muse Spark 1.2 (xhigh)Reasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelDeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelGPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model68%66%62%62%62%61%59%59%57%56%56%55%54%54%54%53%53%51%50%50%49%48%46%44%33%15%

Agentic tool use

Qwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelGPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelDeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelClaude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelMuse Spark 1.2 (xhigh)Logo of Muse Spark 1.2 (xhigh)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model42%33%33%33%33%33%32%32%31%31%31%30%29%29%29%28%28%27%27%27%26%25%24%14%13%12%

Agentic coding & terminal use

GPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelClaude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelGPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelMuse Spark 1.2 (xhigh)Logo of Muse Spark 1.2 (xhigh)Reasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelDeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model90%89%88%88%88%88%87%86%86%85%85%85%84%83%82%81%81%80%80%79%79%78%78%65%54%26%

Coding

Claude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelGPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelMuse Spark 1.2 (xhigh)Logo of Muse Spark 1.2 (xhigh)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelClaude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelDeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model60%59%57%56%56%56%56%56%56%56%55%55%54%54%54%54%53%53%53%52%51%50%50%45%40%39%

Reasoning & knowledge

Claude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelClaude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelMuse Spark 1.2 (xhigh)Logo of Muse Spark 1.2 (xhigh)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelGPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelDeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model53%53%53%51%49%47%46%45%44%44%44%44%43%42%41%40%40%40%40%40%40%38%37%37%27%18%

Scientific reasoning

GPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelClaude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelDeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelMuse Spark 1.2 (xhigh)Logo of Muse Spark 1.2 (xhigh)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model94%94%94%94%94%93%93%93%93%93%93%93%93%93%93%93%92%92%91%91%91%91%90%89%87%78%

Physics reasoning

GPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelClaude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelGPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelGPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelMuse Spark 1.2 (xhigh)Logo of Muse Spark 1.2 (xhigh)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelDeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model32%30%29%29%29%28%28%27%27%27%26%25%23%23%21%21%20%18%17%17%15%12%11%4%3%1%
Claude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelClaude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelMuse Spark 1.2 (xhigh)Logo of Muse Spark 1.2 (xhigh)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelDeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning model61%59%58%57%57%57%56%54%53%53%52%51%50%47%46%46%46%45%38%38%37%31%25%22%22%15%
MiniMax-M3Logo of MiniMax-M3Reasoning modelMuse Spark 1.2 (xhigh)Logo of Muse Spark 1.2 (xhigh)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelClaude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelDeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelGPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model84%72%72%71%64%64%63%60%50%50%49%48%48%46%46%45%16%15%14%14%13%13%12%11%11%9%

Long context reasoning

Kimi K3 (max)Logo of Kimi K3 (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelGPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelClaude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelGPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelDeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelMuse Spark 1.2 (xhigh)Logo of Muse Spark 1.2 (xhigh)Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model75%74%74%74%74%73%71%71%71%71%70%70%70%70%69%69%69%68%68%68%67%67%67%66%65%51%

Agentic knowledge work, Elo

Claude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model1719169116061574154115021470138513421314127712521150110710999628748

Agentic SaaS workflows

Kimi K3 (max)Logo of Kimi K3 (max)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning model53%51%51%51%49%49%46%42%40%39%28%16%6%

Legal agentic work, criterion pass rate

Kimi K3 (max)Logo of Kimi K3 (max)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model95%94%92%91%91%90%88%87%86%85%82%14%

Agentic business operations

Claude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelClaude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model51%47%47%45%45%44%43%43%41%32%29%26%

Instruction following

MiniMax-M3Logo of MiniMax-M3Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelGPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning model83%81%76%73%73%72%71%71%70%69%69%66%63%62%59%

Long-horizon agentic tasks

Kimi K3 (max)Logo of Kimi K3 (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model41%39%38%34%3%

Kubernetes incident root-cause analysis

GPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model56%51%48%47%46%43%6%

Visual reasoning

Claude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelGPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelGPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning model85%84%83%83%83%82%82%82%82%81%81%81%81%80%80%79%79%79%77%
Reasoning models are indicated by a lightbulb icon

While model intelligence generally translates across use cases, specific evaluations may be more relevant for certain use cases.

Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

AA-Omniscience

AA-Omniscience Index

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.
Claude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelClaude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelMuse Spark 1.2 (xhigh)Logo of Muse Spark 1.2 (xhigh)Reasoning modelGPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelGPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelDeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model4031302827262626242222212020191818154410−1−3−16−50
Reasoning models are indicated by a lightbulb icon

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.

Intelligence Index Comparisons

Intelligence Index vs. Cost per Intelligence Index Task

Artificial Analysis Intelligence Index · Weighted average cost (USD) per Artificial Analysis Intelligence Index task
Most attractive quadrant
Pareto line
$0.03$0.04$0.06$0.08$0.1$0.2$0.3$0.4$0.6$0.8$1$2$3$4$5$6$7$8$10$20Cost per Task (USD, Log Scale)20253035404550556065Artificial Analysis Intelligence Index
Reasoning models are indicated by a lightbulb icon

Weighted average cost per Intelligence Index task. Each evaluation’s cost is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight.

Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Token Use

Output Tokens per Intelligence Index Task

Weighted average number of output tokens used to run one task in the Artificial Analysis Intelligence Index
GPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelGPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelMuse Spark 1.2 (xhigh)Logo of Muse Spark 1.2 (xhigh)Reasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelClaude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelDeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning model7k6k7k7k8k11k13k15k12k15k12k14k19k18k19k22k27k29k27k26k35k39k37k58k5k8k11k11k12k12k15k17k17k21k21k24k25k26k26k30k31k34k36k37k38k40k45k45k46k72k4k5k5k6k7k6k4k6k9k9k13k12k7k13k12k11k9k8k11k15k10k7k9k15k
Reasoning models are indicated by a lightbulb icon

The number of tokens required per Intelligence Index task. This is calculated by multiplying the output tokens per eval by the relative weights of each benchmark in the Intelligence Index, then dividing by task count (excluding repeats).

Cost

Cost per Intelligence Index Task

Weighted average cost (USD) per Artificial Analysis Intelligence Index task, segmented by token type. Lower is better
DeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelGPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelMuse Spark 1.2 (xhigh)Logo of Muse Spark 1.2 (xhigh)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelGPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelClaude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning model$0.31$0.21$0.19$0.37$0.41$0.21$0.31$1.26$0.03$0.08$0.14$0.31$0.36$0.38$0.39$0.40$0.51$0.55$0.56$0.57$0.72$0.80$0.83$0.86$1.14$1.17$1.23$1.23$1.72$1.80$2.03$2.22$2.34$3.15
Reasoning models are indicated by a lightbulb icon

Weighted average cost per Intelligence Index task. Each evaluation’s cost is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight.

Cost to Run Artificial Analysis Intelligence Index

Cost (USD) to run all evaluations in the Artificial Analysis Intelligence Index
DeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelGPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelMuse Spark 1.2 (xhigh)Logo of Muse Spark 1.2 (xhigh)Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelClaude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning model$370$470$613$569$434$850$680$743$514$412$579$367$540$373$572$500$545$894$788$72$96$204$530$582$592$593$638$727$939$956$1,115$1,403$1,543$1,719$1,749$1,974$2,437$2,778$2,824$2,910$3,738$3,753$3,836$4,010$5,631$353$424$508
Reasoning models are indicated by a lightbulb icon

The cost to run the evaluations in the Artificial Analysis Intelligence Index, calculated using the model's input, cache hit, cache write, reasoning, and answer token prices and the number of tokens used across evaluations (excluding repeats).

Pricing: Cache Hit, Input, and Output

Price (USD per M Tokens)
DeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelMuse Spark 1.2 (xhigh)Logo of Muse Spark 1.2 (xhigh)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelClaude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning model<0.010.150.060.20.150.230.250.30.150.20.20.20.30.50.50.50.50.50.50.50.50.50.50.50.510.140.150.30.61.251.35221.52223555555555555100.280.61.22.754.254.29667.51012121525252525252530303030303050
Reasoning models are indicated by a lightbulb icon

Price per token for cached prompts (previously processed), typically offering a significant discount compared to regular input price, represented as USD per million tokens. The values shown here are the cache hit price; cache write and cache storage are billed separately and vary by provider — see "Cache pricing by provider" for detail.

Context Window

Context Window

Context window: tokens limit · Higher is better
Kimi K3 (max)Logo of Kimi K3 (max)Reasoning modelMuse Spark 1.2 (xhigh)Logo of Muse Spark 1.2 (xhigh)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelClaude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelDeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelGPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model1M1M1M1M1M1M1M1M1M1M1M1M1M1M1M1M1M1M1M1M1M922k922k500k262k131k
Reasoning models are indicated by a lightbulb icon

Larger context windows are relevant to RAG (Retrieval Augmented Generation) LLM workflows which typically involve reasoning and information retrieval of large amounts of data.

Maximum number of combined input & output tokens. Output tokens commonly have a significantly lower limit (varied by model).

Speed

Measured by Output Speed (tokens per second)

Output Speed

Output tokens per second · Higher is better
Gemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelDeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelGPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelGPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelClaude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning model2131761451431261161111048275726864646362615959555554545139
Reasoning models are indicated by a lightbulb icon

Tokens per second received while the model is generating tokens (ie. after first chunk has been received from the API for models which support streaming).

Figures represent performance of the model's first-party API (e.g. OpenAI for o1) or the median across providers where a first-party API is not available (e.g. Meta's Llama models).

Time per Intelligence Index Task

Weighted average decode time (minutes) per task; excludes TTFT and overhead time · Lower is better
GPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelDeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelClaude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning model1.31.71.92.12.32.42.62.72.93.23.33.84.04.04.44.54.85.45.76.67.77.88.79.110.1
Reasoning models are indicated by a lightbulb icon

The weighted average time (seconds) per Artificial Analysis Intelligence Index task. This is calculated by dividing output tokens per task by output speed, weighted by the relative weights of each benchmark in the Intelligence Index.

Latency

Measured by Time (seconds) to First Token

Latency: Time To First Answer Token

Seconds to first answer token received · Accounts for reasoning model 'thinking' time
GPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelDeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelGPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelClaude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning model10.913.016.822.323.027.534.036.237.548.156.879.3129.6140.7179.3195.17.27.510.912.213.015.516.817.919.520.522.323.027.534.035.136.237.548.154.856.879.3129.6140.7179.3195.111.414.015.718.119.232.551.7
Reasoning models are indicated by a lightbulb icon

Time to first answer token received, in seconds, after API request sent. For reasoning models, this includes the 'thinking' time of the model before providing an answer. For models which do not support streaming, this represents time to receive the completion.

End-to-End Response Time

Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed

End-to-End Response Time

Seconds to output 500 tokens, including reasoning model 'thinking' time · Lower is better
gpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelGPT-5.6 Sol (medium)Logo of GPT-5.6 Sol (medium)Reasoning modelClaude Opus 5 (medium)Logo of Claude Opus 5 (medium)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGemini 3.6 FlashLogo of Gemini 3.6 FlashReasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGPT-5.6 Sol (high)Logo of GPT-5.6 Sol (high)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelDeepSeek V4 Flash 0731 (max)Logo of DeepSeek V4 Flash 0731 (max)Reasoning modelClaude Opus 5 (high)Logo of Claude Opus 5 (high)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGPT-5.6 Terra (xhigh)Logo of GPT-5.6 Terra (xhigh)Reasoning modelGPT-5.5 (high)Logo of GPT-5.5 (high)Reasoning modelQwen3.8 MaxLogo of Qwen3.8 MaxReasoning modelClaude Opus 5 (xhigh)Logo of Claude Opus 5 (xhigh)Reasoning modelGPT-5.6 Sol (xhigh)Logo of GPT-5.6 Sol (xhigh)Reasoning modelClaude Opus 5 (max)Logo of Claude Opus 5 (max)Reasoning modelKimi K3 (max)Logo of Kimi K3 (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning model16.810.913.022.323.027.534.036.237.548.156.879.3129.6140.7179.3195.111.414.015.718.119.232.551.715.115.516.618.919.119.420.821.424.125.331.532.736.138.343.143.246.656.066.067.785.9136.9148.5183.2201.212.9
Reasoning models are indicated by a lightbulb icon

Seconds to receive a 500 token response. Key components:

  • Input time: Time to receive the first response token
  • Thinking time (only for reasoning models): Time reasoning models spend outputting tokens to reason prior to providing an answer. Amount of tokens based on the average reasoning tokens across a diverse set of 60 prompts (methodology details).
  • Answer time: Time to generate 500 output tokens, based on output speed

Figures represent performance of the model's first-party API (e.g. OpenAI for o1) or the median across providers where a first-party API is not available (e.g. Meta's Llama models).

Frequently Asked Questions

Common questions about Qwen3.8 Max

Qwen3.8 Max was released on August 3, 2026.

Qwen3.8 Max was created by Alibaba.

Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, placing it well above average among other reasoning models in a similar price tier (median: 32).

Qwen3.8 Max generates output at 61.5 tokens per second (based on Alibaba's API), which is below average compared to other reasoning models in a similar price tier (median: 71.1 t/s).

Qwen3.8 Max has a time to first token (TTFT) of 2.56s (based on Alibaba's API), which is better than average compared to other reasoning models in a similar price tier (median: 2.84s).

Qwen3.8 Max costs $2.00 per 1M input tokens (somewhat higher than average, median: $1.75) and $6.00 per 1M output tokens (better than average, median: $10.00), based on Alibaba's API.

Qwen3.8 Max costs $2.00 per 1M input tokens and $6.00 per 1M output tokens (based on Alibaba's API). For a blended rate (7:2:1 cache hit/input/output ratio), this is $1.18 per 1M tokens. Pricing may vary by provider. Compare provider pricing

When evaluated on the Intelligence Index, Qwen3.8 Max generated 150M output tokens, which is at the higher end compared to other reasoning models in a similar price tier (median: 66M).

Yes, Qwen3.8 Max is a reasoning model. It uses extended thinking or chain-of-thought reasoning to work through complex problems before providing an answer.

Qwen3.8 Max supports text, image, and video input.

Qwen3.8 Max supports text output.

Yes, Qwen3.8 Max supports image input and can analyze, describe, and answer questions about images.

Yes, Qwen3.8 Max is multimodal. It can process text, image, and video input and generate text output.

Qwen3.8 Max has a context window of 1.0M tokens. This determines how much text and conversation history the model can process in a single request.

No, Qwen3.8 Max is proprietary. The model weights are not publicly available.

Qwen3.8 Max is a proprietary model and Alibaba has not disclosed the model size or parameter count.

Qwen3.8 Max achieves a score of 56 on the Artificial Analysis Intelligence Index. This composite benchmark evaluates models across reasoning, knowledge, mathematics, and coding.

Yes, Qwen3.8 Max is available via API through 1 provider. Compare API providers

Qwen3.8 Max is available through 1 API provider. Compare providers