Last updated: April 5, 2026 | Reviewed by Daniel Ashford

The LLM Judge Index

Independent, multi-dimensional AI model evaluation by Daniel Ashford. 646+ models ranked. Methodology

Full Leaderboard - 646 Models

Data by Artificial Analysis | Updated hourly
#ModelIntelGPQACodeInput $/MLicense
1
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic
53.493.7%81.6$10.0Prop.
2
Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Anthropic
53.293.4%80.7$10.0Prop.
3
GPT-6 Astra (max)OpenAI
52.896.1%76.9$10.0Prop.
4
GPT-6 Astra (xhigh)OpenAI
52.596.3%75.9$10.0Prop.
5
Claude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)Anthropic
51.290.6%79.1$10.0Prop.
6
GPT-6 Astra (high)OpenAI
51.094.9%77.1$10.0Prop.
7
Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic
50.793.2%78.0$5.0Prop.
8
GPT-6 Astra (medium)OpenAI
49.793.9%76.7$10.0Prop.
9
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic
49.792.6%76.5$10.0Prop.
10
Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)Anthropic
49.793.7%77.0$5.0Prop.
11
Claude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback)Anthropic
49.188.6%77.1$10.0Prop.
12
Muse Spark 1.3 (max)Meta
48.293.5%75.8$1.3Prop.
13
Claude Opus 5 (Adaptive Reasoning, High Effort)Anthropic
48.293.7%76.5$5.0Prop.
14
GPT-5.6 Sol (max)OpenAI
47.194.1%77.4$4.0Prop.
15
Claude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback)Anthropic
47.088.1%75.2$10.0Prop.
16
GPT-6 Astra (low)OpenAI
46.093.1%75.7$10.0Prop.
17
Muse Spark 1.3 (xhigh)Meta
45.294.1%76.5$1.3Prop.
18
Claude Opus 5 (Adaptive Reasoning, Medium Effort)Anthropic
45.191.9%74.3$5.0Prop.
19
GLM-5.3 (max)Z AI
44.991.7%74.8$1.4Prop.
20
Grok 4.6 (high)SpaceXAI
44.494.9%76.8$2.0Prop.
21
Grok 4.6 (xhigh)SpaceXAI
44.393.5%75.9$2.0Prop.
22
GPT-5.6 Sol (xhigh)OpenAI
44.193.1%78.3$4.0Prop.
23
Kimi K3 (max)Kimi
43.893.5%76.2$3.0Prop.
24
Grok 4.6 (medium)SpaceXAI
43.093.5%74.4$2.0Prop.
25
GPT-5.6 Sol (high)OpenAI
42.592.8%77.2$4.0Prop.
26
GPT-5.6 Terra (max)OpenAI
42.392.5%76.7$2.0Prop.
27
Claude Opus 4.8 (Adaptive Reasoning, Max Effort)Anthropic
42.092%74.3$5.0Prop.
28
GLM-5.3-FlashZ AI
41.991.2%71.5$0.15Prop.
29
Gemini 3.8 Flash (high)Google
41.295.3%76.3$0.75Prop.
30
Claude Opus 4.7 (Adaptive Reasoning, Max Effort)Anthropic
40.791.4%73.6$5.0Prop.
31
Qwen3.8 MaxAlibaba
40.392.7%71.8$2.0Prop.
32
Gemini 3.8 Flash (medium)Google
40.093.5%74.1$0.75Prop.
33
Qwen3.8 2.4T A95BAlibaba
40.093.5%71.9$2.0Prop.
34
Qwen3.8-Flash-NextAlibaba
39.992.3%73.1$0.15Prop.
35
Muse Spark 1.2 (xhigh)Meta
39.890.4%72.2$1.3Prop.
36
Claude Opus 5 (Adaptive Reasoning, Low Effort)Anthropic
39.888.9%66.9$5.0Prop.
37
Gemini 3.7 Flash (medium)Google
39.692.1%71.5$0.75Prop.
38
DeepSeek V4.1 Flash (Reasoning, Max Effort)DeepSeek
39.5--$0.30Prop.
39
GPT-5.6 Sol (medium)OpenAI
39.592.6%76.3$4.0Prop.
40
Gemini 3.7 Flash (high)Google
39.494.5%76.1$0.75Prop.
41
Grok 4.5 (high)SpaceXAI
39.193.1%72.4$2.0Prop.
42
GPT-5.4 (xhigh)OpenAI
39.092%71.1$2.5Prop.
43
GPT-5.5 (xhigh)OpenAI
38.693.5%74.9$5.0Prop.
44
Claude Sonnet 5 (Adaptive Reasoning, Max Effort)Anthropic
38.491.1%71.5$2.0Prop.
45
GPT-5.6 Terra (xhigh)OpenAI
38.290.8%70.6$2.0Prop.
46
GPT-5.6 Luna (max)OpenAI
37.591.1%71.4$0.20Prop.
47
GPT-5.5 (high)OpenAI
37.393.2%71.6$5.0Prop.
48
Gemini 3.7 Flash (low)Google
36.990.1%71.0$0.75Prop.
49
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)DeepSeek
36.392.8%68.8$1.3Prop.
50
Agnes 3.0 FlashSapiens AI
35.592.4%-$0.05Prop.
Showing top 50 of 646 models. Full data powered by Artificial Analysis.

Best LLM By Use Case

💻
Best for Code Generation
Weighted ranking
💬
Best for Customer Chatbot
Weighted ranking
✍️
Best for Content Writing
Weighted ranking
📊
Best for Data Analysis
Weighted ranking
🔬
Best for Research & RAG
Weighted ranking
🛡️
Best for Safety-Critical
Weighted ranking

Best LLM By Industry

🎓
Education
Schools, tutoring and edtech
🏥
Healthcare
Hospitals, clinics and health tech
🏦
Financial Services
Banking, investment and fintech
⚖️
Legal
Law firms, contracts and legal tech
💬
Customer Support
Help desks, chatbots and CX

Popular Comparisons

Claude Opus 4 vs GPT-5.3 Codex
Full comparison
Claude Opus 4 vs Gemini 2.5 Ultra
Full comparison
Claude Opus 4 vs Claude Sonnet 4
Full comparison
Claude Opus 4 vs GPT-4o
Full comparison
Claude Opus 4 vs Llama 4 405B
Full comparison
Claude Opus 4 vs Mistral Large 3
Full comparison
Claude Opus 4 vs Qwen 3.5 Plus
Full comparison
Claude Opus 4 vs DeepSeek V3
Full comparison

LLM Glossary

47 AI and language model terms explained. Browse all

Large Language Model (LLM)
An AI system trained on massive text data to understand and generate human language.
Tokens
The basic units of text that LLMs process — roughly 3/4 of a word.
Context Window
The maximum amount of text an LLM can process in a single request.
Hallucination
When an LLM generates plausible-sounding but factually incorrect information.
Inference
The process of an LLM generating a response to your input.
Prompt
The text input you send to an LLM to get a response.
Fine-Tuning
Customizing a pre-trained LLM on your specific data to improve performance for your use case.
RAG (Retrieval-Augmented Generation)
A technique that gives LLMs access to external documents to improve accuracy and reduce hallucination.

By Provider

Anthropic (3)OpenAI (3)Google (2)Meta (1)Mistral (1)Alibaba (1)DeepSeek (1)