LLM Leaderboard

Open-weight models (Llama, Qwen, DeepSeek, Mistral…) ranked side by side with the proprietary frontier (GPT, Claude, Gemini, Grok) — live pricing, context, benchmarks and capabilities. No marketing — just the facts.

Live data Loading…
🕐 Time machine
#ModelCreatorIntelligenceCoding
1Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Anthropic53.481.6
2GPT-6 Astra (max)OpenAI52.876.9
3Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic50.778.0
4Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic49.776.5
5Muse Spark 1.3 (max)Meta48.275.8
6GPT-5.6 Sol (max)OpenAI47.177.4
7GLM-5.3 (max)Z AI44.974.8
8Grok 4.6 (high)SpaceXAI44.476.8
9Kimi K3 (max)Kimi43.876.2
10GPT-5.6 Terra (max)OpenAI42.376.7
11Claude Opus 4.8 (Adaptive Reasoning, Max Effort)Anthropic42.074.3
12GLM-5.3-FlashZ AI41.971.5
13Gemini 3.8 Flash (high)Google41.276.3
14Claude Opus 4.7 (Adaptive Reasoning, Max Effort)Anthropic40.773.6
15Qwen3.8 MaxAlibaba40.371.8
16Qwen3.8 2.4T A95BAlibaba40.071.9
17Qwen3.8-Flash-NextAlibaba39.973.1
18Muse Spark 1.2 (xhigh)Meta39.872.2
19Gemini 3.7 Flash (medium)Google39.671.5
20DeepSeek V4.1 Flash (Reasoning, Max Effort)DeepSeek39.5
21Grok 4.5 (high)SpaceXAI39.172.4
22GPT-5.4 (xhigh)OpenAI39.071.1
23GPT-5.5 (xhigh)OpenAI38.674.9
24Claude Sonnet 5 (Adaptive Reasoning, Max Effort)Anthropic38.471.5
25GPT-5.6 Luna (max)OpenAI37.571.4
26DeepSeek V4 Pro 0813 (Reasoning, Max Effort)DeepSeek36.368.8
27Agnes 3.0 FlashSapiens AI35.5
28Agnes 2.5 Pro BetaSapiens AI35.262.3
29DeepSeek V4 Flash Vision (Reasoning, Max Effort)DeepSeek35.065.0
30DeepSeek V4 Flash 0731 (Reasoning, Max Effort)DeepSeek34.569.1
31Gemini 3.6 Flash (high)Google34.369.2
32Muse Spark 1.1 (xhigh)Meta34.371.3
33GLM-5.2 (max)Z AI34.068.8
34Qwen3.8 27B (xhigh)Alibaba33.968.1
35Gemini 3.5 Flash (medium)Google33.6
36Motif 3Motif Technologies33.663.5
37GPT-5.3 Codex (xhigh)OpenAI32.5
38Claude Opus 4.6 (Adaptive Reasoning, Max Effort)Anthropic31.9
39Kimi K2.6Kimi31.361.8
40Muse SparkMeta31.358.6
41DeepSeek V4 Pro (Reasoning, Max Effort)DeepSeek30.959.4
42K2 Horizon 375B A23BMBZUAI Institute of Foundation Models30.861.5
43Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic30.563.0
44Apodex 1.1Apodex30.460.8
45GPT-5.2 (xhigh)OpenAI30.4
46Gemini 3.1 Pro PreviewGoogle30.468.8
47Qwen3.7 MaxAlibaba29.966.0
48MiniMax-M3MiniMax29.658.6
49Claude Opus 4.5 (Reasoning)Anthropic29.1
50MiMo-V2-ProXiaomi28.6
51GPT-5.2 Codex (xhigh)OpenAI28.5
52Qwen3.6 Max PreviewAlibaba28.4
53Nex-N2-ProNex AGI28.259.1
54Solar Pro 4Upstage28.252.7
55Gemini 3 Pro Preview (high)Google28.0
56GLM-5 (Reasoning)Z AI27.9
57Agnes 2.5 Pro AlphaSapiens AI27.858.8
58JT-4.1 Flash 236B A21BChina Mobile27.352.4
59Grok Build 0.1 0616SpaceXAI27.251.5
60Quasar 438B (max, based on GLM-5.2)Multiverse Computing27.161.2
Loading the leaderboard…

How this is ranked: models are ordered by real usage popularity (weekly tokens processed across apps), via the OpenRouter API. Only open-weight models are included (Llama, DeepSeek, Qwen, Mistral, Gemma, Kimi, GLM, Phi, Nemotron, and more). Pricing is per million tokens; context is the maximum window. Movement arrows compare today's rank to the previous snapshot. Data refreshes daily.