PromptLoop
News Analyse Werkstatt Generative Medien Originals Glossar KI-Modelle Vergleich Kosten-Rechner
📊 Live Benchmark

KI-Modelle Leaderboard 2026

Alle relevanten Large Language Models auf einen Blick — sortiert nach Quality Index, Geschwindigkeit, Latenz und Preis. Datenquelle: Artificial Analysis.

200 aktive Modelle aus 12 Vendor-Familien · Letzte Synchronisation:

Wie liest du dieses Leaderboard?

Quality Index ist ein zusammengesetzter Wert von Artificial Analysis aus über zehn unabhängigen Benchmarks (MMLU-Pro, GPQA, HLE, LiveCodeBench, SciCode, AIME u. a.). Höher = besser. Speed misst Output-Tokens pro Sekunde im Median über alle Hosting-Provider. Latency ist die Time-to-First-Token in Sekunden — wichtig für Streaming-UIs. Preise sind US-Dollar pro 1 Million Tokens, separat für Input und Output. Sortiere nach deiner Priorität, filter nach Anbieter oder Preis-Bucket — und prüfe konkrete Kosten direkt im Token-Rechner oder zwei Modelle im Head-to-Head-Vergleich.

Modalität
# Modell Vendor Quality Speed Latency Preis (USD/1M)
1
OpenAI: GPT-6 Astra
openai/gpt-6-astra
OpenAI 54,3 57,5 t/s 61,93 s
$10.00 in
$50.00 out
2
Anthropic: Claude Fable 5.1
anthropic/claude-fable-5.1
Anthropic 53,8 51,7 t/s 34,81 s
$10.00 in
$50.00 out
3
Anthropic: Claude Fable 5
anthropic/claude-fable-5
Anthropic 53,2 59,4 t/s 53,36 s
$10.00 in
$50.00 out
4
Meta: Muse Spark 1.3
meta/muse-spark-1.3
Meta 51,6 107,9 t/s 45,19 s
$1.25 in
$4.25 out
5
Claude Opus 5
anthropic/claude-opus-5
Anthropic 49,5 49,1 t/s 9,07 s
$5.00 in
$25.00 out
6
SpaceXAI: Grok 4.6
x-ai/grok-4.6
xAI 48,4 55,4 t/s 35,79 s
$2.00 in
$6.00 out
7
OpenAI: GPT-5.6 Sol
openai/gpt-5.6-sol
OpenAI 48,3 73,5 t/s 6,91 s
$4.00 in
$20.00 out
8
Google: Gemini 3.8 Flash
google/gemini-3.8-flash
Google 46,8 392 t/s 5,04 s
$0.75 in
$3.75 out
9
Meta: Muse Spark 1.2
meta/muse-spark-1.2
Meta 46,8 223,6 t/s 7,37 s
$1.25 in
$4.25 out
10
Anthropic: Claude Opus 4.8
anthropic/claude-opus-4.8
Anthropic 46,4 0 t/s 0 ms
$5.00 in
$25.00 out
11
SpaceXAI: Grok 4.5
x-ai/grok-4.5
xAI 45,5 48,4 t/s 11,13 s
$2.00 in
$6.00 out
12
OpenAI: GPT-5.5
openai/gpt-5.5
OpenAI 44,1 0 t/s 0 ms
$5.00 in
$30.00 out
13
Google: Gemini 3.7 Flash
google/gemini-3.7-flash
Google 43,4 294,3 t/s 4,37 s
$0.75 in
$3.75 out
14
Meta: Muse Spark 1.1
meta/muse-spark-1.1
Meta 43,3 0 t/s 0 ms
$1.25 in
$4.25 out
15
DeepSeek: DeepSeek V4 Pro 0423
deepseek/deepseek-v4-pro
DeepSeek 42,1 59,5 t/s 1,11 s
$1.32 in
$3.96 out
16
OpenAI: GPT-5.6 Terra
openai/gpt-5.6-terra
OpenAI 41,3 87,0 t/s 2,71 s
$2.00 in
$12.00 out
17
Google: Gemini 3.6 Flash
google/gemini-3.6-flash
Google 40,3 185,7 t/s 11,51 s
$0.75 in
$3.75 out
18
Google: Gemini 3.5 Flash
google/gemini-3.5-flash
Google 39,7 0 t/s 0 ms
$1.50 in
$9.00 out
19
OpenAI: GPT-5.3-Codex
openai/gpt-5.3-codex
OpenAI 36,9 120,1 t/s 31,50 s
$1.75 in
$14.00 out
20
Google: Gemini 3.1 Pro Preview
google/gemini-3.1-pro-preview
Google 36,7 117,4 t/s 20,28 s
$2.00 in
$12.00 out
21
MiniMax: MiniMax M3
minimax/minimax-m3
MiniMax 35,7 84,6 t/s 1,12 s
$0.30 in
$1.20 out
22
Anthropic: Claude Opus 4.7
anthropic/claude-opus-4.7
Anthropic 35,4 0 t/s 0 ms
$5.00 in
$25.00 out
23
Anthropic: Claude Sonnet 5
anthropic/claude-sonnet-5
Anthropic 33,5 60,0 t/s 1,25 s
$2.00 in
$10.00 out
24
Xiaomi: MiMo-V2-Pro
xiaomi/mimo-v2-pro
Xiaomi 33,1 0 t/s 0 ms
$0 in
$0 out
25
DeepSeek: DeepSeek V4 Flash 0423
deepseek/deepseek-v4-flash
DeepSeek 33,1 0 t/s 0 ms
$0.13 in
$0.28 out
26
OpenAI: GPT-5.2-Codex
openai/gpt-5.2-codex
OpenAI 33 0 t/s 0 ms
$1.75 in
$14.00 out
27
Upstage: Solar Pro 4
upstage/solar-pro4
Upstage 32,7 59,1 t/s 1,22 s
$0.30 in
$1.20 out
28
Xiaomi: MiMo-V2.5-Pro
xiaomi/mimo-v2.5-pro
Xiaomi 32,6 34,3 t/s 4,02 s
$0.435 in
$0.87 out
29
OpenAI: GPT-5.2
openai/gpt-5.2
OpenAI 30,9 0 t/s 0 ms
$1.75 in
$14.00 out
30
Anthropic: Claude Opus 4.6
anthropic/claude-opus-4.6
Anthropic 30,8 0 t/s 0 ms
$5.00 in
$25.00 out
31
OpenAI: GPT-5.6 Luna
openai/gpt-5.6-luna
OpenAI 30,2 102,4 t/s 2,18 s
$0.20 in
$1.20 out
32
MiniMax: MiniMax M2.7
minimax/minimax-m2.7
MiniMax 30,1 0 t/s 0 ms
$0.30 in
$1.20 out
33
NVIDIA: Nemotron 3 Ultra
nvidia/nemotron-3-ultra-550b-a55b
NVIDIA 29,6 156,4 t/s 1,21 s
$0.60 in
$2.60 out
34
SpaceXAI: Grok 4.3
x-ai/grok-4.3
xAI 29,3 0 t/s 0 ms
$1.25 in
$2.50 out
35
OpenAI: GPT-5 Codex
openai/gpt-5-codex
OpenAI 29,2 0 t/s 0 ms
$1.25 in
$10.00 out
36
Anthropic: Claude Sonnet 4.6
anthropic/claude-sonnet-4.6
Anthropic 29 0 t/s 0 ms
$3.00 in
$15.00 out
37
Xiaomi: MiMo-V2-Omni
xiaomi/mimo-v2-omni
Xiaomi 28,1 0 t/s 0 ms
$0 in
$0 out
38
OpenAI: GPT-5.1-Codex
openai/gpt-5.1-codex
OpenAI 27,9 0 t/s 0 ms
$1.25 in
$10.00 out
39
Anthropic: Claude Opus 4.5
anthropic/claude-opus-4.5
Anthropic 27,8 0 t/s 0 ms
$5.00 in
$25.00 out
40
Google: Gemini 3.5 Flash Lite
google/gemini-3.5-flash-lite
Google 27,6 315,5 t/s 6,75 s
$0.30 in
$2.50 out
41
inclusionAI: Ling 3.0 Flash
inclusionai/ling-3.0-flash
Inclusionai 27,4 347,2 t/s 1,60 s
$0.075 in
$0.22 out
42
MiniMax: MiniMax M2.5
minimax/minimax-m2.5
MiniMax 26,8 0 t/s 0 ms
$0.30 in
$1.20 out
43
Tencent: Hy3 preview
tencent/hy3-preview
Tencent 26,8 0 t/s 0 ms
$0.063 in
$0.21 out
44
xAI: Grok 4
x-ai/grok-4
xAI 26,5 0 t/s 0 ms
$3.00 in
$15.00 out
45
OpenAI: o3 Pro
openai/o3-pro
OpenAI 25,7 0 t/s 0 ms
$20.00 in
$80.00 out
46
MiniMax: MiniMax M2.1
minimax/minimax-m2.1
MiniMax 24,6 0 t/s 0 ms
$0.30 in
$1.20 out
47
Xiaomi: MiMo-V2-Flash
xiaomi/mimo-v2-flash
Xiaomi 24,5 0 t/s 0 ms
$0.10 in
$0.30 out
48
OpenAI: GPT-5
openai/gpt-5
OpenAI 24,4 0 t/s 0 ms
$1.25 in
$10.00 out
49
OpenAI: GPT-5 Mini
openai/gpt-5-mini
OpenAI 24,2 0 t/s 0 ms
$0.25 in
$2.00 out
50
OpenAI: GPT-5.1-Codex-Mini
openai/gpt-5.1-codex-mini
OpenAI 24 0 t/s 0 ms
$0.25 in
$2.00 out
51
OpenAI: o3
openai/o3
OpenAI 23,7 122,5 t/s 4,41 s
$2.00 in
$8.00 out
52
inclusionAI: Ring-2.6-1T
inclusionai/ring-2.6-1t
Inclusionai 23,7 121,8 t/s 2,06 s
$0.30 in
$2.50 out
53
OpenAI: GPT-5.4 Nano
openai/gpt-5.4-nano
OpenAI 23,5 0 t/s 0 ms
$0.20 in
$1.25 out
54
DeepSeek: DeepSeek V3.1 Terminus
deepseek/deepseek-v3.1-terminus
DeepSeek 23,5 0 t/s 0 ms
$1.635 in
$2.75 out
55
StepFun: Step 3.7 Flash
stepfun/step-3.7-flash
Stepfun 22,9 77,6 t/s 2,23 s
$0.20 in
$1.15 out
56
Anthropic: Claude Sonnet 4.5
anthropic/claude-sonnet-4.5
Anthropic 22,7 0 t/s 0 ms
$3.00 in
$15.00 out
57
Mistral: Mistral Medium 3.5
mistralai/mistral-medium-3-5
Mistral 22,7 139,3 t/s 710 ms
$1.50 in
$7.50 out
58
MiniMax: MiniMax M2
minimax/minimax-m2
MiniMax 21,7 0 t/s 0 ms
$0.30 in
$1.20 out
59
Anthropic: Claude Opus 4.1
anthropic/claude-opus-4.1
Anthropic 21,7 0 t/s 0 ms
$15.00 in
$75.00 out
60
OpenAI: GPT-5.4
openai/gpt-5.4
OpenAI 21,1 0 t/s 0 ms
$2.50 in
$15.00 out
61
xAI: Grok 4 Fast
x-ai/grok-4-fast
xAI 20,8 0 t/s 0 ms
$0.20 in
$0.50 out
62
inclusionAI: Ling-2.6-1T
inclusionai/ling-2.6-1t
Inclusionai 19,6 0 t/s 0 ms
$0.30 in
$2.50 out
63
Tencent: Hy3
tencent/hy3
Tencent 19,6 0 t/s 0 ms
$0.063 in
$0.21 out
64
StepFun: Step 3.5 Flash
stepfun/step-3.5-flash
Stepfun 19,5 0 t/s 0 ms
$0.10 in
$0.30 out
65
Anthropic: Claude Opus 4
anthropic/claude-opus-4
Anthropic 19,1 0 t/s 0 ms
$15.00 in
$75.00 out
66
OpenAI: o4 Mini
openai/o4-mini
OpenAI 19,1 0 t/s 0 ms
$1.10 in
$4.40 out
67
Anthropic: Claude Sonnet 4
anthropic/claude-sonnet-4
Anthropic 19,1 0 t/s 0 ms
$3.00 in
$15.00 out
68
Google: Gemini 2.5 Pro
google/gemini-2.5-pro
Google 19 0 t/s 0 ms
$1.25 in
$10.00 out
69
Google: Gemini 3.1 Flash Lite Preview
google/gemini-3.1-flash-lite-preview
Google 18,7 0 t/s 0 ms
$0.25 in
$1.50 out
70
NVIDIA: Nemotron 3 Super
nvidia/nemotron-3-super-120b-a12b
NVIDIA 18,6 141,9 t/s 1,12 s
$0.20 in
$0.80 out
71
DeepSeek: DeepSeek V3.2
deepseek/deepseek-v3.2
DeepSeek 18,3 0 t/s 0 ms
$0.28 in
$0.42 out
72
Anthropic: Claude Haiku 4.5
anthropic/claude-haiku-4.5
Anthropic 17,4 82,1 t/s 598 ms
$1.00 in
$5.00 out
73
Anthropic: Claude 3.7 Sonnet
anthropic/claude-3.7-sonnet
Anthropic 17,2 0 t/s 0 ms
$3.00 in
$15.00 out
74
OpenAI: o1
openai/o1
OpenAI 17,1 0 t/s 0 ms
$15.00 in
$60.00 out
75
NVIDIA: Nemotron 3.5 Lightning
nvidia/nemotron-3.5-lightning
NVIDIA 16,4 295,6 t/s 461 ms
$0.06 in
$0.20 out
76
xAI: Grok 3 Mini
x-ai/grok-3-mini
xAI 16,2 0 t/s 0 ms
$0.30 in
$0.50 out
77
DeepSeek: DeepSeek V3.2 Speciale
deepseek/deepseek-v3.2-speciale
DeepSeek 16 0 t/s 0 ms
$0 in
$0 out
78
SpaceXAI: Grok 4.20
x-ai/grok-4.20
xAI 15,6 0 t/s 0 ms
$1.25 in
$2.50 out
79
xAI: Grok Code Fast 1
x-ai/grok-code-fast-1
xAI 15,4 0 t/s 0 ms
$0 in
$0 out
80
Inception: Mercury 2
inception/mercury-2
Inception 14,9 794,9 t/s 3,93 s
$0.25 in
$0.75 out
81
OpenAI: GPT-5.1
openai/gpt-5.1
OpenAI 14,2 0 t/s 0 ms
$1.25 in
$10.00 out
82
Google: Gemini 2.5 Flash
google/gemini-2.5-flash
Google 13,9 0 t/s 0 ms
$0.30 in
$2.50 out
83
DeepSeek: R1
deepseek/deepseek-r1
DeepSeek 13,9 0 t/s 0 ms
$1.35 in
$3.00 out
84
OpenAI: GPT-4.1
openai/gpt-4.1
OpenAI 13,2 0 t/s 0 ms
$2.00 in
$8.00 out
85
OpenAI: GPT-5 Nano
openai/gpt-5-nano
OpenAI 12,9 0 t/s 0 ms
$0.05 in
$0.40 out
86
OpenAI: o1-pro
openai/o1-pro
OpenAI 12,8 0 t/s 0 ms
$150.00 in
$600.00 out
87
Perplexity: Sonar Reasoning Pro
perplexity/sonar-reasoning-pro
Perplexity 11,8 0 t/s 0 ms
$0 in
$0 out
88
xAI: Grok 4.1 Fast
x-ai/grok-4.1-fast
xAI 10,9 0 t/s 0 ms
$0 in
$0 out
89
OpenAI: GPT-5.4 Mini
openai/gpt-5.4-mini
OpenAI 10,6 0 t/s 0 ms
$0.75 in
$4.50 out
90
Prime Intellect: INTELLECT-3
prime-intellect/intellect-3
Prime-intellect 9,6 0 t/s 0 ms
$0 in
$0 out
91
OpenAI: o3 Mini
openai/o3-mini
OpenAI 9,6 0 t/s 0 ms
$1.10 in
$4.40 out
92
xAI: Grok 3
x-ai/grok-3
xAI 9,2 0 t/s 0 ms
$0 in
$0 out
93
OpenAI: gpt-oss-120b
openai/gpt-oss-120b
OpenAI 8,9 211,8 t/s 456 ms
$0.15 in
$0.545 out
94
OpenAI: GPT-4.1 Mini
openai/gpt-4.1-mini
OpenAI 8,8 0 t/s 0 ms
$0.40 in
$1.60 out
95
Meta: Llama 4 Maverick
meta-llama/llama-4-maverick
Meta 8,5 84,5 t/s 582 ms
$0.26 in
$0.91 out
96
Upstage: Solar Pro 3
upstage/solar-pro-3
Upstage 8,5 139,4 t/s 1,17 s
$0.15 in
$0.60 out
97
Mistral: Mistral Medium 3.1
mistralai/mistral-medium-3.1
Mistral 8,4 0 t/s 0 ms
$0.40 in
$2.00 out
98
OpenAI: gpt-oss-20b
openai/gpt-oss-20b
OpenAI 8,4 232,4 t/s 553 ms
$0.07 in
$0.20 out
99
inclusionAI: Ling-2.6-flash
inclusionai/ling-2.6-flash
Inclusionai 7,9 0 t/s 0 ms
$0 in
$0 out
100
Google: Gemini 2.5 Flash Lite Preview 09-2025
google/gemini-2.5-flash-lite-preview-09-2025
Google 7,3 0 t/s 0 ms
$0.10 in
$0.40 out
101
Mistral: Mistral Medium 3
mistralai/mistral-medium-3
Mistral 6,7 0 t/s 0 ms
$0.40 in
$2.00 out
102
Anthropic: Claude 3.5 Haiku
anthropic/claude-3.5-haiku
Anthropic 6,5 0 t/s 0 ms
$0 in
$0 out
103
Perplexity: Sonar
perplexity/sonar
Perplexity 5,9 0 t/s 0 ms
$0 in
$0 out
104
OpenAI: GPT-4o
openai/gpt-4o
OpenAI 5,4 0 t/s 0 ms
$2.50 in
$10.00 out
105
DeepSeek: R1 Distill Qwen 32B
deepseek/deepseek-r1-distill-qwen-32b
DeepSeek 5,3 0 t/s 0 ms
$0 in
$0 out
106
Meta: Llama 4 Scout
meta-llama/llama-4-scout
Meta 4,6 85,2 t/s 611 ms
$0.19 in
$0.68 out
107
DeepSeek: R1 Distill Llama 70B
deepseek/deepseek-r1-distill-llama-70b
DeepSeek 4,2 0 t/s 0 ms
$0.70 in
$1.10 out
108
OpenAI: GPT-4.1 Nano
openai/gpt-4.1-nano
OpenAI 4,1 0 t/s 0 ms
$0.10 in
$0.40 out
109
OpenAI: GPT-4o (2024-08-06)
openai/gpt-4o-2024-08-06
OpenAI 3,9 0 t/s 0 ms
$2.50 in
$10.00 out
110
Meta: Llama 3.3 70B Instruct
meta-llama/llama-3.3-70b-instruct
Meta 3,7 90,4 t/s 617 ms
$0.655 in
$0.72 out
111
Mistral: Devstral Small 1.1
mistralai/devstral-small
Mistral 3,6 0 t/s 0 ms
$0 in
$0 out
112
Perplexity: Sonar Pro
perplexity/sonar-pro
Perplexity 3,6 0 t/s 0 ms
$0 in
$0 out
113
NVIDIA: Llama 3.1 Nemotron Ultra 253B v1
nvidia/llama-3.1-nemotron-ultra-253b-v1
NVIDIA 3,4 52,8 t/s 685 ms
$0.60 in
$1.80 out
114
Baidu: ERNIE 4.5 300B A47B
baidu/ernie-4.5-300b-a47b
Baidu 3,4 0 t/s 0 ms
$0.28 in
$1.10 out
115
NVIDIA: Nemotron Nano 12B 2 VL
nvidia/nemotron-nano-12b-v2-vl
NVIDIA 3,3 25,5 t/s 10,10 s
$0.20 in
$0.60 out
116
Google: Gemini 2.0 Flash Lite
google/gemini-2.0-flash-lite-001
Google 3,2 0 t/s 0 ms
$0 in
$0 out
117
NVIDIA: Nemotron Nano 9B V2
nvidia/nemotron-nano-9b-v2
NVIDIA 3,2 94,5 t/s 4,37 s
$0.04 in
$0.16 out
118
OpenAI: GPT-4o (2024-05-13)
openai/gpt-4o-2024-05-13
OpenAI 3 0 t/s 0 ms
$5.00 in
$15.00 out
119
Mistral: Pixtral Large 2411
mistralai/pixtral-large-2411
Mistral 2,5 0 t/s 0 ms
$0 in
$0 out
120
OpenAI: GPT-4 Turbo
openai/gpt-4-turbo
OpenAI 2,3 0 t/s 0 ms
$10.00 in
$30.00 out
121
Cohere: Command A
cohere/command-a
Cohere 2,1 58,8 t/s 323 ms
$2.50 in
$10.00 out
122
NVIDIA: Llama 3.1 Nemotron 70B Instruct
nvidia/llama-3.1-nemotron-70b-instruct
NVIDIA 2,1 19,3 t/s 17,25 s
$1.20 in
$1.20 out
123
Meta: Llama 3.1 8B Instruct
meta-llama/llama-3.1-8b-instruct
Meta 2 0 t/s 0 ms
$0.02 in
$0.05 out
124
NVIDIA: Nemotron 3 Nano 30B A3B
nvidia/nemotron-3-nano-30b-a3b
NVIDIA 1,8 165,8 t/s 400 ms
$0.05 in
$0.20 out
125
Mistral Large 2407
mistralai/mistral-large-2407
Mistral 1,7 0 t/s 0 ms
$2.00 in
$6.00 out
126
OpenAI: GPT-4
openai/gpt-4
OpenAI 1,5 0 t/s 0 ms
$30.00 in
$60.00 out
127
Google: Gemini 2.5 Flash Lite
google/gemini-2.5-flash-lite
Google 1,4 0 t/s 0 ms
$0.10 in
$0.40 out
128
OpenAI: GPT-4o-mini
openai/gpt-4o-mini
OpenAI 1,4 0 t/s 0 ms
$0.15 in
$0.60 out
129
Meta: Llama 3.1 70B Instruct
meta-llama/llama-3.1-70b-instruct
Meta 1,2 0 t/s 0 ms
$0.56 in
$0.56 out
130
Mistral: Mixtral 8x7B Instruct
mistralai/mixtral-8x7b-instruct
Mistral 1 0 t/s 0 ms
$0.45 in
$0.70 out
131
Meta: Llama 3 70B Instruct
meta-llama/llama-3-70b-instruct
Meta 1 0 t/s 0 ms
$0.65 in
$2.75 out
132
Mistral: Saba
mistralai/mistral-saba
Mistral 1 0 t/s 0 ms
$0 in
$0 out
133
Meta: Llama 3 8B Instruct
meta-llama/llama-3-8b-instruct
Meta 1 0 t/s 0 ms
$0.045 in
$0.145 out
134
Mistral Large
mistralai/mistral-large
Mistral 1 0 t/s 0 ms
$4.00 in
$12.00 out
135
Meta: Llama 3.2 11B Vision Instruct
meta-llama/llama-3.2-11b-vision-instruct
Meta 1 26,8 t/s 583 ms
$0.345 in
$0.345 out
136
Meta: Llama 3.2 1B Instruct
meta-llama/llama-3.2-1b-instruct
Meta 1 0 t/s 0 ms
$0 in
$0 out
137
Meta: Llama 3.2 3B Instruct
meta-llama/llama-3.2-3b-instruct
Meta 1 0 t/s 0 ms
$0 in
$0 out
138
Anthropic: Claude 3 Haiku
anthropic/claude-3-haiku
Anthropic 1 0 t/s 0 ms
$0.25 in
$1.25 out
139
Qwen: Qwen3.8 Max (0902)
qwen/qwen3.8-max-0902
Alibaba / Qwen — t/s
in
out
140
OpenAI: GPT-6 Astra (batch)
openai/gpt-6-astra:batch
OpenAI — t/s
in
out
141
OpenAI: GPT-6 Astra Pro
openai/gpt-6-astra-pro
OpenAI — t/s
in
out
142
Elephant
openrouter/elephant-alpha
Openrouter — t/s
in
out
143
inclusionAI: Ring-2.6-1T (free)
inclusionai/ring-2.6-1t:free
Inclusionai — t/s
in
out
144
OpenAI: GPT-6 Astra Pro (batch)
openai/gpt-6-astra-pro:batch
OpenAI — t/s
in
out
145
inclusionAI: Ling 3.0 Flash Sante (free)
inclusionai/ling-3.0-flash-sante:free
Inclusionai — t/s
in
out
146
Nex AGI: Nex-N2-Pro (free)
nex-agi/nex-n2-pro:free
Nex-agi — t/s
in
out
147
Ox Alpha
stealth/ox-alpha
Stealth — t/s
in
out
148
Baidu Qianfan: CoBuddy (free)
baidu/cobuddy:free
Baidu — t/s
in
out
149
DeepSeek: DeepSeek V4 Flash (free)
deepseek/deepseek-v4-flash:free
DeepSeek — t/s
in
out
150
MoonshotAI: Kimi K3
moonshotai/kimi-k3
Moonshot AI — t/s
in
out
151
Meta: Muse Spark 1.3 Contributor
meta/muse-spark-1.3-contributor
Meta — t/s
in
out
152
Inception: Mercury 2.5 Preview
inception/mercury-2.5-preview
Inception — t/s
in
out
153
inclusionAI: Ling 3.0 Tiny (free)
inclusionai/ling-3.0-tiny:free
Inclusionai — t/s
in
out
154
IBM: Granite 4.2 8B
ibm-granite/granite-4.2-8b
Ibm-granite — t/s
in
out
155
Google: Gemini 3.8 Flash (batch)
google/gemini-3.8-flash:batch
Google — t/s
in
out
156
Tencent: Hy4 preview
tencent/hy4-preview
Tencent — t/s
in
out
157
Anthropic: Claude Fable 5.1 (batch)
anthropic/claude-fable-5.1:batch
Anthropic — t/s
in
out
158
Z.ai: GLM Flash Latest
~z-ai/glm-flash-latest
~z-ai — t/s
in
out
159
Tencent: Hy3 (free)
tencent/hy3:free
Tencent — t/s
in
out
160
inclusionAI: Ling 3.0 Flash Fin
inclusionai/ling-3.0-flash-fin
Inclusionai — t/s
in
out
161
inclusionAI: Ling-2.6-flash (free)
inclusionai/ling-2.6-flash:free
Inclusionai — t/s
in
out
162
Owl Alpha
openrouter/owl-alpha
Openrouter — t/s
in
out
163
DeepSeek: DeepSeek V4 Flash Vision Exp
deepseek/deepseek-v4-flash-vision-exp
DeepSeek — t/s
in
out
164
Tencent: Hy-MT2-1.8B
tencent/hy-mt2-1.8b
Tencent — t/s
in
out
165
inclusionAI: Ling-2.6-1T (free)
inclusionai/ling-2.6-1t:free
Inclusionai — t/s
in
out
166
Tencent: Hy3 preview (free)
tencent/hy3-preview:free
Tencent — t/s
in
out
167
Ling-3.0-flash (free)
inclusionai/ling-3.0-flash:free
Inclusionai — t/s
in
out
168
Z.ai: GLM Latest
~z-ai/glm-latest
~z-ai — t/s
in
out
169
inclusionAI: Ling 3.0 Flash Fin (free)
inclusionai/ling-3.0-flash-fin:free
Inclusionai — t/s
in
out
170
Qwen: Qwen3.8 Flash
qwen/qwen3.8-flash
Alibaba / Qwen — t/s
in
out
171
Z.ai: GLM 5.3 Flash
z-ai/glm-5.3-flash
Z-ai — t/s
in
out
172
Z.ai: GLM 5.3 Flash (batch)
z-ai/glm-5.3-flash:batch
Z-ai — t/s
in
out
173
Meta: Muse Spark 1.2 Contributor
meta/muse-spark-1.2-contributor
Meta — t/s
in
out
174
Baidu: Qianfan-OCR-Fast
baidu/qianfan-ocr-fast
Baidu — t/s
in
out
175
Poolside: Laguna XS.2 (free)
poolside/laguna-xs.2:free
Poolside — t/s
in
out
176
Poolside: Laguna XS.2
poolside/laguna-xs.2
Poolside — t/s
in
out
177
Baidu: Qianfan-OCR-Fast (free)
baidu/qianfan-ocr-fast:free
Baidu — t/s
in
out
178
Google: Gemini 3.7 Flash (batch)
google/gemini-3.7-flash:batch
Google — t/s
in
out
179
Arcee AI: Trinity Large Thinking (free)
arcee-ai/trinity-large-thinking:free
Arcee-ai — t/s
in
out
180
Tencent: Hy-MT2-30B-A3B
tencent/hy-mt2-30b-a3b
Tencent — t/s
in
out
181
Tencent: Hy-MT2-7B
tencent/hy-mt2-7b
Tencent — t/s
in
out
182
Z.ai: GLM 5.3
z-ai/glm-5.3
Z-ai — t/s
in
out
183
Qwen: Qwen3.8 27B
qwen/qwen3.8-27b
Alibaba / Qwen — t/s
in
out
184
Dots Studio: Dots3-Note Preview (free)
dots-studio/dots-3-note-preview:free
Dots-studio — t/s
in
out
185
Qwen: Qwen3.6 Plus (free)
qwen/qwen3.6-plus:free
Alibaba / Qwen — t/s
in
out
186
Anthropic: Claude Opus 4.6 (Fast)
anthropic/claude-opus-4.6-fast
Anthropic — t/s
in
out
187
MoonshotAI: Kimi K2.6 (free)
moonshotai/kimi-k2.6:free
Moonshot AI — t/s
in
out
188
Claude Opus 5 (Fast)
anthropic/claude-opus-5-fast
Anthropic — t/s
in
out
189
LiquidAI: LFM2.5-2.6B (free)
liquid/lfm-2.5-2.6b:free
Liquid — t/s
in
out
190
ByteDance Seed: Seed 2.1 Turbo
bytedance-seed/seed-2-1-turbo
Bytedance-seed — t/s
in
out
191
Qwen: Qwen3.8 2.4T A95B
qwen/qwen3.8-2.4t-a95b
Alibaba / Qwen — t/s
in
out
192
Qwen: Qwen3.8 2.4T A95B (batch)
qwen/qwen3.8-2.4t-a95b:batch
Alibaba / Qwen — t/s
in
out
193
ByteDance Seed: Seed-2.0-Code
bytedance-seed/seed-2.0-code
Bytedance-seed — t/s
in
out
194
DeepSeek: DeepSeek V4 Pro 0813
deepseek/deepseek-v4-pro-0813
DeepSeek — t/s
in
out
195
DeepSeek: DeepSeek V4 Pro 0813 (batch)
deepseek/deepseek-v4-pro-0813:batch
DeepSeek — t/s
in
out
196
MiniMax: MiniMax M2.5 (free)
minimax/minimax-m2.5:free
MiniMax — t/s
in
out
197
Arcee AI: Trinity Large Preview
arcee-ai/trinity-large-preview
Arcee-ai — t/s
in
out
198
DeepSeek: DeepSeek V4 Flash 0731
deepseek/deepseek-v4-flash-0731
DeepSeek — t/s
in
out
199
NVIDIA: Nemotron 3.5 Lightning (free)
nvidia/nemotron-3.5-lightning:free
NVIDIA — t/s
in
out
200
AllenAI: Olmo 3.1 32B Instruct
allenai/olmo-3.1-32b-instruct
Allenai — t/s
in
out
🧮 Tool

Pricing-Calculator

Schätze deine monatlichen API-Kosten — gib dein erwartetes Token-Volumen ein, wir rechnen für die Top-15 Modelle.

#ModellVendorQualityKosten/Monat (USD)
📈 Tool

Quality vs. Preis (Pareto-Chart)

Wo liegt der beste Tradeoff? Modelle oben-links sind die Pareto-Optima: hohe Quality, niedriger Preis.

Pareto-Optimum Andere Modelle

❓ Häufige Fragen zum KI-Modelle-Leaderboard

Woher stammen die Leaderboard-Daten?
Quality Index, Speed (Output-Tokens/s), Latenz (Time-to-First-Token) und Preise (USD pro 1 Mio. Tokens) kommen direkt von Artificial Analysis. Synchronisation täglich um 04:00 UTC. Wir speichern keine eigenen Benchmark-Werte und führen keine eigene Bewertung durch.
Was misst der Quality Index genau?
Der Artificial-Analysis-Intelligence-Index ist ein zusammengesetzter Score aus über zehn unabhängigen Benchmarks: MMLU-Pro (Wissen), GPQA & HLE (Reasoning), LiveCodeBench & SciCode (Coding), AIME (Mathematik), IFBench (Instruction-Following), LCR (Long-Context-Recall) und τ² (Tool-Use). Skala 0–100 — höher = besser.
Warum sind manche bekannte Modelle nicht im Leaderboard?
Wir zeigen nur Modelle, die in der Registry als is_active=true markiert sind und für die Artificial Analysis vollständige Benchmark-Daten liefert. Reine Bild- oder Audio-Modelle, deprecatete Versionen sowie Closed-Beta-Modelle ohne öffentliche API erscheinen nicht.
Wie nutze ich Sortierung und Filter sinnvoll?
Sortiere nach Quality für maximale Genauigkeit, nach Speed für hohen Durchsatz, nach Latency für reaktive Streaming-UIs und nach Preis für kostensensitive Workloads. Kombiniere die Vendor- und Preis-Filter, um z. B. „nur OpenAI unter 5 USD/Mio. Output“ zu finden.
Wie berechne ich konkrete Token-Kosten?
Im Token-Kostenrechner kannst du Input- und Output-Volumen einsetzen und die monatlichen Kosten für jedes Modell live durchspielen. Für direkten Head-to-Head-Vergleich zweier Modelle nutze die Vergleichs-Seiten.

Methodik. Der Quality Index ist der Artificial-Analysis-Intelligence-Index (kombiniertes Ranking aus MMLU-Pro, GPQA, HLE, LiveCodeBench, SciCode, AIME, IFBench, LCR und τ²). Speed und Latency sind Median-Werte über alle Provider, Preise pro 1 Mio Tokens (Input/Output). Synchronisation täglich um 04:00 UTC.

PromptLoop ist nicht mit Artificial Analysis verbunden. Alle Marken sind Eigentum ihrer jeweiligen Inhaber.

📬 KI-News direkt ins Postfach