PromptLoop
News Analyse Werkstatt Generative Medien Originals Glossar KI-Modelle Vergleich Kosten-Rechner
📊 Live Benchmark

KI-Modelle Leaderboard 2026

Alle relevanten Large Language Models auf einen Blick — sortiert nach Quality Index, Geschwindigkeit, Latenz und Preis. Datenquelle: Artificial Analysis.

200 aktive Modelle aus 12 Vendor-Familien · Letzte Synchronisation:

Wie liest du dieses Leaderboard?

Quality Index ist ein zusammengesetzter Wert von Artificial Analysis aus über zehn unabhängigen Benchmarks (MMLU-Pro, GPQA, HLE, LiveCodeBench, SciCode, AIME u. a.). Höher = besser. Speed misst Output-Tokens pro Sekunde im Median über alle Hosting-Provider. Latency ist die Time-to-First-Token in Sekunden — wichtig für Streaming-UIs. Preise sind US-Dollar pro 1 Million Tokens, separat für Input und Output. Sortiere nach deiner Priorität, filter nach Anbieter oder Preis-Bucket — und prüfe konkrete Kosten direkt im Token-Rechner oder zwei Modelle im Head-to-Head-Vergleich.

Modalität
# Modell Vendor Quality Speed Latency Preis (USD/1M)
1
Anthropic: Claude Fable 5
anthropic/claude-fable-5
Anthropic 59,9 72,1 t/s 71,29 s
$10.00 in
$50.00 out
2
OpenAI: GPT-5.6 Sol
openai/gpt-5.6-sol
OpenAI 57,7 69,0 t/s 9,13 s
$5.00 in
$30.00 out
3
Anthropic: Claude Opus 4.8
anthropic/claude-opus-4.8
Anthropic 55,7 57,1 t/s 30,42 s
$5.00 in
$25.00 out
4
xAI: Grok 4.5
x-ai/grok-4.5
xAI 53,8 108,9 t/s 13,09 s
$2.00 in
$6.00 out
5
Anthropic: Claude Opus 4.7
anthropic/claude-opus-4.7
Anthropic 53,5 48,5 t/s 17,78 s
$5.00 in
$25.00 out
6
Google: Gemini 3.5 Flash
google/gemini-3.5-flash
Google 50,2 236,7 t/s 11,02 s
$1.50 in
$9.00 out
7
OpenAI: GPT-5.6 Terra
openai/gpt-5.6-terra
OpenAI 49 116,5 t/s 1,65 s
$2.50 in
$15.00 out
8
Google: Gemini 3.1 Pro Preview
google/gemini-3.1-pro-preview
Google 46,5 131,2 t/s 19,75 s
$2.00 in
$12.00 out
9
MiniMax: MiniMax M3
minimax/minimax-m3
MiniMax 44,4 106,3 t/s 1,40 s
$0.30 in
$1.20 out
10
OpenAI: GPT-5.3-Codex
openai/gpt-5.3-codex
OpenAI 44,3 83,3 t/s 81,50 s
$1.75 in
$14.00 out
11
OpenAI: GPT-5.5
openai/gpt-5.5
OpenAI 43,5 60,7 t/s 1,75 s
$5.00 in
$30.00 out
12
OpenAI: GPT-5.2
openai/gpt-5.2
OpenAI 42,2 75,9 t/s 75,51 s
$1.75 in
$14.00 out
13
Anthropic: Claude Sonnet 5
anthropic/claude-sonnet-5
Anthropic 41,7 59,3 t/s 1,29 s
$2.00 in
$10.00 out
14
Xiaomi: MiMo-V2-Pro
xiaomi/mimo-v2-pro
Xiaomi 40,3 0 t/s 0 ms
$1.00 in
$3.00 out
15
OpenAI: GPT-5.2-Codex
openai/gpt-5.2-codex
OpenAI 40,1 129,1 t/s 1,15 s
$1.75 in
$14.00 out
16
OpenAI: GPT-5.6 Luna
openai/gpt-5.6-luna
OpenAI 38,1 217,9 t/s 1,45 s
$1.00 in
$6.00 out
17
MiniMax: MiniMax M2.7
minimax/minimax-m2.7
MiniMax 38,1 49,7 t/s 1,42 s
$0.30 in
$1.20 out
18
Anthropic: Claude Opus 4.6
anthropic/claude-opus-4.6
Anthropic 37,8 42,2 t/s 2,02 s
$5.00 in
$25.00 out
19
NVIDIA: Nemotron 3 Ultra
nvidia/nemotron-3-ultra-550b-a55b
NVIDIA 37,8 170,7 t/s 849 ms
$0.675 in
$2.675 out
20
xAI: Grok 4.3
x-ai/grok-4.3
xAI 37,6 109,5 t/s 36,51 s
$1.25 in
$2.50 out
21
OpenAI: GPT-5 Codex
openai/gpt-5-codex
OpenAI 36,1 162,4 t/s 6,95 s
$1.25 in
$10.00 out
22
Anthropic: Claude Sonnet 4.6
anthropic/claude-sonnet-4.6
Anthropic 35,9 46,5 t/s 1,05 s
$3.00 in
$15.00 out
23
Xiaomi: MiMo-V2-Omni
xiaomi/mimo-v2-omni
Xiaomi 35 0 t/s 0 ms
$0 in
$0 out
24
OpenAI: GPT-5.1-Codex
openai/gpt-5.1-codex
OpenAI 34,7 184,6 t/s 3,56 s
$1.25 in
$10.00 out
25
OpenAI: GPT-5
openai/gpt-5
OpenAI 34,7 86,7 t/s 74,17 s
$1.25 in
$10.00 out
26
Anthropic: Claude Opus 4.5
anthropic/claude-opus-4.5
Anthropic 34,7 53,0 t/s 1,30 s
$5.00 in
$25.00 out
27
MiniMax: MiniMax M2.5
minimax/minimax-m2.5
MiniMax 33,7 75,9 t/s 1,40 s
$0.30 in
$1.20 out
28
Tencent: Hy3
tencent/hy3
Tencent 33,6 196,0 t/s 1,74 s
$0.123 in
$0.43 out
29
xAI: Grok 4
x-ai/grok-4
xAI 33,3 0 t/s 0 ms
$5.50 in
$27.50 out
30
OpenAI: o3 Pro
openai/o3-pro
OpenAI 32,5 24,2 t/s 72,57 s
$20.00 in
$80.00 out
31
DeepSeek: DeepSeek V3.2
deepseek/deepseek-v3.2
DeepSeek 32 0 t/s 0 ms
$0.28 in
$0.42 out
32
MiniMax: MiniMax M2.1
minimax/minimax-m2.1
MiniMax 31,4 77,2 t/s 1,18 s
$0.30 in
$1.20 out
33
DeepSeek: DeepSeek V4 Pro
deepseek/deepseek-v4-pro
DeepSeek 31,2 61,7 t/s 1,05 s
$0.435 in
$0.87 out
34
Xiaomi: MiMo-V2-Flash
xiaomi/mimo-v2-flash
Xiaomi 31,2 0 t/s 0 ms
$0.10 in
$0.30 out
35
OpenAI: GPT-5 Mini
openai/gpt-5-mini
OpenAI 30,9 103,4 t/s 14,60 s
$0.25 in
$2.00 out
36
inclusionAI: Ring-2.6-1T
inclusionai/ring-2.6-1t
Inclusionai 30,6 125,3 t/s 1,81 s
$0.30 in
$2.50 out
37
OpenAI: GPT-5.1-Codex-Mini
openai/gpt-5.1-codex-mini
OpenAI 30,6 211,6 t/s 3,33 s
$0.25 in
$2.00 out
38
DeepSeek: DeepSeek V3.1 Terminus
deepseek/deepseek-v3.1-terminus
DeepSeek 30,4 0 t/s 0 ms
$1.635 in
$2.75 out
39
OpenAI: o3
openai/o3
OpenAI 30,4 107,1 t/s 6,62 s
$2.00 in
$8.00 out
40
StepFun: Step 3.7 Flash
stepfun/step-3.7-flash
Stepfun 30,3 401,5 t/s 733 ms
$0.20 in
$1.15 out
41
OpenAI: GPT-5.4 Nano
openai/gpt-5.4-nano
OpenAI 30,2 167,9 t/s 3,37 s
$0.20 in
$1.25 out
42
Mistral: Mistral Medium 3.5
mistralai/mistral-medium-3-5
Mistral 29,9 92,1 t/s 761 ms
$1.50 in
$7.50 out
43
Anthropic: Claude Sonnet 4.5
anthropic/claude-sonnet-4.5
Anthropic 29,3 45,8 t/s 1,48 s
$3.00 in
$15.00 out
44
DeepSeek: DeepSeek V4 Flash
deepseek/deepseek-v4-flash
DeepSeek 28,7 105 t/s 906 ms
$0.14 in
$0.28 out
45
MiniMax: MiniMax M2
minimax/minimax-m2
MiniMax 28,3 76,5 t/s 1,22 s
$0.30 in
$1.20 out
46
Anthropic: Claude Opus 4.1
anthropic/claude-opus-4.1
Anthropic 28,2 36,5 t/s 2,03 s
$15.00 in
$75.00 out
47
Xiaomi: MiMo-V2.5-Pro
xiaomi/mimo-v2.5-pro
Xiaomi 27,9 63,4 t/s 1,34 s
$0.661 in
$1.722 out
48
OpenAI: GPT-5.4
openai/gpt-5.4
OpenAI 27,7 103,9 t/s 680 ms
$2.50 in
$15.00 out
49
xAI: Grok 4 Fast
x-ai/grok-4-fast
xAI 27,4 0 t/s 0 ms
$0.20 in
$0.50 out
50
inclusionAI: Ling-2.6-1T
inclusionai/ling-2.6-1t
Inclusionai 26,1 0 t/s 0 ms
$0.30 in
$2.50 out
51
StepFun: Step 3.5 Flash
stepfun/step-3.5-flash
Stepfun 26 195,5 t/s 889 ms
$0.10 in
$0.30 out
52
Google: Gemini 2.5 Pro
google/gemini-2.5-pro
Google 25,8 131,1 t/s 20,52 s
$1.25 in
$10.00 out
53
OpenAI: o4 Mini
openai/o4-mini
OpenAI 25,6 161,6 t/s 20,34 s
$1.10 in
$4.40 out
54
Anthropic: Claude Opus 4
anthropic/claude-opus-4
Anthropic 25,5 0 t/s 0 ms
$15.00 in
$75.00 out
55
Anthropic: Claude Sonnet 4
anthropic/claude-sonnet-4
Anthropic 25,5 0 t/s 0 ms
$3.00 in
$15.00 out
56
NVIDIA: Nemotron 3 Super
nvidia/nemotron-3-super-120b-a12b
NVIDIA 25,4 145,9 t/s 933 ms
$0.25 in
$0.775 out
57
Inception: Mercury 2
inception/mercury-2
Inception 25,3 1042,5 t/s 4,10 s
$0.25 in
$0.75 out
58
Google: Gemini 3.1 Flash Lite Preview
google/gemini-3.1-flash-lite-preview
Google 25 308,8 t/s 5,36 s
$0.25 in
$1.50 out
59
OpenAI: gpt-oss-120b
openai/gpt-oss-120b
OpenAI 23,8 293,6 t/s 515 ms
$0.15 in
$0.60 out
60
Anthropic: Claude Haiku 4.5
anthropic/claude-haiku-4.5
Anthropic 23,7 96,9 t/s 781 ms
$1.00 in
$5.00 out
61
Anthropic: Claude 3.7 Sonnet
anthropic/claude-3.7-sonnet
Anthropic 23,5 0 t/s 0 ms
$3.00 in
$15.00 out
62
OpenAI: o1
openai/o1
OpenAI 23,4 138,5 t/s 16,68 s
$15.00 in
$60.00 out
63
xAI: Grok 3 Mini
x-ai/grok-3-mini
xAI 22,5 47,1 t/s 632 ms
$0.30 in
$0.50 out
64
DeepSeek: DeepSeek V3.2 Speciale
deepseek/deepseek-v3.2-speciale
DeepSeek 22,2 0 t/s 0 ms
$0 in
$0 out
65
xAI: Grok 4.20
x-ai/grok-4.20
xAI 21,8 147,3 t/s 440 ms
$2.00 in
$6.00 out
66
xAI: Grok Code Fast 1
x-ai/grok-code-fast-1
xAI 21,6 0 t/s 0 ms
$0 in
$0 out
67
OpenAI: GPT-5.1
openai/gpt-5.1
OpenAI 20,4 97,9 t/s 774 ms
$1.25 in
$10.00 out
68
Google: Gemini 2.5 Flash
google/gemini-2.5-flash
Google 20,1 232,5 t/s 14,08 s
$0.30 in
$2.50 out
69
DeepSeek: R1
deepseek/deepseek-r1
DeepSeek 20,1 0 t/s 0 ms
$1.35 in
$4.20 out
70
OpenAI: GPT-5 Nano
openai/gpt-5-nano
OpenAI 19,9 151,3 t/s 95,13 s
$0.05 in
$0.40 out
71
OpenAI: GPT-4.1
openai/gpt-4.1
OpenAI 19,4 115,5 t/s 686 ms
$2.00 in
$8.00 out
72
OpenAI: o3 Mini
openai/o3-mini
OpenAI 19 221,3 t/s 4,84 s
$1.10 in
$4.40 out
73
OpenAI: o1-pro
openai/o1-pro
OpenAI 18,9 0 t/s 0 ms
$150.00 in
$600.00 out
74
Perplexity: Sonar Reasoning Pro
perplexity/sonar-reasoning-pro
Perplexity 17,8 0 t/s 0 ms
$0 in
$0 out
75
xAI: Grok 4.1 Fast
x-ai/grok-4.1-fast
xAI 16,9 0 t/s 0 ms
$0 in
$0 out
76
OpenAI: GPT-5.4 Mini
openai/gpt-5.4-mini
OpenAI 16,6 162,6 t/s 625 ms
$0.75 in
$4.50 out
77
Prime Intellect: INTELLECT-3
prime-intellect/intellect-3
Prime-intellect 15,6 0 t/s 0 ms
$0 in
$0 out
78
xAI: Grok 3
x-ai/grok-3
xAI 15,1 0 t/s 0 ms
$0 in
$0 out
79
OpenAI: gpt-oss-20b
openai/gpt-oss-20b
OpenAI 14,9 225,6 t/s 406 ms
$0.05 in
$0.20 out
80
OpenAI: GPT-4.1 Mini
openai/gpt-4.1-mini
OpenAI 14,8 73,9 t/s 606 ms
$0.40 in
$1.60 out
81
Mistral: Mistral Medium 3.1
mistralai/mistral-medium-3.1
Mistral 14,7 84,0 t/s 504 ms
$0.40 in
$2.00 out
82
Meta: Llama 4 Maverick
meta-llama/llama-4-maverick
Meta 14,3 114,4 t/s 633 ms
$0.35 in
$0.85 out
83
inclusionAI: Ling-2.6-flash
inclusionai/ling-2.6-flash
Inclusionai 14,1 180,9 t/s 740 ms
$0.10 in
$0.30 out
84
Upstage: Solar Pro 3
upstage/solar-pro-3
Upstage 14,1 0 t/s 0 ms
$0 in
$0 out
85
Google: Gemini 2.5 Flash Lite Preview 09-2025
google/gemini-2.5-flash-lite-preview-09-2025
Google 13,1 0 t/s 0 ms
$0.10 in
$0.40 out
86
Mistral: Mistral Medium 3
mistralai/mistral-medium-3
Mistral 12,5 46,7 t/s 535 ms
$0.40 in
$2.00 out
87
Anthropic: Claude 3.5 Haiku
anthropic/claude-3.5-haiku
Anthropic 12,3 0 t/s 0 ms
$0.80 in
$4.00 out
88
Google: Gemini 2.5 Flash Lite
google/gemini-2.5-flash-lite
Google 11,4 244,4 t/s 19,48 s
$0.10 in
$0.40 out
89
OpenAI: GPT-4o
openai/gpt-4o
OpenAI 11,2 199,7 t/s 598 ms
$2.50 in
$10.00 out
90
DeepSeek: R1 Distill Qwen 32B
deepseek/deepseek-r1-distill-qwen-32b
DeepSeek 11 0 t/s 0 ms
$0 in
$0 out
91
Meta: Llama 4 Scout
meta-llama/llama-4-scout
Meta 10 106,0 t/s 644 ms
$0.17 in
$0.66 out
92
DeepSeek: R1 Distill Llama 70B
deepseek/deepseek-r1-distill-llama-70b
DeepSeek 9,9 29,9 t/s 494 ms
$0.70 in
$1.05 out
93
OpenAI: GPT-4.1 Nano
openai/gpt-4.1-nano
OpenAI 9,6 135,4 t/s 526 ms
$0.10 in
$0.40 out
94
OpenAI: GPT-4o (2024-08-06)
openai/gpt-4o-2024-08-06
OpenAI 9,6 109,8 t/s 616 ms
$2.50 in
$10.00 out
95
Perplexity: Sonar
perplexity/sonar
Perplexity 9,5 0 t/s 0 ms
$0 in
$0 out
96
Meta: Llama 3.3 70B Instruct
meta-llama/llama-3.3-70b-instruct
Meta 9,4 90,2 t/s 620 ms
$0.58 in
$0.71 out
97
Mistral: Devstral Small 1.1
mistralai/devstral-small
Mistral 9,3 0 t/s 0 ms
$0.10 in
$0.30 out
98
Perplexity: Sonar Pro
perplexity/sonar-pro
Perplexity 9,3 0 t/s 0 ms
$0 in
$0 out
99
NVIDIA: Llama 3.1 Nemotron Ultra 253B v1
nvidia/llama-3.1-nemotron-ultra-253b-v1
NVIDIA 9,1 51,8 t/s 714 ms
$0.60 in
$1.80 out
100
Baidu: ERNIE 4.5 300B A47B
baidu/ernie-4.5-300b-a47b
Baidu 9 0 t/s 0 ms
$0.28 in
$1.10 out
101
Google: Gemini 2.0 Flash Lite
google/gemini-2.0-flash-lite-001
Google 8,8 0 t/s 0 ms
$0 in
$0 out
102
NVIDIA: Nemotron Nano 9B V2
nvidia/nemotron-nano-9b-v2
NVIDIA 8,8 82,3 t/s 4,56 s
$0.04 in
$0.16 out
103
OpenAI: GPT-4o (2024-05-13)
openai/gpt-4o-2024-05-13
OpenAI 8,6 90,5 t/s 686 ms
$5.00 in
$15.00 out
104
Mistral: Pixtral Large 2411
mistralai/pixtral-large-2411
Mistral 8,1 0 t/s 0 ms
$2.00 in
$6.00 out
105
OpenAI: GPT-4 Turbo
openai/gpt-4-turbo
OpenAI 7,9 33,6 t/s 1,32 s
$10.00 in
$30.00 out
106
Cohere: Command A
cohere/command-a
Cohere 7,7 56,4 t/s 444 ms
$2.50 in
$10.00 out
107
NVIDIA: Llama 3.1 Nemotron 70B Instruct
nvidia/llama-3.1-nemotron-70b-instruct
NVIDIA 7,6 301 t/s 2,49 s
$1.20 in
$1.20 out
108
Meta: Llama 3.1 8B Instruct
meta-llama/llama-3.1-8b-instruct
Meta 7,6 156,8 t/s 510 ms
$0.075 in
$0.095 out
109
NVIDIA: Nemotron 3 Nano 30B A3B
nvidia/nemotron-3-nano-30b-a3b
NVIDIA 7,4 118,1 t/s 306 ms
$0.05 in
$0.20 out
110
Mistral Large 2407
mistralai/mistral-large-2407
Mistral 7,3 0 t/s 0 ms
$2.00 in
$6.00 out
111
OpenAI: GPT-4
openai/gpt-4
OpenAI 7 31,1 t/s 1,05 s
$30.00 in
$60.00 out
112
OpenAI: GPT-4o-mini
openai/gpt-4o-mini
OpenAI 6,9 70,6 t/s 667 ms
$0.15 in
$0.60 out
113
Meta: Llama 3.1 70B Instruct
meta-llama/llama-3.1-70b-instruct
Meta 6,8 34,1 t/s 606 ms
$0.56 in
$0.56 out
114
Mistral: Saba
mistralai/mistral-saba
Mistral 6,4 0 t/s 0 ms
$0 in
$0 out
115
NVIDIA: Nemotron Nano 12B 2 VL
nvidia/nemotron-nano-12b-v2-vl
NVIDIA 4,6 215,3 t/s 744 ms
$0.20 in
$0.60 out
116
Mistral Large
mistralai/mistral-large
Mistral 4,4 0 t/s 0 ms
$4.00 in
$12.00 out
117
Meta: Llama 3.2 3B Instruct
meta-llama/llama-3.2-3b-instruct
Meta 4,2 52,7 t/s 609 ms
$0.15 in
$0.15 out
118
Anthropic: Claude 3 Haiku
anthropic/claude-3-haiku
Anthropic 3,9 0 t/s 0 ms
$0.25 in
$1.25 out
119
Meta: Llama 3 70B Instruct
meta-llama/llama-3-70b-instruct
Meta 3,5 45,6 t/s 670 ms
$0.65 in
$2.75 out
120
Meta: Llama 3.2 11B Vision Instruct
meta-llama/llama-3.2-11b-vision-instruct
Meta 3,3 67,7 t/s 544 ms
$0.345 in
$0.345 out
121
Mistral: Mixtral 8x7B Instruct
mistralai/mixtral-8x7b-instruct
Mistral 2,4 0 t/s 0 ms
$0.45 in
$0.70 out
122
Meta: Llama 3 8B Instruct
meta-llama/llama-3-8b-instruct
Meta 1,2 78,0 t/s 492 ms
$0.045 in
$0.145 out
123
Meta: Llama 3.2 1B Instruct
meta-llama/llama-3.2-1b-instruct
Meta 1,1 86,6 t/s 592 ms
$0.05 in
$0.05 out
124
OpenAI: GPT-5.6 Luna Pro
openai/gpt-5.6-luna-pro
OpenAI — t/s
in
out
125
OpenAI: GPT-5.6 Terra Pro
openai/gpt-5.6-terra-pro
OpenAI — t/s
in
out
126
Elephant
openrouter/elephant-alpha
Openrouter — t/s
in
out
127
Qwen: Qwen3.7 Max
qwen/qwen3.7-max
Alibaba / Qwen — t/s
in
out
128
OpenAI: GPT-5.6 Sol Pro
openai/gpt-5.6-sol-pro
OpenAI — t/s
in
out
129
xAI: Grok Build 0.1
x-ai/grok-build-0.1
xAI — t/s
in
out
130
Anthropic: Claude Opus 4.7 (Fast)
anthropic/claude-opus-4.7-fast
Anthropic — t/s
in
out
131
inclusionAI: Ring-2.6-1T (free)
inclusionai/ring-2.6-1t:free
Inclusionai — t/s
in
out
132
Poolside: Laguna M.1
poolside/laguna-m.1
Poolside — t/s
in
out
133
Anthropic Claude Haiku Latest
~anthropic/claude-haiku-latest
~anthropic — t/s
in
out
134
Nex AGI: Nex-N2-Pro (free)
nex-agi/nex-n2-pro:free
Nex-agi — t/s
in
out
135
Baidu Qianfan: CoBuddy (free)
baidu/cobuddy:free
Baidu — t/s
in
out
136
DeepSeek: DeepSeek V4 Flash (free)
deepseek/deepseek-v4-flash:free
DeepSeek — t/s
in
out
137
Nex AGI: Nex-N2-Mini
nex-agi/nex-n2-mini
Nex-agi — t/s
in
out
138
Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)
google/gemini-3.1-flash-lite-image
Google — t/s
in
out
139
xAI: Grok Latest
~x-ai/grok-latest
~x-ai — t/s
in
out
140
AionLabs: Aion-3.0-Mini
aion-labs/aion-3.0-mini
Aion-labs — t/s
in
out
141
AionLabs: Aion-3.0
aion-labs/aion-3.0
Aion-labs — t/s
in
out
142
OpenAI GPT Mini Latest
~openai/gpt-mini-latest
~openai — t/s
in
out
143
Google Gemini Pro Latest
~google/gemini-pro-latest
~google — t/s
in
out
144
Tencent: Hy3 (free)
tencent/hy3:free
Tencent — t/s
in
out
145
Poolside: Laguna XS 2.1 (free)
poolside/laguna-xs-2.1:free
Poolside — t/s
in
out
146
Poolside: Laguna XS 2.1
poolside/laguna-xs-2.1
Poolside — t/s
in
out
147
inclusionAI: Ling-2.6-flash (free)
inclusionai/ling-2.6-flash:free
Inclusionai — t/s
in
out
148
Owl Alpha
openrouter/owl-alpha
Openrouter — t/s
in
out
149
inclusionAI: Ling-2.6-1T (free)
inclusionai/ling-2.6-1t:free
Inclusionai — t/s
in
out
150
Tencent: Hy3 preview (free)
tencent/hy3-preview:free
Tencent — t/s
in
out
151
MoonshotAI Kimi Latest
~moonshotai/kimi-latest
~moonshotai — t/s
in
out
152
OpenAI: GPT-5.4 Image 2
openai/gpt-5.4-image-2
OpenAI — t/s
in
out
153
Sakana: Fugu Ultra
sakana/fugu-ultra
Sakana — t/s
in
out
154
Google: Nano Banana 2 (Gemini 3.1 Flash Image)
google/gemini-3.1-flash-image
Google — t/s
in
out
155
Google: Nano Banana Pro (Gemini 3 Pro Image)
google/gemini-3-pro-image
Google — t/s
in
out
156
Cohere: North Mini Code (free)
cohere/north-mini-code:free
Cohere — t/s
in
out
157
Z.ai: GLM 5.2
z-ai/glm-5.2
Z-ai — t/s
in
out
158
OpenRouter: Fusion
openrouter/fusion
Openrouter — t/s
in
out
159
MoonshotAI: Kimi K2.7 Code
moonshotai/kimi-k2.7-code
Moonshot AI — t/s
in
out
160
Anthropic: Claude Fable Latest
~anthropic/claude-fable-latest
~anthropic — t/s
in
out
161
Baidu: Qianfan-OCR-Fast
baidu/qianfan-ocr-fast
Baidu — t/s
in
out
162
Poolside: Laguna XS.2 (free)
poolside/laguna-xs.2:free
Poolside — t/s
in
out
163
Poolside: Laguna XS.2
poolside/laguna-xs.2
Poolside — t/s
in
out
164
Baidu: Qianfan-OCR-Fast (free)
baidu/qianfan-ocr-fast:free
Baidu — t/s
in
out
165
Nex AGI: Nex-N2-Pro
nex-agi/nex-n2-pro
Nex-agi — t/s
in
out
166
Arcee AI: Trinity Large Thinking (free)
arcee-ai/trinity-large-thinking:free
Arcee-ai — t/s
in
out
167
NVIDIA: Nemotron 3.5 Content Safety (free)
nvidia/nemotron-3.5-content-safety:free
NVIDIA — t/s
in
out
168
NVIDIA: Nemotron 3 Ultra (free)
nvidia/nemotron-3-ultra-550b-a55b:free
NVIDIA — t/s
in
out
169
Qwen: Qwen3.7 Plus
qwen/qwen3.7-plus
Alibaba / Qwen — t/s
in
out
170
Anthropic: Claude Opus 4.8 (Fast)
anthropic/claude-opus-4.8-fast
Anthropic — t/s
in
out
171
Perceptron: Perceptron Mk1
perceptron/perceptron-mk1
Perceptron — t/s
in
out
172
Qwen: Qwen3.6 Plus (free)
qwen/qwen3.6-plus:free
Alibaba / Qwen — t/s
in
out
173
Anthropic: Claude Opus 4.6 (Fast)
anthropic/claude-opus-4.6-fast
Anthropic — t/s
in
out
174
MoonshotAI: Kimi K2.6 (free)
moonshotai/kimi-k2.6:free
Moonshot AI — t/s
in
out
175
Google: Gemini 3.1 Flash Lite
google/gemini-3.1-flash-lite
Google — t/s
in
out
176
Poolside: Laguna M.1 (free)
poolside/laguna-m.1:free
Poolside — t/s
in
out
177
OpenAI: GPT Chat Latest
openai/gpt-chat-latest
OpenAI — t/s
in
out
178
IBM: Granite 4.1 8B
ibm-granite/granite-4.1-8b
Ibm-granite — t/s
in
out
179
NVIDIA: Nemotron 3 Nano Omni (free)
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free
NVIDIA — t/s
in
out
180
OpenAI GPT Latest
~openai/gpt-latest
~openai — t/s
in
out
181
Google Gemini Flash Latest
~google/gemini-flash-latest
~google — t/s
in
out
182
Tencent: Hy3 preview
tencent/hy3-preview
Tencent — t/s
in
out
183
Qwen: Qwen3.5 Plus 2026-04-20
qwen/qwen3.5-plus-20260420
Alibaba / Qwen — t/s
in
out
184
MiniMax: MiniMax M2.5 (free)
minimax/minimax-m2.5:free
MiniMax — t/s
in
out
185
Arcee AI: Trinity Large Preview
arcee-ai/trinity-large-preview
Arcee-ai — t/s
in
out
186
Anthropic Claude Sonnet Latest
~anthropic/claude-sonnet-latest
~anthropic — t/s
in
out
187
AllenAI: Olmo 3.1 32B Instruct
allenai/olmo-3.1-32b-instruct
Allenai — t/s
in
out
188
Qwen: Qwen3.6 Flash
qwen/qwen3.6-flash
Alibaba / Qwen — t/s
in
out
189
Qwen: Qwen3.6 35B A3B
qwen/qwen3.6-35b-a3b
Alibaba / Qwen — t/s
in
out
190
Xiaomi: MiMo-V2.5
xiaomi/mimo-v2.5
Xiaomi — t/s
in
out
191
Qwen: Qwen3.6 Max Preview
qwen/qwen3.6-max-preview
Alibaba / Qwen — t/s
in
out
192
Qwen: Qwen3.6 27B
qwen/qwen3.6-27b
Alibaba / Qwen — t/s
in
out
193
Google: Gemma 4 31B (free)
google/gemma-4-31b-it:free
Google — t/s
in
out
194
Google: Gemma 4 26B A4B (free)
google/gemma-4-26b-a4b-it:free
Google — t/s
in
out
195
Anthropic: Claude Opus Latest
~anthropic/claude-opus-latest
~anthropic — t/s
in
out
196
Google: Gemma 4 26B A4B
google/gemma-4-26b-a4b-it
Google — t/s
in
out
197
Pareto Code Router
openrouter/pareto-code
Openrouter — t/s
in
out
198
MoonshotAI: Kimi K2.6
moonshotai/kimi-k2.6
Moonshot AI — t/s
in
out
199
Z.ai: GLM 5.1
z-ai/glm-5.1
Z-ai — t/s
in
out
200
Meituan: LongCat Flash Chat
meituan/longcat-flash-chat
Meituan — t/s
in
out
🧮 Tool

Pricing-Calculator

Schätze deine monatlichen API-Kosten — gib dein erwartetes Token-Volumen ein, wir rechnen für die Top-15 Modelle.

#ModellVendorQualityKosten/Monat (USD)
📈 Tool

Quality vs. Preis (Pareto-Chart)

Wo liegt der beste Tradeoff? Modelle oben-links sind die Pareto-Optima: hohe Quality, niedriger Preis.

Pareto-Optimum Andere Modelle

❓ Häufige Fragen zum KI-Modelle-Leaderboard

Woher stammen die Leaderboard-Daten?
Quality Index, Speed (Output-Tokens/s), Latenz (Time-to-First-Token) und Preise (USD pro 1 Mio. Tokens) kommen direkt von Artificial Analysis. Synchronisation täglich um 04:00 UTC. Wir speichern keine eigenen Benchmark-Werte und führen keine eigene Bewertung durch.
Was misst der Quality Index genau?
Der Artificial-Analysis-Intelligence-Index ist ein zusammengesetzter Score aus über zehn unabhängigen Benchmarks: MMLU-Pro (Wissen), GPQA & HLE (Reasoning), LiveCodeBench & SciCode (Coding), AIME (Mathematik), IFBench (Instruction-Following), LCR (Long-Context-Recall) und τ² (Tool-Use). Skala 0–100 — höher = besser.
Warum sind manche bekannte Modelle nicht im Leaderboard?
Wir zeigen nur Modelle, die in der Registry als is_active=true markiert sind und für die Artificial Analysis vollständige Benchmark-Daten liefert. Reine Bild- oder Audio-Modelle, deprecatete Versionen sowie Closed-Beta-Modelle ohne öffentliche API erscheinen nicht.
Wie nutze ich Sortierung und Filter sinnvoll?
Sortiere nach Quality für maximale Genauigkeit, nach Speed für hohen Durchsatz, nach Latency für reaktive Streaming-UIs und nach Preis für kostensensitive Workloads. Kombiniere die Vendor- und Preis-Filter, um z. B. „nur OpenAI unter 5 USD/Mio. Output“ zu finden.
Wie berechne ich konkrete Token-Kosten?
Im Token-Kostenrechner kannst du Input- und Output-Volumen einsetzen und die monatlichen Kosten für jedes Modell live durchspielen. Für direkten Head-to-Head-Vergleich zweier Modelle nutze die Vergleichs-Seiten.

Methodik. Der Quality Index ist der Artificial-Analysis-Intelligence-Index (kombiniertes Ranking aus MMLU-Pro, GPQA, HLE, LiveCodeBench, SciCode, AIME, IFBench, LCR und τ²). Speed und Latency sind Median-Werte über alle Provider, Preise pro 1 Mio Tokens (Input/Output). Synchronisation täglich um 04:00 UTC.

PromptLoop ist nicht mit Artificial Analysis verbunden. Alle Marken sind Eigentum ihrer jeweiligen Inhaber.

📬 KI-News direkt ins Postfach