AI 大模型排行榜 (Artificial Analysis LLM Ranking)

信息查询
2.3k 次浏览
100% 有帮助 · 1 人反馈

值品工具箱提供的 AI 大模型排行榜聚合了来自 Artificial Analysis 的权威数据,实时追踪并排名超过 100 个主流大语言模型。

AI 大模型排行榜数据中心

重置
排名 模型名称 综合指数 ▼ 编程 价格 ($/1M)
1 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) 59.9 76.5 $20
2 GPT-5.6 Sol (max) 58.9 77.4 $11.25
3 GPT-5.6 Sol (xhigh) 57.7 78.3 $11.25
4 GPT-5.6 Sol (high) 55.9 77.2 $11.25
5 Claude Opus 4.8 (Adaptive Reasoning, Max Effort) 55.7 74.3 $10
6 GPT-5.6 Terra (max) 55 76.7 $5.625
7 GPT-5.5 (xhigh) 54.8 74.9 $11.25
8 Grok 4.5 (high) 53.8 72.4 $3
9 GPT-5.6 Sol (medium) 53.6 76.3 $11.25
10 Claude Opus 4.7 (Adaptive Reasoning, Max Effort) 53.5 73.6 $10
11 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) 53.4 71.5 $4
12 GPT-5.5 (high) 53.1 71.6 $11.25
13 GPT-5.6 Terra (xhigh) 51.6 70.6 $5.625
14 GPT-5.4 (xhigh) 51.4 71.1 $5.625
15 GPT-5.6 Luna (max) 51.2 71.4 $2.25
16 GLM-5.2 (max) 51.1 68.8 $2.15
17 Muse Spark 1.1 (xhigh) 50.6 71.3 $2
18 GPT-5.5 (medium) 50.4 71.5 $11.25
19 Gemini 3.5 Flash (high) 50.2 70.1 $3.375
20 GPT-5.6 Sol (low) 49.4 69.7 $11.25
21 GPT-5.6 Luna (xhigh) 49.1 68.6 $2.25
22 GPT-5.6 Terra (high) 49 67.1 $5.625
23 Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) 47.2 63 $6
24 Gemini 3.1 Pro Preview 46.5 68.8 $4.5
25 GPT-5.6 Luna (high) 46.1 63.3 $2.25
26 Qwen3.7 Max 46 66 $3.75
27 GPT-5.6 Terra (medium) 45.6 64.7 $5.625
28 Gemini 3.5 Flash (medium) 45.4 - $3.375
29 MiniMax-M3 44.4 58.6 $0.525
30 GPT-5.3 Codex (xhigh) 44.3 - $4.813
31 DeepSeek V4 Pro (Reasoning, Max Effort) 44.3 59.4 $0.544
32 Kimi K2.6 44.2 61.8 $1.712
33 Claude Opus 4.6 (Adaptive Reasoning, Max Effort) 43.7 - $10
34 GPT-5.5 (low) 43.5 60.9 $11.25
35 Muse Spark 43.1 58.6 $0
36 Claude Opus 4.7 (Non-reasoning, High Effort) 42.7 - $10
37 MiMo-V2.5-Pro 42.2 60.2 $0.544
38 GPT-5.2 (xhigh) 42.2 - $4.813
39 Kimi K2.7 Code 41.9 60.8 $1.712
40 Claude Sonnet 5 (Non-reasoning, High Effort) 41.7 66.4 $4
41 GPT-5.6 Sol (Non-reasoning) 41.2 65.1 $11.25
42 Nex-N2-Pro 41 59.1 $1
43 DeepSeek V4 Pro (Reasoning, High Effort) 40.8 - $0.544
44 Claude Opus 4.5 (Reasoning) 40.8 - $10
45 GPT-5.6 Terra (low) 40.5 58.1 $5.625
46 DeepSeek V4 Flash (Reasoning, Max Effort) 40.3 56.2 $0.175
47 MiMo-V2-Pro 40.3 - $1.5
48 GLM-5.1 (Reasoning) 40.2 55.8 $2.15
49 GPT-5.2 Codex (xhigh) 40.1 - $4.813
50 GPT-5.4 mini (xhigh) 40 56.1 $1.688
51 Qwen3.6 Max Preview 40 - $2.925
52 Grok Build 0.1 0616 39.8 51.5 $1.25
53 Qwen3.6 Plus 39.6 54.5 $1.125
54 Gemini 3 Pro Preview (high) 39.6 - $4.5
55 GLM-5 (Reasoning) 39.5 - $1.55
56 GPT-5.4 (low) 39.1 - $5.625
57 Qwen3.7 Plus 39 55.9 $0.7
58 JT-4.1 Flash 236B A21B 38.8 52.4 $0
59 GPT-5.4 nano (xhigh) 38.2 56.1 $0.463
60 GPT-5.6 Luna (medium) 38.1 50.7 $2.25
61 MiniMax-M2.7 38.1 52.6 $0.525
62 GLM-5-Turbo 38.1 - $0
63 Kimi K2.5 (Reasoning) 38.1 - $1.2
64 GPT-5.2 (medium) 38 - $4.813
65 Nemotron 3 Ultra 550B A55B (Reasoning) 37.8 49.3 $1.175
66 Gemini 3 Flash Preview (Reasoning) 37.8 - $1.125
67 Claude Opus 4.6 (Non-reasoning, High Effort) 37.8 - $10
68 Grok 4.3 (high) 37.6 42.2 $1.563
69 DeepSeek V4 Flash (Reasoning, High Effort) 37.4 - $0.175
70 MiMo-V2.5 37.2 56.8 $0.175
71 Qwen3.6 27B (Reasoning) 37.1 53.7 $1.35
72 Grok 4.20 0309 v2 (Reasoning) 37 - $3
73 GPT-5.1 (high) 36.9 49.4 $3.438
74 Grok 4.20 0309 (Reasoning) 36.5 - $3
75 MiMo-V2-Omni-0327 36.4 - $0.8
76 Claude 4.5 Sonnet (Reasoning) 36.4 52.1 $6
77 GPT-5 Codex (high) 36.1 - $3.438
78 Grok 4.3 (medium) 36 - $1.563
79 Claude Sonnet 4.6 (Non-reasoning, High Effort) 35.9 - $6
80 GPT-5.5 (Non-reasoning) 35.4 56.5 $11.25
81 Grok 4.3 (low) 35.4 - $1.563
82 KAT Coder Pro V2 35.4 - $0.525
83 GLM-5.1 (Non-reasoning) 35.4 - $2.15
84 MiMo-V2-Omni 35 - $0
85 Gemini 3.5 Flash (minimal) 34.9 - $3.375
86 GPT-5 (high) 34.7 37.8 $3.438
87 GPT-5.1 Codex (high) 34.7 - $3.438
88 Claude Opus 4.5 (Non-reasoning) 34.7 - $10
89 Kimi K2.6 (Non-reasoning) 34.6 - $1.712
90 KAT-Coder-Pro V1 34.6 58.9 $0.525
91 GLM 5V Turbo (Reasoning) 34.5 - $0
92 Claude Sonnet 4.6 (Non-reasoning, Low Effort) 34.3 - $6
93 GLM-5.2 (Non-reasoning) 34.1 46.5 $3.162
94 GPT-5.6 Terra (Non-reasoning) 34 52.3 $5.625
95 Qwen3.5 27B (Reasoning) 33.8 - $0.825
96 Qwen3.5 397B A17B (Reasoning) 33.7 48.2 $1.35
97 GPT-5 (medium) 33.7 - $3.438
98 Claude 4.1 Opus (Reasoning) 33.7 - $30
99 MiniMax-M2.5 33.7 - $0.525
100 GLM-4.7 (Reasoning) 33.7 45.3 $1
101 Hy3-preview (Reasoning) 33.6 - $0.2
102 GPT-5.5 Instant (May 2026) 33.5 - $11.25
103 GPT-5.6 Luna (low) 33.3 44.2 $2.25
104 Grok 4 33.3 - $11
105 MiMo-V2-Flash (Feb 2026) 33.2 - $0.15
106 Gemini 3 Pro Preview (low) 33.1 - $4.5
107 Kimi K2 Thinking 32.7 - $1.075
108 o3-pro 32.5 - $35
109 GLM-5 (Non-reasoning) 32.4 - $1.55
110 Qwen3.5 122B A10B (Reasoning) 32.3 45.7 $1.1
111 Qwen3.5 397B A17B (Non-reasoning) 32 - $1.35
112 DeepSeek V3.2 (Reasoning) 32 44.2 $0.315
113 Qwen3 Max Thinking 31.7 - $0
114 Qwen3.6 35B A3B (Reasoning) 31.6 41.9 $0.557
115 MiniMax-M2.1 31.4 - $0.525
116 DeepSeek V4 Pro (Non-reasoning) 31.2 - $0.544
117 GPT-5 (low) 31.2 - $3.438
118 MiMo-V2-Flash (Reasoning) 31.2 - $0.15
119 Claude 4 Opus (Reasoning) 31 - $30
120 GPT-5 mini (medium) 30.9 - $0.688
121 Qwen3.5 Omni Plus 30.6 - $1.5
122 Ring-2.6-1T 30.6 42.8 $0.85
123 GPT-5.1 Codex mini (high) 30.6 - $0.688
124 Grok 4.1 Fast (Reasoning) 30.6 - $0
125 Qwen3.6 27B (Non-reasoning) 30.5 46.6 $1.35
126 o3 30.4 - $3.5
127 DeepSeek V3.1 Terminus (Reasoning) 30.4 43.5 $1.914
128 Step 3.7 Flash 30.3 39.6 $0.438
129 GPT-5.4 nano (medium) 30.2 - $0.463
130 Mistral Medium 3.5 29.9 46.9 $3
131 GPT-5.4 mini (medium) 29.8 - $1.688
132 Claude 4.5 Haiku (Reasoning) 29.6 43.9 $2
133 Gemma 4 31B (Reasoning) 29.4 43.4 $0
134 Kimi K2.5 (Non-reasoning) 29.4 - $1.2
135 Claude 4.5 Sonnet (Non-reasoning) 29.3 - $6
136 Qwen3.5 35B A3B (Reasoning) 29.3 - $0.688
137 Qwen3.5 27B (Non-reasoning) 29.3 - $0.825
138 GPT-5.5 Instant (June 2026) 28.9 39.4 $11.25
139 Claude 4 Sonnet (Reasoning) 28.9 37.6 $6
140 DeepSeek V4 Flash (Non-reasoning) 28.7 - $0.175
141 GLM-4.6 (Reasoning) 28.7 45.8 $0.963
142 JT-35B-Flash 28.4 - $0
143 MiniMax-M2 28.3 - $0.525
144 Claude 4.1 Opus (Non-reasoning) 28.2 - $30
145 MiMo-V2.5-Pro (Non-reasoning) 27.9 - $0.926
146 GPT-5.4 (Non-reasoning) 27.7 - $5.625
147 Qwen3.5 122B A10B (Non-reasoning) 27.6 43.3 $1.1
148 Gemini 3 Flash Preview (Non-reasoning) 27.4 - $1.125
149 Grok 4 Fast (Reasoning) 27.4 - $0.275
150 Claude 3.7 Sonnet (Reasoning) 27.1 36.4 $0
151 GPT-5.6 Luna (Non-reasoning) 26.6 39.3 $2.25
152 GLM-4.7 (Non-reasoning) 26.6 - $1
153 Hy3-preview (Non-reasoning) 26.1 - $0.2
154 Ling-2.6-1T 26.1 - $0.85
155 Step 3.5 Flash 2603 26 - $0.15
156 Doubao Seed Code 26 - $0
157 GPT-5.2 (Non-reasoning) 26 - $4.813
158 Gemini 2.5 Pro 25.8 33.3 $3.438
159 Gemma 4 26B A4B (Reasoning) 25.7 39.3 $0.198
160 o4-mini (high) 25.6 - $1.925
161 Claude 4 Sonnet (Non-reasoning) 25.5 - $6
162 Claude 4 Opus (Non-reasoning) 25.5 - $30
163 Step 3.5 Flash 25.5 - $0.15
164 NVIDIA Nemotron 3 Super 120B A12B (Reasoning) 25.4 37.7 $0.381
165 DeepSeek V3.2 Exp (Reasoning) 25.4 - $0.315
166 Mercury 2 25.3 - $0.375
167 GPT-5 mini (high) 25.3 15.6 $0.688
168 Gemini 3.1 Flash-Lite 25 34.7 $0.563
169 Qwen3 Max Thinking (Preview) 25 - $2.4
170 Grok 4.3 (Non-reasoning) 24.8 35.2 $1.563
171 K-EXAONE (Reasoning) 24.7 - $0
172 MiMo-V2-Flash (Non-reasoning) 24.7 49.8 $0.15
173 DeepSeek V3.2 (Non-reasoning) 24.7 - $0.315
174 Trinity Large Thinking 24.5 - $0.395
175 Qwen3.6 35B A3B (Non-reasoning) 24.2 - $0.844
176 Qwen3 Max 24 - $2.4
177 gpt-oss-120b (high) 23.8 30.4 $0.262
178 Gemini 2.5 Flash Preview (Sep '25) (Reasoning) 23.8 - $0
179 Claude 4.5 Haiku (Non-reasoning) 23.7 - $2
180 Claude 3.7 Sonnet (Non-reasoning) 23.5 - $6
181 Kimi K2 0905 23.5 - $1.075
182 Qwen3.5 35B A3B (Non-reasoning) 23.4 - $0.688
183 o1 23.4 39.7 $26.25
184 EXAONE 4.5 33B 23 - $0
185 Gemini 2.5 Pro Preview (Mar' 25) 23 46.7 $0
186 GLM-4.6 (Non-reasoning) 23 - $0.981
187 GLM-4.7-Flash (Reasoning) 22.9 - $0.153
188 Command A+ 22.5 27.8 $0
189 Grok 3 mini Reasoning (high) 22.5 - $0.35
190 Grok 4.20 0309 (Non-reasoning) 22.5 - $3
191 Gemini 2.5 Pro Preview (May' 25) 22.3 - $3.438
192 DeepSeek V3.2 Speciale 22.2 - $0
193 Gemma 4 12B (Reasoning) 22 - $0.15
194 ERNIE 5.0 Thinking Preview 21.9 - $0
195 Gemma 4 31B (Non-reasoning) 21.8 33.2 $0.205
196 Nova 2.0 Pro Preview (medium) 21.8 34 $3.438
197 Grok 4.20 0309 v2 (Non-reasoning) 21.8 - $3
198 Grok Code Fast 1 21.6 - $0
199 Qwen3.5 9B (Reasoning) 21.4 28.7 $0.113
200 DeepSeek V3.1 Terminus (Non-reasoning) 21.4 - $0.453
201 Nemotron Cascade 2 30B A3B 21.3 - $0
202 DeepSeek V3.2 Exp (Non-reasoning) 21.3 - $0.315
203 Apriel-v1.5-15B-Thinker 21.2 - $0
204 Qwen3 Coder Next 21.1 36.2 $0.563
205 DeepSeek V3.1 (Non-reasoning) 21 - $0.84
206 Nova 2.0 Omni (medium) 20.9 - $0.85
207 DeepSeek V3.1 (Reasoning) 20.7 - $0.865
208 North Mini Code 20.6 36.5 $0
209 Qwen3 VL 235B A22B (Reasoning) 20.6 - $2.625
210 Apriel-v1.6-15B-Thinker 20.5 - $0
211 GPT-5.1 (Non-reasoning) 20.4 - $3.438
212 Qwen3.5 9B (Non-reasoning) 20.3 23.5 $0
213 Gemma 4 26B A4B (Non-reasoning) 20.1 - $0.198
214 Qwen3.5 4B (Reasoning) 20.1 22.6 $0.06
215 Gemini 2.5 Flash (Reasoning) 20.1 - $0.85
216 DeepSeek R1 0528 (May '25) 20.1 - $2.063
217 GPT-5 nano (high) 19.9 - $0.138
218 Mistral Small 4 (Reasoning) 19.6 26.6 $0.262
219 Nova 2.0 Pro Preview (low) 19.6 25.9 $3.438
220 Qwen3 235B A22B 2507 (Reasoning) 19.6 22.1 $2.625
221 GLM-4.5 (Reasoning) 19.5 - $1
222 GPT-4.1 19.4 - $3.5
223 Kimi K2 19.4 - $1.039
224 Devstral 2 19.2 31.3 $0
225 Qwen3 Max (Preview) 19.2 - $2.4
226 Nova 2.0 Lite (medium) 19 - $0.85
227 Qwen3.5 Omni Flash 19 - $0.275
228 o3-mini 19 - $1.925
229 GPT-5 nano (medium) 19 - $0.138
230 o1-pro 18.9 - $262.5
231 Gemini 2.5 Flash Preview (Sep '25) (Non-reasoning) 18.8 - $0
232 JT-MINI 18.5 - $0
233 DeepSeek R1 (Jan '25) 18.5 24.6 $2.431
234 Grok 3 18.4 - $8
235 Seed-OSS-36B-Instruct 18.3 - $0.3
236 Nova 2.0 Lite (high) 18.2 23 $0.85
237 Qwen3 235B A22B 2507 Instruct 18.2 - $1.225
238 Qwen3 Coder 480B A35B Instruct 18 - $3
239 Magistral Medium 1.2 17.9 21.3 $2.75
240 Qwen3 VL 32B (Reasoning) 17.9 - $2.625
241 Nova 2.0 Lite (low) 17.8 - $0.85
242 HyperNova 60B 2605 17.8 23.2 $0.065
243 Sonar Reasoning Pro 17.8 - $0
244 gpt-oss-120b (low) 17.7 - $0.262
245 MiniMax M1 80k 17.7 - $0.963
246 GPT-5.4 nano (Non-Reasoning) 17.6 - $0.463
247 Gemini 2.5 Flash Preview (Reasoning) 17.5 - $0
248 Devstral Small 2 17.4 29.3 $0
249 K2 Think V2 17.3 21 $0
250 LongCat Flash Lite 17.2 - $0
251 GPT-5 (minimal) 17.2 - $3.438
252 HyperCLOVA X SEED Think (32B) 17 - $0
253 o1-preview 17 34 $28.875
254 Grok 4.1 Fast (Non-reasoning) 16.9 - $0
255 GLM-4.6V (Reasoning) 16.8 - $0.45
256 K-EXAONE (Non-reasoning) 16.7 - $0
257 Qwen3 Next 80B A3B (Reasoning) 16.7 17.4 $1.875
258 GPT-5.4 mini (Non-Reasoning) 16.6 - $1.688
259 Nova 2.0 Omni (low) 16.6 - $0.85
260 Grok 4 Fast (Non-reasoning) 16.5 - $0.275
261 GLM-4.5-Air 16.5 - $0.372
262 Mi:dm K 2.5 Pro 16.4 - $0
263 Ring-1T 16.2 - $0
264 Qwen3.5 4B (Non-reasoning) 16 20.3 $0.06
265 Mistral Large 3 15.9 20.1 $0.75
266 INTELLECT-3 15.6 - $0
267 o3-mini (high) 15.6 16.3 $1.925
268 GLM-4.7-Flash (Non-reasoning) 15.5 - $0.153
269 DeepSeek V3 0324 15.4 21.2 $1.209
270 GPT-5 (ChatGPT) 15.3 - $3.438
271 Solar Open 100B (Reasoning) 15.1 - $0
272 Gemini 2.5 Flash-Lite Preview (Sep '25) (Reasoning) 15.1 - $0.175
273 Grok 3 Reasoning Beta 15.1 - $0
274 gpt-oss-20b (high) 14.9 20.7 $0.088
275 Nemotron 3 Nano Omni 30B A3B Reasoning 14.9 - $0.131
276 GPT-4.1 mini 14.8 20.2 $0.7
277 Mistral Small 3.1 14.7 26.3 $0.15
278 Mistral Medium 3.1 14.7 20.5 $0.8
279 Nova 2.0 Pro Preview (Non-reasoning) 14.4 20.9 $3.438
280 MiniMax M1 40k 14.4 - $0
281 Qwen3 30B A3B 2507 (Reasoning) 14.4 12.1 $0.75
282 gpt-oss-20b (low) 14.3 - $0.095
283 Llama 4 Maverick 14.3 16.3 $0.475
284 GPT-5 mini (minimal) 14.3 - $0.688
285 Qwen3 VL 235B A22B Instruct 14.3 - $1.225
286 NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) 14.2 14.4 $0.088
287 K2-V2 (high) 14.2 - $0
288 DeepSeek V3 (Dec '24) 14.2 23 $0.523
289 Solar Pro 3 14.1 16.2 $0
290 Ling 2.6 Flash 14.1 25.3 $0.15
291 Gemini 2.5 Flash (Non-reasoning) 14.1 - $0.85
292 o1-mini 14 - $0
293 Qwen3 Next 80B A3B Instruct 13.7 - $0.875
294 Tri-21B-think Preview 13.6 - $0
295 GPT-4.5 (Preview) 13.6 - $0
296 Qwen3 Coder 30B A3B Instruct 13.6 - $0.9
297 DiffusionGemma 26B A4B 13.5 19.7 $0
298 QwQ 32B 13.4 - $0.745
299 Qwen3 235B A22B (Reasoning) 13.4 - $2.625
300 Gemini 2.0 Flash Thinking Experimental (Jan '25) 13.3 24.1 $0
301 Qwen3 VL 30B A3B (Reasoning) 13.3 - $0.75
302 Gemma 4 12B (Non-reasoning) 13.2 - $0.15
303 Gemini 2.5 Flash-Lite Preview (Sep '25) (Non-reasoning) 13.1 - $0.175
304 Motif-2-12.7B-Reasoning 12.8 - $0
305 Ling-1T 12.8 - $0
306 Nova Premier 12.7 - $5
307 Gemma 4 E4B (Reasoning) 12.5 - $0
308 Magistral Medium 1 12.5 - $0
309 Mistral Medium 3 12.5 - $0.8
310 Solar Pro 2 (Preview) (Reasoning) 12.5 - $0
311 Mistral Small 4 (Non-reasoning) 12.4 - $0.262
312 Llama Nemotron Super 49B v1.5 (Reasoning) 12.4 - $0.4
313 K2-V2 (medium) 12.4 - $0
314 Tri-21B-Think 12.4 - $0
315 Devstral Medium 12.4 - $0.8
316 GPT-4o (March 2025, chatgpt-4o-latest) 12.3 - $0
317 Gemini 2.0 Flash (Feb '25) 12.3 - $0.262
318 Claude 3.5 Haiku 12.3 15.9 $1.6
319 Llama 3.3 Nemotron Super 49B v1 (Reasoning) 12.2 - $0
320 MiniCPM5-1B (Reasoning) 12 - $0
321 Qwen3 4B 2507 (Reasoning) 12 - $0
322 Sarvam 105B (high) 11.9 - $0.074
323 Nova 2.0 Lite (Non-reasoning) 11.8 - $0.85
324 Gemini 2.0 Pro Experimental (Feb '25) 11.8 25.5 $0
325 Claude 3 Opus 11.8 19.5 $30
326 Devstral Small (May '25) 11.8 - $0
327 MiniCPM5-1B (Non-reasoning) 11.7 - $0
328 Gemini 2.5 Flash Preview (Non-reasoning) 11.7 - $0
329 Sonar Reasoning 11.7 - $0
330 Qwen3 32B (Reasoning) 11.5 15.3 $2.625
331 Gemini 2.5 Flash-Lite (Reasoning) 11.4 - $0.175
332 Magistral Small 1.2 11.3 14.7 $0.75
333 GPT-4o (Nov '24) 11.2 - $4.375
334 Ministral 3 14B 11.1 14.4 $0.2
335 Nanbeige4.1-3B 11.1 9.6 $0
336 Qwen3 VL 32B Instruct 11.1 - $1.225
337 DeepSeek R1 Distill Qwen 32B 11 - $0
338 GLM-4.6V (Non-reasoning) 11 - $0.45
339 Qwen3 235B A22B (Non-reasoning) 10.9 - $1.225
340 Gemini 2.0 Flash (experimental) 10.7 - $0
341 Magistral Small 1 10.7 - $0
342 EXAONE 4.0 32B (Reasoning) 10.6 - $0
343 Mistral Small 3.2 10.6 12.5 $0.15
344 Qwen3 VL 8B (Reasoning) 10.6 - $0.66
345 Nova 2.0 Omni (Non-reasoning) 10.5 - $0.85
346 DeepSeek R1 0528 Qwen3 8B 10.4 - $0
347 Qwen3 14B (Reasoning) 10.4 13.8 $1.313
348 Qwen2.5 Max 10.2 - $0
349 Llama 4 Scout 10 8.2 $0.292
350 Hermes 4 - Llama-3.1 70B (Reasoning) 10 - $0.198
351 Gemini 1.5 Pro (Sep '24) 10 23.6 $0
352 Solar Pro 2 (Preview) (Non-reasoning) 10 - $0
353 Qwen3 VL 30B A3B Instruct 10 - $0.35
354 Claude 3.5 Sonnet (Oct '24) 9.9 30.2 $6
355 DeepSeek R1 Distill Llama 70B 9.9 - $0.787
356 Falcon-H1R-7B 9.8 - $0
357 DeepSeek R1 Distill Qwen 14B 9.8 - $0
358 Ling-flash-2.0 9.7 - $0.247
359 Qwen3 Omni 30B A3B (Reasoning) 9.6 - $0.43
360 GPT-4o (Aug '24) 9.6 - $4.375
361 GPT-4.1 nano 9.6 11.1 $0.175
362 Qwen2.5 Instruct 72B 9.6 - $0.48
363 Step3 VL 10B 9.5 - $0
364 Sonar 9.5 - $0
365 Llama 3.3 Instruct 70B 9.4 11.9 $0.612
366 Gemma 4 E2B (Reasoning) 9.3 - $0
367 Devstral Small (Jul '25) 9.3 - $0.15
368 Sonar Pro 9.3 - $0
369 Qwen3 30B A3B (Reasoning) 9.3 - $0.75
370 QwQ 32B-Preview 9.2 - $0
371 Llama 3.1 Nemotron Ultra 253B v1 (Reasoning) 9.1 - $0.9
372 Mistral Large 2 (Nov '24) 9.1 - $3
373 GLM-4.5V (Reasoning) 9.1 - $0.9
374 Qwen3 30B A3B 2507 Instruct 9.1 - $0.35
375 Ministral 3 8B 9 9.7 $0.15
376 Solar Pro 2 (Reasoning) 9 - $0
377 NVIDIA Nemotron Nano 12B v2 VL (Reasoning) 9 - $0.3
378 Hermes 4 - Llama-3.1 405B (Reasoning) 9 - $1.5
379 ERNIE 4.5 300B A47B 9 - $0.485
380 Gemma 4 E4B (Non-reasoning) 8.9 - $0
381 Granite 4.1 30B 8.9 10.4 $0
382 NVIDIA Nemotron Nano 9B V2 (Reasoning) 8.8 - $0.07
383 Hermes 4 - Llama-3.1 405B (Non-reasoning) 8.8 - $1.5
384 Gemini 2.0 Flash-Lite (Feb '25) 8.8 - $0
385 NVIDIA Nemotron 3 Nano 4B 8.7 8 $0
386 Llama Nemotron Super 49B v1.5 (Non-reasoning) 8.7 - $0.4
387 K2-V2 (low) 8.6 - $0
388 GPT-4o (May '24) 8.6 24.2 $7.5
389 Gemini 2.0 Flash-Lite (Preview) 8.6 - $0
390 Qwen3 32B (Non-reasoning) 8.6 - $1.225
391 Llama 3.1 Instruct 405B 8.5 - $3.688
392 Kimi Linear 48B A3B Instruct 8.5 - $0
393 Llama 3.1 Nemotron Nano 4B v1.1 (Reasoning) 8.5 - $0
394 Llama 3.3 Nemotron Super 49B v1 (Non-reasoning) 8.5 - $0
395 Qwen3 4B (Reasoning) 8.4 - $0
396 Qwen3 VL 8B Instruct 8.4 - $0.31
397 LFM2.5-8B-A1B 8.3 - $0
398 Claude 3.5 Sonnet (June '24) 8.3 26 $6
399 Llama 3.1 Tulu3 405B 8.3 - $0
400 Qwen3 8B (Reasoning) 8.3 9 $0.66
401 Ring-flash-2.0 8.2 - $0.247
402 GPT-4o (ChatGPT) 8.2 - $0
403 Olmo 3.1 32B Think 8.1 - $0
404 Pixtral Large 8.1 - $3
405 GPT-5 nano (minimal) 8 - $0.138
406 Gemini 1.5 Flash (Sep '24) 8 - $0
407 Grok 2 (Dec '24) 8 - $0
408 GPT-4 Turbo 7.9 21.5 $15
409 Qwen3 VL 4B (Reasoning) 7.9 - $0
410 Solar Pro 2 (Non-reasoning) 7.8 - $0
411 Command A 7.7 - $4.375
412 Nova Pro 7.7 - $1.4
413 Llama 3.1 Nemotron Instruct 70B 7.6 - $1.2
414 Qwen3.5 2B (Reasoning) 7.6 2.9 $0
415 Llama 3.1 Instruct 8B 7.6 5.4 $0.08
416 Grok Beta 7.5 - $0
417 Qwen2.5 Instruct 32B 7.5 - $0
418 NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning) 7.4 - $0.088
419 NVIDIA Nemotron Nano 9B V2 (Non-reasoning) 7.4 - $0.086
420 Gemma 3 27B Instruct 7.4 10.1 $0
421 Mistral Large 2 (Jul '24) 7.3 - $3
422 Qwen2.5 Coder Instruct 32B 7.1 - $0
423 Qwen3 4B 2507 Instruct 7.1 - $0
424 GPT-4 7 13.1 $37.5
425 GLM-4.5V (Non-reasoning) 7 - $0.9
426 Qwen3 14B (Non-reasoning) 7 - $0.612
427 Hermes 4 - Llama-3.1 70B (Non-reasoning) 6.9 - $0.198
428 GPT-4o mini 6.9 11.4 $0.262
429 Gemini 2.5 Flash-Lite (Non-reasoning) 6.9 - $0.175
430 Mistral Small 3 6.9 - $0.15
431 Nova Lite 6.9 - $0.105
432 Ministral 3 3B 6.8 4.8 $0.1
433 Llama 3.1 Instruct 70B 6.8 - $0.56
434 DeepSeek-V2.5 (Dec '24) 6.8 - $0
435 Qwen3 30B A3B (Non-reasoning) 6.8 - $0.35
436 Qwen3 4B (Non-reasoning) 6.8 - $0
437 Granite 4.1 8B 6.7 9.5 $0.063
438 Sarvam 30B (high) 6.6 - $0.047
439 Gemini 2.0 Flash Thinking Experimental (Dec '24) 6.6 - $0
440 DeepSeek-V2.5 6.6 - $0
441 Olmo 3.1 32B Instruct 6.5 - $0
442 Gemma 4 E2B (Non-reasoning) 6.4 - $0
443 Mistral Saba 6.4 - $0
444 DeepSeek R1 Distill Llama 8B 6.4 - $0
445 Olmo 3 32B Think 6.4 - $0
446 R1 1776 6.3 - $0
447 Gemini 1.5 Pro (May '24) 6.3 19.8 $0
448 Reka Flash (Sep '24) 6.3 - $0.35
449 Qwen2.5 Turbo 6.3 - $0.088
450 Llama 3.2 Instruct 90B (Vision) 6.2 - $1.38
451 Solar Mini 6.2 - $0.15
452 Grok-1 6 - $0
453 Phi-4 Mini Instruct 6 3.8 $0
454 EXAONE 4.0 32B (Non-reasoning) 6 - $0
455 Qwen2 Instruct 72B 6 - $0
456 Qwen3.5 2B (Non-reasoning) 5.6 2.4 $0
457 Qwen3.5 0.8B (Reasoning) 5.5 0 $0
458 Gemini 1.5 Flash-8B 5.5 - $0
459 Gemma 3 12B Instruct 5.5 5.8 $0
460 DeepHermes 3 - Mistral 24B Preview (Non-reasoning) 5.3 - $0
461 Jamba 1.7 Large 5.3 - $3.5
462 Granite 4.0 H Small 5.2 - $0.107
463 Qwen3 Omni 30B A3B Instruct 5.1 - $0.43
464 DeepSeek-Coder-V2 5.1 - $0
465 Hermes 3 - Llama-3.1 70B 5.1 - $0.7
466 Jamba 1.5 Large 5.1 - $3.5
467 Qwen3 8B (Non-reasoning) 5.1 - $0.31
468 OLMo 2 32B 5 - $0
469 Jamba 1.6 Large 5 - $3.5
470 Phi-4 4.9 - $0.219
471 LFM2 24B A2B 4.9 - $0.052
472 Gemini 1.5 Flash (May '24) 4.9 - $0
473 Nova Micro 4.7 - $0.061
474 Granite 4.1 3B 4.7 4.7 $0
475 Claude 3 Sonnet 4.7 - $6
476 Mistral Small (Sep '24) 4.7 - $0.3
477 NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) 4.6 - $0.3
478 Gemini 1.0 Ultra 4.6 17.6 $0
479 Gemma 3n E4B Instruct Preview (May '25) 4.6 - $0
480 Phi-3 Mini Instruct 3.8B 4.6 - $0
481 Phi-4 Multimodal Instruct 4.5 - $0
482 Qwen2.5 Coder Instruct 7B 4.5 - $0
483 Qwen3.5 0.8B (Non-reasoning) 4.4 1.2 $0
484 Mixtral 8x22B Instruct 4.4 - $0
485 Mistral Large (Feb '24) 4.4 - $6
486 Llama 2 Chat 7B 4.3 - $0.1
487 MiniCPM-V 4.6 1.3B 4.2 0.7 $0
488 Llama 3.2 Instruct 3B 4.2 - $0.15
489 Reka Flash 3 4.1 - $0.35
490 Jamba Reasoning 3B 4.1 - $0
491 Qwen1.5 Chat 110B 4.1 - $0
492 Qwen3 VL 4B Instruct 4.1 - $0
493 Olmo 3 7B Think 4 - $0
494 Claude 3 Haiku 3.9 - $0.5
495 Claude 2.1 3.9 14 $0
496 OLMo 2 7B 3.9 - $0
497 Molmo 7B-D 3.8 - $0
498 Ling-mini-2.0 3.8 - $0
499 DeepSeek R1 Distill Qwen 1.5B 3.7 - $0
500 GPT-3.5 Turbo 3.6 10.7 $0.75
501 Claude 2.0 3.6 12.9 $0
502 Mistral Small (Feb '24) 3.6 - $1.5
503 Mistral Medium 3.6 - $4.088
504 DeepSeek-V2-Chat 3.6 - $0
505 Llama 3 Instruct 70B 3.5 - $1.175
506 LFM 40B 3.4 - $0
507 Arctic Instruct 3.4 - $0
508 Qwen Chat 72B 3.4 - $0
509 Llama 3.2 Instruct 11B (Vision) 3.3 - $0.345
510 PALM-2 3.2 4.6 $0
511 Gemini 1.0 Pro 3.1 - $0
512 DeepSeek Coder V2 Lite Instruct 3.1 - $0
513 Llama 2 Chat 70B 3 - $0
514 Llama 2 Chat 13B 3 - $0
515 DeepSeek LLM 67B Chat (V1) 3 - $0
516 OpenChat 3.5 (1210) 3 - $0
517 DBRX Instruct 3 - $0
518 Sarvam M (Reasoning) 3 - $0
519 Command-R+ (Apr '24) 3 - $6
520 Exaone 4.0 1.2B (Reasoning) 2.9 - $0
521 Olmo 3 7B Instruct 2.8 - $0.125
522 Exaone 4.0 1.2B (Non-reasoning) 2.8 - $0
523 LFM2 2.6B 2.7 - $0
524 LFM2.5-1.2B-Thinking 2.7 - $0
525 LFM2.5-1.2B-Instruct 2.7 - $0
526 Granite 4.0 H 1B 2.7 - $0
527 Jamba 1.7 Mini 2.7 - $0
528 Jamba 1.5 Mini 2.7 - $0.25
529 Jamba 1.6 Mini 2.6 - $0.25
530 Qwen3 1.7B (Reasoning) 2.6 - $0
531 Gemma 3 270M 2.4 - $0
532 Granite 4.0 Micro 2.4 - $0
533 Apertus 70B Instruct 2.4 - $1.345
534 Mixtral 8x7B Instruct 2.4 - $0.512
535 DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning) 2.3 - $0
536 Llama 65B 2.1 - $0
537 Qwen Chat 14B 2.1 - $0
538 Granite 4.0 1B 2.1 - $0
539 Claude Instant 2.1 7.8 $0
540 Mistral 7B Instruct 2.1 - $0.25
541 Command-R (Mar '24) 2.1 - $0.75
542 Molmo2-8B 2 - $0
543 LFM2 8B A1B 1.8 - $0
544 Granite 3.3 8B (Non-reasoning) 1.8 - $0.085
545 Qwen3 1.7B (Non-reasoning) 1.5 - $0
546 Qwen3 0.6B (Reasoning) 1.3 - $0
547 Llama 3 Instruct 8B 1.2 - $0.07
548 Gemma 3n E4B Instruct 1.2 3.2 $0.025
549 Llama 3.2 Instruct 1B 1.1 - $0.05
550 Gemma 3 4B Instruct 1.1 2.7 $0
551 LFM2 1.2B 1.1 - $0
552 LFM2.5-VL-1.6B 1 - $0
553 Granite 4.0 H 350M 1 - $0
554 Granite 4.0 350M 1 - $0
555 Apertus 8B Instruct 1 - $0.125
556 Tiny Aya Global 1 - $0
557 Gemma 3 1B Instruct 1 - $0
558 Gemma 3n E2B Instruct 1 - $0
559 Qwen3 0.6B (Non-reasoning) 1 - $0
560 GPT-5.5 Pro (xhigh) - - $0
561 Gemini 3 Deep Think - - $0
562 Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort) - - $4
563 Claude Sonnet 5 (Adaptive Reasoning, Medium Effort) - - $4
564 Claude Sonnet 5 (Adaptive Reasoning, High Effort) - - $4
565 Claude Sonnet 5 (Adaptive Reasoning, Low Effort) - - $4
566 EXAONE 4.5 33B (Non-reasoning) - - $0
567 Cogito v2.1 (Reasoning) - - $1.25
568 Mi:dm K 2.5 Pro Preview - - $0
569 GPT-5.4 Pro (xhigh) - - $67.5
570 GPT-4o Realtime (Dec '24) - - $0
571 GPT-4o mini Realtime (Dec '24) - - $0
572 GPT-3.5 Turbo (0613) - - $0

榜单解读建议

参考 AI 大模型排行榜 时,应综合考虑“综合指数”与“成本价格”。如果您是开发者,编程能力 (Coding) 是更核心的指标。

值品工具箱同步的 AI 大模型排行榜 数据每 24 小时更新,确保您获取到最新的模型性能对比。

指标说明

  • 综合指数:评估通用理解与逻辑。
  • 价格 $/1M:混合 3:1 输入输出比的平均成本。
  • 编程能力:衡量代码生成的准确性。

AI 大模型排行榜 常见问题 (FAQ)

Q1: AI 大模型排行榜 的数据多久更新?

AI 大模型排行榜 数据每 24 小时自动抓取一次,确保最新模型加入列表。

Q2: 这个 AI 大模型排行榜 包含国产模型吗?

是的,只要国产模型通过了 Artificial Analysis 的全球测评,就会出现在 AI 大模型排行榜 中。

Q3: 综合指数在 AI 大模型排行榜 中代表什么?

它代表模型的全能表现。AI 大模型排行榜 通过加权算法给出这个综合评分。

Q4: 如何在 AI 大模型排行榜 中查找性价比最高的游戏?

在 AI 大模型排行榜 页面中,您可以点击“价格”标题进行排序,寻找低价高分的模型。

Q5: AI 大模型排行榜 的编程能力测试准吗?

AI 大模型排行榜 参考了 LiveCodeBench 等权威基准测试,具有极高的参考价值。

Q6: 为什么有的新模型没进入 AI 大模型排行榜?

模型进入 AI 大模型排行榜 需要经过一系列测试,通常在新模型发布后数日内会完成更新。

Q7: AI 大模型排行榜 中的价格计算标准是什么?

价格是基于百万 Token 的调用成本,由 AI 大模型排行榜 统一混合计算得出。

Q8: 手机上能查看 AI 大模型排行榜 吗?

当然可以。AI 大模型排行榜 进行了移动端响应式深度优化。

Q9: AI 大模型排行榜 这个工具免费吗?

是的,由值品工具箱免费提供 AI 大模型排行榜 信息查询服务。

Q10: 我该怎么利用 AI 大模型排行榜 做选型?

如果您需要智能客服,参考 AI 大模型排行榜 的综合指数;如果做翻译,参考编程外的语言指标。

发表评论

请友善文明留言