AI 大模型排行榜 (Artificial Analysis LLM Ranking)

信息查询
3.3k 次浏览
100% 有帮助 · 1 人反馈

值品工具箱提供的 AI 大模型排行榜聚合了来自 Artificial Analysis 的权威数据,实时追踪并排名超过 100 个主流大语言模型。

AI 大模型排行榜数据中心

重置
排名 模型名称 综合指数 ▼ 编程 价格 ($/1M)
1 Claude Opus 5 (Adaptive Reasoning, Max Effort) 63.1 78 $10
2 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) 62.5 77 $10
3 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) 62.1 76.5 $20
4 Claude Opus 5 (Adaptive Reasoning, High Effort) 61.5 76.5 $10
5 GPT-5.6 Sol (max) 60.9 77.4 $11.25
6 Grok 4.6 (high) 60.9 76.8 $3
7 Kimi K3 (max) 59.7 76.2 $6
8 GPT-5.6 Sol (xhigh) 59 78.3 $11.25
9 Claude Opus 5 (Adaptive Reasoning, Medium Effort) 58.6 74.3 $10
10 Qwen3.8 Max 58.1 71.8 $3
11 Qwen3.8 2.4T A95B 57.7 71.9 $3
12 GPT-5.6 Sol (high) 57.3 77.2 $11.25
13 Claude Opus 4.8 (Adaptive Reasoning, Max Effort) 57.3 74.3 $10
14 Muse Spark 1.2 (xhigh) 56.8 72.2 $2
15 GPT-5.6 Terra (max) 56.6 76.7 $4.5
16 GPT-5.5 (xhigh) 56.3 74.9 $11.25
17 Gemini 3.7 Flash (high) 56 76.1 $1.5
18 Grok 4.5 (high) 55.8 72.4 $3
19 GPT-5.6 Sol (medium) 55.6 76.3 $11.25
20 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) 55.3 71.5 $4
21 Claude Opus 4.7 (Adaptive Reasoning, Max Effort) 55 73.6 $10
22 GPT-5.5 (high) 54.7 71.6 $11.25
23 Gemini 3.7 Flash (medium) 53.4 71.5 $1.5
24 DeepSeek V4 Pro 0813 (Reasoning, Max Effort) 53.2 68.8 $1.98
25 Muse Spark 1.1 (xhigh) 53.2 71.3 $2
26 GPT-5.4 (xhigh) 53.1 71.1 $5.625
27 GPT-5.6 Terra (xhigh) 52.8 70.6 $4.5
28 GLM-5.2 (max) 52.6 68.8 $2.15
29 Claude Opus 5 (Adaptive Reasoning, Low Effort) 52.5 66.9 $10
30 GPT-5.6 Luna (max) 52.3 71.4 $0.45
31 Qwen3.8 27B 52 68.1 $0
32 Gemini 3.5 Flash (high) 52 70.1 $3.375
33 DeepSeek V4 Flash 0731 (Reasoning, Max Effort) 51.8 69.1 $0.66
34 Gemini 3.6 Flash (high) 51.6 69.2 $3
35 GPT-5.5 (medium) 51.4 71.5 $11.25
36 Gemini 3.7 Flash (low) 50.9 71 $1.5
37 GPT-5.6 Sol (low) 50.7 69.7 $11.25
38 GPT-5.6 Terra (high) 50.1 67.1 $4.5
39 GPT-5.6 Luna (xhigh) 50.1 68.6 $0.45
40 Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) 48.4 63 $6
41 Kimi K3 (low) 48.3 72 $6
42 Gemini 3.1 Pro Preview 47.7 68.8 $4.5
43 Motif 3 47.4 63.5 $0
44 GPT-5.6 Luna (high) 47 63.3 $0.45
45 GPT-5.6 Terra (medium) 46.8 64.7 $4.5
46 Gemini 3.5 Flash (medium) 46.7 - $3.375
47 Qwen3.7 Max 46.7 66 $3.75
48 GPT-5.3 Codex (xhigh) 45.5 - $4.813
49 MiniMax-M3 45.4 58.6 $0.525
50 DeepSeek V4 Pro (Reasoning, Max Effort) 45.3 59.4 $0.544
51 Motif 3 (Beta) 45.3 62 $0
52 Kimi K2.6 45.1 61.8 $1.712
53 Claude Opus 4.6 (Adaptive Reasoning, Max Effort) 44.9 - $10
54 GPT-5.5 (low) 44.5 60.9 $11.25
55 Muse Spark 44.3 58.6 $0
56 Claude Opus 4.7 (Non-reasoning, High Effort) 43.9 - $10
57 DeepSeek V4 Pro (Reasoning, High Effort) 43.7 58.7 $0.544
58 GPT-5.2 (xhigh) 43.3 - $4.813
59 Kimi K2.7 Code 43 60.8 $1.712
60 MiMo-V2.5-Pro 42.9 60.2 $0.544
61 Claude Sonnet 5 (Non-reasoning, High Effort) 42.6 66.4 $4
62 Inkling (xhigh) 42.3 52.1 $1.762
63 Hy3 42.2 58.8 $0.241
64 DeepSeek V4 Flash (Reasoning, Max Effort) 42.1 56.2 $0.168
65 GPT-5.6 Sol (Non-reasoning) 41.9 65.1 $11.25
66 Claude Opus 4.5 (Reasoning) 41.9 - $10
67 Nex-N2-Pro 41.7 59.1 $1
68 Solar Pro 4 41.6 52.7 $0.525
69 MiMo-V2-Pro 41.4 - $0
70 GPT-5.6 Terra (low) 41.3 58.1 $4.5
71 Inkling Small 41.2 52.9 $0.525
72 GPT-5.2 Codex (xhigh) 41.2 - $4.813
73 Qwen3.6 Max Preview 41.1 - $2.925
74 GLM-5.1 (Reasoning) 41 55.8 $2.135
75 GPT-5.4 mini (xhigh) 40.9 56.1 $1.688
76 Grok Build 0.1 0616 40.7 51.5 $1.25
77 Gemini 3 Pro Preview (high) 40.6 - $4.5
78 GLM-5 (Reasoning) 40.6 - $1.55
79 Qwen3.6 Plus 40.5 54.5 $1.125
80 GPT-5.4 (low) 40.2 - $5.625
81 JT-4.1 Flash 236B A21B 39.9 52.4 $0
82 Agnes 2.5 Pro Alpha 39.7 58.8 $0.563
83 GPT-5.4 nano (xhigh) 39.7 56.1 $0.463
84 Qwen3.7 Plus 39.4 55.9 $0.7
85 GLM-5-Turbo 39.1 - $0
86 DeepSeek V4 Flash (Reasoning, High Effort) 39 52 $0.175
87 GPT-5.6 Luna (medium) 38.9 50.7 $0.45
88 GPT-5.2 (medium) 38.9 - $4.813
89 MiniMax-M2.7 38.9 52.6 $0.525
90 Claude Opus 4.6 (Non-reasoning, High Effort) 38.8 - $10
91 Gemini 3 Flash Preview (Reasoning) 38.7 - $1.125
92 Nemotron 3 Ultra 550B A55B (Reasoning) 38.3 49.3 $1.137
93 MiMo-V2.5 38 56.8 $0.175
94 Grok 4.20 0309 v2 (Reasoning) 38 - $1.563
95 Grok 4.3 (high) 37.9 42.2 $1.563
96 Ling 3.0 Flash 37.8 50.6 $0.111
97 Qwen3.6 27B (Reasoning) 37.7 53.7 $1.35
98 GPT-5.1 (high) 37.5 49.4 $3.438
99 Gemini 3.5 Flash-Lite 37.4 49.3 $0.85
100 Solar Open2 250B 37.4 44.7 $0
101 Claude 4.5 Sonnet (Reasoning) 37.4 52.1 $6
102 Grok 4.20 0309 (Reasoning) 37.4 - $3
103 MiMo-V2-Omni-0327 37.3 - $0
104 GPT-5 Codex (high) 37 - $3.438
105 Grok 4.3 (medium) 36.9 - $1.563
106 Claude Sonnet 4.6 (Non-reasoning, High Effort) 36.8 - $6
107 Grok 4.3 (low) 36.3 - $1.563
108 GLM-5.1 (Non-reasoning) 36.3 - $2.135
109 Kimi K2.5 (Reasoning) 36 46.8 $1.2
110 MiMo-V2-Omni 35.9 - $0
111 Gemini 3.5 Flash (minimal) 35.8 - $3.375
112 GPT-5.5 (Non-reasoning) 35.8 56.5 $11.25
113 GPT-5.1 Codex (high) 35.6 - $3.438
114 Claude Opus 4.5 (Non-reasoning) 35.6 - $10
115 Kimi K2.6 (Non-reasoning) 35.4 - $1.712
116 GPT-5 (high) 35.3 37.8 $3.438
117 GLM 5V Turbo (Reasoning) 35.3 - $0
118 Muse Glimmer (high) 35.1 49 $0.581
119 Claude Sonnet 4.6 (Non-reasoning, Low Effort) 35.1 - $6
120 A.X-K2 35 38.8 $0
121 GLM-5.2 (Non-reasoning) 34.8 46.5 $2.15
122 GPT-5.6 Terra (Non-reasoning) 34.6 52.3 $4.5
123 GPT-5 (medium) 34.6 - $3.438
124 Qwen3.5 27B (Reasoning) 34.6 - $0.825
125 Claude 4.1 Opus (Reasoning) 34.5 - $30
126 MiniMax-M2.5 34.5 - $0.525
127 GLM-4.7 (Reasoning) 34.5 45.3 $1
128 Hy3-preview (Reasoning) 34.4 - $0.1
129 Qwen3.5 397B A17B (Reasoning) 34.3 48.2 $1.35
130 GPT-5.5 Instant (May 2026) 34.3 - $11.25
131 Grok 4 34.1 - $6
132 MiMo-V2-Flash (Feb 2026) 34 - $0
133 LongCat 2.0 34 45.3 $1.3
134 GPT-5.6 Luna (low) 33.9 44.2 $0.45
135 Gemini 3 Pro Preview (low) 33.9 - $4.5
136 KAT Coder Pro V2 33.7 59.5 $0.525
137 Kimi K2 Thinking 33.5 - $1.075
138 o3-pro 33.3 - $35
139 GLM-5 (Non-reasoning) 33.2 - $1.55
140 Qwen3.5 122B A10B (Reasoning) 32.8 45.7 $1.1
141 DeepSeek V3.2 (Reasoning) 32.8 44.2 $0.315
142 Qwen3.5 397B A17B (Non-reasoning) 32.7 - $1.35
143 Qwen3 Max Thinking 32.5 - $0
144 Qwen3.6 35B A3B (Reasoning) 32.1 41.9 $0.844
145 MiniMax-M2.1 32.1 - $0.525
146 DeepSeek V4 Pro (Non-reasoning) 31.9 - $0.544
147 GPT-5 (low) 31.9 - $3.438
148 MiMo-V2-Flash (Reasoning) 31.9 - $0.15
149 Ring-2.6-1T 31.7 42.8 $0.85
150 Claude 4 Opus (Reasoning) 31.7 - $30
151 G9v3-39A5B 31.6 31.7 $0
152 GPT-5 mini (medium) 31.6 - $0.688
153 Qwen3.5 Omni Plus 31.3 - $1.5
154 Qwen3.6 27B (Non-reasoning) 31.3 46.6 $1.35
155 GPT-5.1 Codex mini (high) 31.3 - $0.688
156 Grok 4.1 Fast (Reasoning) 31.3 - $0
157 o3 31.1 - $3.5
158 DeepSeek V3.1 Terminus (Reasoning) 31.1 43.5 $1.914
159 K-EXAONE 2.0 0803 31 40.6 $0
160 Step 3.7 Flash 30.9 39.6 $0.438
161 GPT-5.4 nano (medium) 30.8 - $0.463
162 GPT-5.4 mini (medium) 30.5 - $1.688
163 Mistral Medium 3.5 30.4 46.9 $3
164 Kimi K2.5 (Non-reasoning) 30.1 - $1.2
165 Qwen3.5 27B (Non-reasoning) 30 - $0.825
166 Claude 4.5 Haiku (Reasoning) 29.9 43.9 $2
167 Claude 4.5 Sonnet (Non-reasoning) 29.9 - $6
168 Qwen3.5 35B A3B (Reasoning) 29.9 - $0.688
169 Claude 4 Sonnet (Reasoning) 29.8 37.6 $6
170 Gemma 4 31B (Reasoning) 29.7 43.4 $0
171 DeepSeek V4 Flash (Non-reasoning) 29.3 - $0.175
172 GLM-4.6 (Reasoning) 29.3 45.8 $0.963
173 GPT-5.5 Instant (June 2026) 29.2 39.4 $11.25
174 JT-35B-Flash 29 - $0
175 MiniMax-M2 28.9 - $0.525
176 KAT-Coder-Pro V1 28.9 - $0
177 Claude 4.1 Opus (Non-reasoning) 28.8 - $30
178 MiMo-V2.5-Pro (Non-reasoning) 28.4 - $0.544
179 GPT-5.4 (Non-reasoning) 28.3 - $5.625
180 Qwen3.5 122B A10B (Non-reasoning) 28.2 43.3 $1.1
181 Gemini 3 Flash Preview (Non-reasoning) 27.9 - $1.125
182 Grok 4 Fast (Reasoning) 27.9 - $0.275
183 Claude 3.7 Sonnet (Reasoning) 27.6 36.4 $0
184 GLM-4.7 (Non-reasoning) 27.1 - $1
185 GPT-5.6 Luna (Non-reasoning) 26.8 39.3 $0.45
186 Hy3-preview (Non-reasoning) 26.6 - $0.1
187 Ling-2.6-1T 26.6 - $0.85
188 Doubao Seed Code 26.5 - $0
189 GPT-5.2 (Non-reasoning) 26.5 - $4.813
190 Step 3.5 Flash 2603 26.5 - $0.15
191 Gemma 4 26B A4B (Reasoning) 26.1 39.3 $0.198
192 o4-mini (high) 26.1 - $1.925
193 Claude 4 Opus (Non-reasoning) 26 - $30
194 Claude 4 Sonnet (Non-reasoning) 26 - $6
195 Step 3.5 Flash 26 - $0.15
196 Gemini 2.5 Pro 25.9 33.3 $3.438
197 DeepSeek V3.2 Exp (Reasoning) 25.9 - $0.315
198 GPT-5 mini (high) 25.8 15.6 $0.688
199 Nemotron 3 Super 120B A12B (Reasoning) 25.7 37.7 $0.35
200 Gemini 3.1 Flash-Lite 25.6 34.7 $0.563
201 Qwen3 Max Thinking (Preview) 25.5 - $2.4
202 MiMo-V2-Flash (Non-reasoning) 25.1 49.8 $0
203 DeepSeek V3.2 (Non-reasoning) 25.1 - $0.315
204 Grok 4.3 (Non-reasoning) 25 35.2 $1.563
205 Qwen3.6 35B A3B (Non-reasoning) 24.6 28.1 $0.844
206 Ling 3.0 Tiny 24.5 26.5 $0
207 Qwen3 Max 24.5 - $2.4
208 Qwen3.5 35B A3B (Non-reasoning) 24.3 37 $0.688
209 Gemini 2.5 Flash Preview (Sep '25) (Reasoning) 24.2 - $0
210 gpt-oss-120b (high) 24.1 30.4 $0.262
211 Claude 4.5 Haiku (Non-reasoning) 24.1 - $2
212 Kimi K2 0905 24 - $1.075
213 o1 23.9 39.7 $26.25
214 Claude 3.7 Sonnet (Non-reasoning) 23.9 - $6
215 Nemotron 3.5 Lightning 23.6 26.8 $0.088
216 Gemini 2.5 Pro Preview (Mar' 25) 23.4 46.7 $0
217 GLM-4.6 (Non-reasoning) 23.4 - $0.981
218 GLM-4.7-Flash (Reasoning) 23.3 - $0.153
219 Grok 4.20 0309 (Non-reasoning) 22.9 - $3
220 Grok 3 mini Reasoning (high) 22.9 - $0.35
221 Command A+ 22.8 27.8 $0
222 Gemini 2.5 Pro Preview (May' 25) 22.7 - $3.438
223 DeepSeek V3.2 Speciale 22.6 - $0
224 K-EXAONE (Reasoning) 22.5 32.1 $0
225 Gemma 4 31B (Non-reasoning) 22.3 33.2 $0.205
226 ERNIE 5.0 Thinking Preview 22.3 - $0
227 Gemma 4 12B (Reasoning) 22.2 31 $0.15
228 Grok 4.20 0309 v2 (Non-reasoning) 22.2 - $1.563
229 Nova 2.0 Pro Preview (medium) 22.1 34 $3.438
230 Grok Code Fast 1 22 - $0
231 Mercury 2 21.9 31.1 $0.375
232 Qwen3.5 9B (Reasoning) 21.8 28.7 $0.151
233 DeepSeek V3.1 Terminus (Non-reasoning) 21.7 - $0.453
234 DeepSeek V3.2 Exp (Non-reasoning) 21.7 - $0.315
235 Apriel-v1.5-15B-Thinker 21.6 - $0
236 DeepSeek V3.1 (Non-reasoning) 21.4 - $0.84
237 Nova 2.0 Omni (medium) 21.3 - $0.85
238 Qwen3 Coder Next 21.3 36.2 $0.563
239 DeepSeek V3.1 (Reasoning) 21 - $0.865
240 Qwen3 VL 235B A22B (Reasoning) 20.9 - $2.625
241 Nova 2.0 Lite (high) 20.8 23 $0.85
242 Apriel-v1.6-15B-Thinker 20.8 - $0
243 GPT-5.1 (Non-reasoning) 20.7 - $3.438
244 Qwen3.5 9B (Non-reasoning) 20.6 23.5 $0.19
245 EXAONE 4.5 33B 20.5 23.6 $0
246 Gemma 4 26B A4B (Non-reasoning) 20.4 - $0.198
247 Qwen3.5 4B (Reasoning) 20.4 22.6 $0.06
248 DeepSeek R1 0528 (May '25) 20.4 - $2.063
249 Gemini 2.5 Flash (Reasoning) 20.3 - $0.85
250 North Mini Code 20.2 36.5 $0
251 GPT-5 nano (high) 20.1 - $0.138
252 Qwen3 235B A22B 2507 (Reasoning) 19.9 22.1 $2.625
253 Nova 2.0 Pro Preview (low) 19.8 25.9 $3.438
254 Mistral Small 4 (Reasoning) 19.7 26.6 $0.262
255 Kimi K2 19.7 - $1.002
256 GLM-4.5 (Reasoning) 19.7 - $0
257 GPT-4.1 19.6 - $3.5
258 Qwen3 Max (Preview) 19.4 - $2.4
259 Devstral 2 19.2 31.3 $0
260 Nova 2.0 Lite (medium) 19.2 - $0.85
261 Qwen3.5 Omni Flash 19.2 - $0.275
262 o3-mini 19.2 - $1.925
263 GPT-5 nano (medium) 19.2 - $0.138
264 o1-pro 19.1 - $262.5
265 Gemini 2.5 Flash Preview (Sep '25) (Non-reasoning) 19.1 - $0
266 JT-MINI 18.8 - $0
267 Trinity Large Thinking 18.7 25.8 $0.395
268 DeepSeek R1 (Jan '25) 18.6 24.6 $2.431
269 Grok 3 18.6 - $8
270 Seed-OSS-36B-Instruct 18.5 - $0.3
271 Qwen3 235B A22B 2507 Instruct 18.4 - $1.225
272 HyperNova 60B 2605 18.3 23.2 $0
273 Qwen3 Coder 480B A35B Instruct 18.2 - $3
274 Qwen3 VL 32B (Reasoning) 18.1 - $2.625
275 Magistral Medium 1.2 18 21.3 $2.75
276 Nova 2.0 Lite (low) 18 - $0.85
277 Sonar Reasoning Pro 18 - $0
278 MiniMax M1 80k 17.9 - $0.963
279 Nemotron Cascade 2 30B A3B 17.8 25.3 $0
280 GPT-5.4 nano (Non-Reasoning) 17.8 - $0.463
281 Devstral Small 2 17.7 29.3 $0
282 Gemini 2.5 Flash Preview (Reasoning) 17.7 - $0
283 K2 Think V2 17.4 21 $0
284 LongCat Flash Lite 17.4 - $0
285 GPT-5 (minimal) 17.3 - $3.438
286 HyperCLOVA X SEED Think (32B) 17.2 - $0
287 o1-preview 17.2 34 $28.875
288 Grok 4.1 Fast (Non-reasoning) 17 - $0
289 K-EXAONE (Non-reasoning) 16.9 - $0
290 Qwen3 Next 80B A3B (Reasoning) 16.9 17.4 $1.875
291 GLM-4.6V (Reasoning) 16.9 - $0.45
292 GPT-5.4 mini (Non-Reasoning) 16.8 - $1.688
293 Nova 2.0 Omni (low) 16.7 - $0.85
294 GLM-4.5-Air 16.7 - $0.372
295 Mi:dm K 2.5 Pro 16.6 - $0
296 Grok 4 Fast (Non-reasoning) 16.6 - $0.275
297 Ring-1T 16.3 - $0
298 G9v3-3B 16.2 9.9 $0
299 Qwen3.5 4B (Non-reasoning) 16.1 20.3 $0.06
300 Mistral Large 3 15.9 20.1 $0.75
301 INTELLECT-3 15.7 - $0
302 o3-mini (high) 15.7 16.3 $1.925
303 GLM-4.7-Flash (Non-reasoning) 15.6 - $0.153
304 GPT-5 (ChatGPT) 15.4 - $0
305 gpt-oss-20b (high) 15.2 20.7 $0.092
306 Solar Open 100B (Reasoning) 15.2 - $0
307 Gemini 2.5 Flash-Lite Preview (Sep '25) (Reasoning) 15.2 - $0.175
308 DeepSeek V3 0324 15.2 21.2 $0.483
309 Grok 3 Reasoning Beta 15.2 - $0
310 Nemotron 3 Nano Omni 30B A3B Reasoning 15 13.8 $0.131
311 gpt-oss-120b (low) 14.9 21.2 $0.261
312 Mistral Small 3.1 14.9 26.3 $0.15
313 GPT-4.1 mini 14.8 20.2 $0.7
314 Mistral Medium 3.1 14.7 20.5 $0.8
315 Qwen3 30B A3B 2507 (Reasoning) 14.6 12.1 $0.75
316 Llama 4 Maverick 14.5 16.3 $0.415
317 Solar Pro 3 14.5 16.2 $0.262
318 NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) 14.5 14.4 $0.088
319 MiniMax M1 40k 14.5 - $0
320 gpt-oss-20b (low) 14.4 - $0.103
321 Nova 2.0 Pro Preview (Non-reasoning) 14.4 20.9 $3.438
322 Qwen3 VL 235B A22B Instruct 14.4 - $1.225
323 GPT-5 mini (minimal) 14.3 - $0.688
324 K2-V2 (high) 14.2 - $0
325 Gemini 2.5 Flash (Non-reasoning) 14.2 - $0.85
326 DeepSeek V3 (Dec '24) 14.2 23 $0.493
327 Ling 2.6 Flash 14.2 25.3 $0.15
328 o1-mini 14 - $0
329 Qwen3 Next 80B A3B Instruct 13.8 - $0.875
330 GPT-4.5 (Preview) 13.6 - $0
331 Tri-21B-think Preview 13.6 - $0
332 Qwen3 Coder 30B A3B Instruct 13.6 - $0.9
333 DiffusionGemma 26B A4B 13.5 19.7 $0
334 Qwen3 235B A22B (Reasoning) 13.5 - $2.625
335 QwQ 32B 13.4 - $0.745
336 Qwen3 VL 30B A3B (Reasoning) 13.4 - $0.75
337 Gemini 2.0 Flash Thinking Experimental (Jan '25) 13.3 24.1 $0
338 Gemma 4 12B (Non-reasoning) 13.2 - $0.15
339 Gemini 2.5 Flash-Lite Preview (Sep '25) (Non-reasoning) 13.1 - $0.175
340 Motif-2-12.7B-Reasoning 12.8 - $0
341 Nova Premier 12.7 - $5
342 Ling-1T 12.7 - $0
343 Magistral Medium 1 12.5 - $0
344 Mistral Medium 3 12.5 - $0.8
345 Solar Pro 2 (Preview) (Reasoning) 12.5 - $0
346 Llama Nemotron Super 49B v1.5 (Reasoning) 12.4 - $0.4
347 K2-V2 (medium) 12.4 - $0
348 Celeris-1 12.4 14.4 $0.325
349 Devstral Medium 12.4 - $0
350 Mistral Small 4 (Non-reasoning) 12.3 - $0.262
351 Tri-21B-Think 12.3 - $0
352 GPT-4o (March 2025, chatgpt-4o-latest) 12.3 - $0
353 Gemma 4 E4B (Reasoning) 12.2 9.4 $0.04
354 Gemini 2.0 Flash (Feb '25) 12.2 - $0
355 Claude 3.5 Haiku 12.2 15.9 $0
356 Llama 3.3 Nemotron Super 49B v1 (Reasoning) 12.2 - $0
357 Sarvam 105B (high) 11.9 - $0.074
358 MiniCPM5-1B (Reasoning) 11.9 - $0
359 Qwen3 4B 2507 (Reasoning) 11.9 - $0
360 Nova 2.0 Lite (Non-reasoning) 11.8 - $0.85
361 Gemini 2.0 Pro Experimental (Feb '25) 11.8 25.5 $0
362 Claude 3 Opus 11.8 19.5 $30
363 Devstral Small (May '25) 11.8 - $0
364 MiniCPM5-1B (Non-reasoning) 11.7 - $0
365 Gemini 2.5 Flash Preview (Non-reasoning) 11.6 - $0
366 Sonar Reasoning 11.6 - $0
367 Magistral Small 1.2 11.5 14.7 $0.75
368 Gemini 2.5 Flash-Lite (Reasoning) 11.4 - $0.175
369 Qwen3 32B (Reasoning) 11.4 15.3 $2.625
370 Ministral 3 14B 11.2 14.4 $0.2
371 GPT-4o (Nov '24) 11.1 - $4.375
372 Nanbeige4.1-3B 11 9.6 $0
373 DeepSeek R1 Distill Qwen 32B 11 - $0
374 Qwen3 VL 32B Instruct 11 - $1.225
375 GLM-4.6V (Non-reasoning) 10.9 - $0.45
376 Qwen3 235B A22B (Non-reasoning) 10.8 - $1.225
377 Mistral Small 3.2 10.7 12.5 $0.15
378 Gemini 2.0 Flash (experimental) 10.6 - $0
379 Magistral Small 1 10.6 - $0
380 EXAONE 4.0 32B (Reasoning) 10.5 - $0
381 Qwen3 VL 8B (Reasoning) 10.5 - $0.66
382 Nova 2.0 Omni (Non-reasoning) 10.4 - $0.85
383 Qwen3 14B (Reasoning) 10.4 13.8 $1.313
384 Llama 4 Scout 10.3 8.2 $0.3
385 DeepSeek R1 0528 Qwen3 8B 10.3 - $0
386 Qwen2.5 Max 10.1 - $0
387 Hermes 4 - Llama-3.1 70B (Reasoning) 9.9 - $0.198
388 Gemini 1.5 Pro (Sep '24) 9.9 23.6 $0
389 Solar Pro 2 (Preview) (Non-reasoning) 9.9 - $0
390 Qwen3 VL 30B A3B Instruct 9.9 - $0.35
391 Claude 3.5 Sonnet (Oct '24) 9.8 30.2 $6
392 DeepSeek R1 Distill Llama 70B 9.8 - $0.8
393 Falcon-H1R-7B 9.7 - $0
394 DeepSeek R1 Distill Qwen 14B 9.7 - $0
395 GPT-4.1 nano 9.6 11.1 $0.175
396 Ling-flash-2.0 9.6 - $0.247
397 Gemma 4 E2B (Reasoning) 9.5 7.2 $0
398 Qwen3 Omni 30B A3B (Reasoning) 9.5 - $0.43
399 GPT-4o (Aug '24) 9.4 - $4.375
400 Sonar 9.4 - $0
401 Qwen2.5 Instruct 72B 9.4 - $0.48
402 Llama 3.3 Instruct 70B 9.3 11.9 $0.671
403 Step3 VL 10B 9.3 - $0
404 Qwen3 30B A3B (Reasoning) 9.2 - $0.75
405 Devstral Small (Jul '25) 9.1 - $0
406 Sonar Pro 9.1 - $0
407 QwQ 32B-Preview 9.1 - $0
408 Ministral 3 8B 9 9.7 $0.15
409 Mistral Large 2 (Nov '24) 9 - $0
410 GLM-4.5V (Reasoning) 9 - $0.9
411 Llama 3.1 Nemotron Ultra 253B v1 (Reasoning) 8.9 - $0.9
412 ERNIE 4.5 300B A47B 8.9 - $0.485
413 Qwen3 30B A3B 2507 Instruct 8.9 - $0.35
414 NVIDIA Nemotron Nano 12B v2 VL (Reasoning) 8.8 - $0.3
415 Hermes 4 - Llama-3.1 405B (Reasoning) 8.8 - $1.5
416 Solar Pro 2 (Reasoning) 8.8 - $0
417 Gemma 4 E4B (Non-reasoning) 8.7 - $0.04
418 NVIDIA Nemotron Nano 9B V2 (Reasoning) 8.7 - $0.07
419 Granite 4.1 30B 8.7 10.4 $0
420 NVIDIA Nemotron 3 Nano 4B 8.6 8 $0
421 Hermes 4 - Llama-3.1 405B (Non-reasoning) 8.6 - $1.5
422 Gemini 2.0 Flash-Lite (Feb '25) 8.6 - $0
423 Llama Nemotron Super 49B v1.5 (Non-reasoning) 8.5 - $0.4
424 Qwen3 32B (Non-reasoning) 8.5 - $1.225
425 Kimi Linear 48B A3B Instruct 8.4 - $0
426 K2-V2 (low) 8.4 - $0
427 GPT-4o (May '24) 8.4 24.2 $7.5
428 Gemini 2.0 Flash-Lite (Preview) 8.4 - $0
429 Llama 3.1 Nemotron Nano 4B v1.1 (Reasoning) 8.4 - $0
430 Llama 3.1 Instruct 405B 8.3 - $0
431 Llama 3.3 Nemotron Super 49B v1 (Non-reasoning) 8.3 - $0
432 Qwen3 8B (Reasoning) 8.3 9 $0.66
433 Qwen3 VL 8B Instruct 8.2 - $0.31
434 Qwen3 4B (Reasoning) 8.2 - $0
435 LFM2.5-8B-A1B 8.1 - $0
436 GPT-4o (ChatGPT) 8.1 - $0
437 Claude 3.5 Sonnet (June '24) 8.1 26 $6
438 Llama 3.1 Tulu3 405B 8.1 - $0
439 Ring-flash-2.0 8 - $0.247
440 Pixtral Large 8 - $0
441 Olmo 3.1 32B Think 7.9 - $0
442 GPT-5 nano (minimal) 7.8 - $0.138
443 Gemini 1.5 Flash (Sep '24) 7.8 - $0
444 Grok 2 (Dec '24) 7.8 - $0
445 GPT-4 Turbo 7.7 21.5 $15
446 Qwen3 VL 4B (Reasoning) 7.7 - $0
447 Solar Pro 2 (Non-reasoning) 7.6 - $0
448 Command A 7.5 - $4.375
449 Nova Pro 7.5 - $1.4
450 Llama 3.1 Nemotron Instruct 70B 7.4 - $1.2
451 Qwen3.5 2B (Reasoning) 7.4 2.9 $0
452 Llama 3.1 Instruct 8B 7.4 5.4 $0.028
453 Gemma 3 27B Instruct 7.4 10.1 $0
454 Grok Beta 7.3 - $0
455 NVIDIA Nemotron Nano 9B V2 (Non-reasoning) 7.2 - $0.086
456 NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning) 7.2 - $0.088
457 Qwen2.5 Instruct 32B 7.2 - $0
458 Ministral 3 3B 7.1 4.8 $0.1
459 Mistral Large 2 (Jul '24) 7 - $3
460 Qwen2.5 Coder Instruct 32B 6.9 - $0
461 Qwen3 4B 2507 Instruct 6.9 - $0
462 GPT-4 6.8 13.1 $37.5
463 GLM-4.5V (Non-reasoning) 6.8 - $0.9
464 Qwen3 14B (Non-reasoning) 6.8 - $0.612
465 Hermes 4 - Llama-3.1 70B (Non-reasoning) 6.7 - $0.198
466 GPT-4o mini 6.7 11.4 $0.262
467 Gemini 2.5 Flash-Lite (Non-reasoning) 6.7 - $0.175
468 Mistral Small 3 6.7 - $0.15
469 Nova Lite 6.7 - $0.105
470 Qwen3 30B A3B (Non-reasoning) 6.6 - $0.35
471 Llama 3.1 Instruct 70B 6.5 - $0.56
472 DeepSeek-V2.5 (Dec '24) 6.5 - $0
473 Qwen3 4B (Non-reasoning) 6.5 - $0
474 Granite 4.1 8B 6.4 9.5 $0.063
475 Sarvam 30B (high) 6.4 - $0.047
476 Gemini 2.0 Flash Thinking Experimental (Dec '24) 6.4 - $0
477 DeepSeek-V2.5 6.4 - $0
478 Gemma 4 E2B (Non-reasoning) 6.2 - $0
479 Olmo 3.1 32B Instruct 6.2 - $0
480 Mistral Saba 6.2 - $0
481 DeepSeek R1 Distill Llama 8B 6.2 - $0
482 Gemini 1.5 Pro (May '24) 6.1 19.8 $0
483 Olmo 3 32B Think 6.1 - $0
484 Llama 3.2 Instruct 90B (Vision) 6 - $0
485 R1 1776 6 - $0
486 Solar Mini 6 - $0.15
487 Reka Flash (Sep '24) 6 - $0.35
488 Qwen2.5 Turbo 6 - $0.088
489 Grok-1 5.8 - $0
490 Phi-4 Mini Instruct 5.7 3.8 $0
491 EXAONE 4.0 32B (Non-reasoning) 5.7 - $0
492 Qwen2 Instruct 72B 5.7 - $0
493 Gemma 3 12B Instruct 5.5 5.8 $0
494 Qwen3.5 2B (Non-reasoning) 5.3 2.4 $0
495 Qwen3.5 0.8B (Reasoning) 5.2 0 $0
496 Gemini 1.5 Flash-8B 5.2 - $0
497 DeepHermes 3 - Mistral 24B Preview (Non-reasoning) 5 - $0
498 Jamba 1.7 Large 5 - $0
499 Granite 4.0 H Small 4.9 - $0.107
500 Qwen3 Omni 30B A3B Instruct 4.8 - $0.43
501 Hermes 3 - Llama-3.1 70B 4.8 - $0.7
502 Jamba 1.5 Large 4.8 - $3.5
503 Qwen3 8B (Non-reasoning) 4.8 - $0.31
504 DeepSeek-Coder-V2 4.7 - $0
505 OLMo 2 32B 4.7 - $0
506 Jamba 1.6 Large 4.7 - $0
507 Phi-4 4.6 - $0.219
508 LFM2 24B A2B 4.6 - $0
509 Gemini 1.5 Flash (May '24) 4.6 - $0
510 Nova Micro 4.4 - $0.061
511 Granite 4.1 3B 4.4 4.7 $0
512 Claude 3 Sonnet 4.4 - $0
513 Gemini 1.0 Ultra 4.3 17.6 $0
514 Mistral Small (Sep '24) 4.3 - $0.3
515 Phi-3 Mini Instruct 3.8B 4.3 - $0
516 Phi-4 Multimodal Instruct 4.2 - $0
517 NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) 4.2 - $0.3
518 Gemma 3n E4B Instruct Preview (May '25) 4.2 - $0
519 Mistral Large (Feb '24) 4.1 - $6
520 Qwen2.5 Coder Instruct 7B 4.1 - $0
521 Mixtral 8x22B Instruct 4 - $0
522 Llama 3.2 Instruct 3B 3.9 - $0
523 Llama 2 Chat 7B 3.9 - $0.1
524 MiniCPM-V 4.6 1.3B 3.8 0.7 $0
525 Jamba Reasoning 3B 3.8 - $0
526 Reka Flash 3 3.7 - $0.35
527 Qwen1.5 Chat 110B 3.7 - $0
528 Qwen3 VL 4B Instruct 3.7 - $0
529 Olmo 3 7B Think 3.6 - $0
530 Claude 3 Haiku 3.5 - $0.5
531 Claude 2.1 3.5 14 $0
532 OLMo 2 7B 3.5 - $0
533 Molmo 7B-D 3.4 - $0
534 Ling-mini-2.0 3.4 - $0
535 Claude 2.0 3.3 12.9 $0
536 DeepSeek R1 Distill Qwen 1.5B 3.3 - $0
537 DeepSeek-V2-Chat 3.3 - $0
538 GPT-3.5 Turbo 3.2 10.7 $0.75
539 Mistral Small (Feb '24) 3.2 - $0.262
540 Mistral Medium 3.2 - $3
541 Llama 3 Instruct 70B 3.1 - $1.175
542 Llama 3.2 Instruct 11B (Vision) 3 - $0.345
543 LFM 40B 3 - $0
544 Arctic Instruct 3 - $0
545 Qwen Chat 72B 3 - $0
546 Qwen3.5 0.8B (Non-reasoning) 2.9 1.2 $0
547 PALM-2 2.8 4.6 $0
548 Gemini 1.0 Pro 2.7 - $0
549 DeepSeek Coder V2 Lite Instruct 2.7 - $0
550 Llama 2 Chat 13B 2.6 - $0
551 Llama 2 Chat 70B 2.6 - $0
552 DeepSeek LLM 67B Chat (V1) 2.6 - $0
553 OpenChat 3.5 (1210) 2.6 - $0
554 DBRX Instruct 2.6 - $0
555 Sarvam M (Reasoning) 2.6 - $0
556 Command-R+ (Apr '24) 2.6 - $6
557 Exaone 4.0 1.2B (Reasoning) 2.5 - $0
558 Olmo 3 7B Instruct 2.4 - $0.125
559 Exaone 4.0 1.2B (Non-reasoning) 2.4 - $0
560 LFM2.5-1.2B-Thinking 2.3 - $0
561 LFM2.5-1.2B-Instruct 2.3 - $0
562 LFM2 2.6B 2.3 - $0
563 Jamba 1.7 Mini 2.3 - $0
564 Jamba 1.5 Mini 2.3 - $0.25
565 Granite 4.0 H 1B 2.2 - $0
566 Qwen3 1.7B (Reasoning) 2.2 - $0
567 Jamba 1.6 Mini 2.1 - $0
568 Gemma 3 270M 2 - $0
569 Granite 4.0 Micro 2 - $0
570 Apertus 70B Instruct 2 - $1.345
571 Mixtral 8x7B Instruct 2 - $0.512
572 DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning) 1.9 - $0
573 Llama 65B 1.7 - $0
574 Qwen Chat 14B 1.7 - $0
575 Claude Instant 1.7 7.8 $0
576 Mistral 7B Instruct 1.7 - $0.25
577 Command-R (Mar '24) 1.7 - $0.75
578 Molmo2-8B 1.6 - $0
579 Granite 4.0 1B 1.6 - $0
580 LFM2 8B A1B 1.3 - $0
581 Granite 3.3 8B (Non-reasoning) 1.3 - $0.085
582 Qwen3 1.7B (Non-reasoning) 1.1 - $0
583 LFM2.5-VL-1.6B 1 - $0
584 Granite 4.0 H 350M 1 - $0
585 Granite 4.0 350M 1 - $0
586 Apertus 8B Instruct 1 - $0.125
587 Tiny Aya Global 1 - $0
588 Llama 3 Instruct 8B 1 - $0.07
589 Llama 3.2 Instruct 1B 1 - $0
590 Gemma 3 4B Instruct 1 2.7 $0
591 Gemma 3 1B Instruct 1 - $0
592 Gemma 3n E4B Instruct 1 3.2 $0.075
593 Gemma 3n E2B Instruct 1 - $0
594 LFM2 1.2B 1 - $0
595 Qwen3 0.6B (Reasoning) 1 - $0
596 Qwen3 0.6B (Non-reasoning) 1 - $0
597 GPT-5.5 Pro (xhigh) - - $0
598 Gemini 3 Deep Think - - $0
599 Claude Sonnet 5 (Adaptive Reasoning, Low Effort) - - $4
600 Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort) - - $4
601 Claude Sonnet 5 (Adaptive Reasoning, High Effort) - - $4
602 Claude Sonnet 5 (Adaptive Reasoning, Medium Effort) - - $4
603 EXAONE 4.5 33B (Non-reasoning) - - $0
604 Cogito v2.1 (Reasoning) - - $1.25
605 GPT-4o mini Realtime (Dec '24) - - $0
606 GPT-5.4 Pro (xhigh) - - $67.5
607 GPT-3.5 Turbo (0613) - - $0
608 GPT-4o Realtime (Dec '24) - - $0
609 Mi:dm K 2.5 Pro Preview - - $0

榜单解读建议

参考 AI 大模型排行榜 时,应综合考虑“综合指数”与“成本价格”。如果您是开发者,编程能力 (Coding) 是更核心的指标。

值品工具箱同步的 AI 大模型排行榜 数据每 24 小时更新,确保您获取到最新的模型性能对比。

指标说明

  • 综合指数:评估通用理解与逻辑。
  • 价格 $/1M:混合 3:1 输入输出比的平均成本。
  • 编程能力:衡量代码生成的准确性。

AI 大模型排行榜 常见问题 (FAQ)

Q1: AI 大模型排行榜 的数据多久更新?

AI 大模型排行榜 数据每 24 小时自动抓取一次,确保最新模型加入列表。

Q2: 这个 AI 大模型排行榜 包含国产模型吗?

是的,只要国产模型通过了 Artificial Analysis 的全球测评,就会出现在 AI 大模型排行榜 中。

Q3: 综合指数在 AI 大模型排行榜 中代表什么?

它代表模型的全能表现。AI 大模型排行榜 通过加权算法给出这个综合评分。

Q4: 如何在 AI 大模型排行榜 中查找性价比最高的游戏?

在 AI 大模型排行榜 页面中,您可以点击“价格”标题进行排序,寻找低价高分的模型。

Q5: AI 大模型排行榜 的编程能力测试准吗?

AI 大模型排行榜 参考了 LiveCodeBench 等权威基准测试,具有极高的参考价值。

Q6: 为什么有的新模型没进入 AI 大模型排行榜?

模型进入 AI 大模型排行榜 需要经过一系列测试,通常在新模型发布后数日内会完成更新。

Q7: AI 大模型排行榜 中的价格计算标准是什么?

价格是基于百万 Token 的调用成本,由 AI 大模型排行榜 统一混合计算得出。

Q8: 手机上能查看 AI 大模型排行榜 吗?

当然可以。AI 大模型排行榜 进行了移动端响应式深度优化。

Q9: AI 大模型排行榜 这个工具免费吗?

是的,由值品工具箱免费提供 AI 大模型排行榜 信息查询服务。

Q10: 我该怎么利用 AI 大模型排行榜 做选型?

如果您需要智能客服,参考 AI 大模型排行榜 的综合指数;如果做翻译,参考编程外的语言指标。

发表评论

请友善文明留言