AI 大模型排行榜 (Artificial Analysis LLM Ranking)

信息查询
3.2k 次浏览
100% 有帮助 · 1 人反馈

值品工具箱提供的 AI 大模型排行榜聚合了来自 Artificial Analysis 的权威数据,实时追踪并排名超过 100 个主流大语言模型。

AI 大模型排行榜数据中心

重置
排名 模型名称 综合指数 ▼ 编程 价格 ($/1M)
1 Claude Opus 5 (Adaptive Reasoning, Max Effort) 63.1 78 $10
2 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) 62.5 77 $10
3 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) 62.1 76.5 $20
4 Claude Opus 5 (Adaptive Reasoning, High Effort) 61.5 76.5 $10
5 GPT-5.6 Sol (max) 60.9 77.4 $11.25
6 Grok 4.6 (high) 60.9 76.8 $3
7 Kimi K3 (max) 59.7 76.2 $6
8 GPT-5.6 Sol (xhigh) 59 78.3 $11.25
9 Claude Opus 5 (Adaptive Reasoning, Medium Effort) 58.6 74.3 $10
10 Qwen3.8 Max 58.1 71.8 $3
11 Qwen3.8 2.4T A95B 57.7 71.9 $3
12 GPT-5.6 Sol (high) 57.3 77.2 $11.25
13 Claude Opus 4.8 (Adaptive Reasoning, Max Effort) 57.3 74.3 $10
14 Muse Spark 1.2 (xhigh) 56.8 72.2 $2
15 GPT-5.6 Terra (max) 56.6 76.7 $4.5
16 GPT-5.5 (xhigh) 56.3 74.9 $11.25
17 Gemini 3.7 Flash (high) 56 76.1 $1.5
18 Grok 4.5 (high) 55.8 72.4 $3
19 GPT-5.6 Sol (medium) 55.6 76.3 $11.25
20 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) 55.3 71.5 $4
21 Claude Opus 4.7 (Adaptive Reasoning, Max Effort) 55 73.6 $10
22 GPT-5.5 (high) 54.7 71.6 $11.25
23 Gemini 3.7 Flash (medium) 53.4 71.5 $1.5
24 DeepSeek V4 Pro 0813 (Reasoning, Max Effort) 53.2 68.8 $1.98
25 Muse Spark 1.1 (xhigh) 53.2 71.3 $2
26 GPT-5.4 (xhigh) 53.1 71.1 $5.625
27 GPT-5.6 Terra (xhigh) 52.8 70.6 $4.5
28 GLM-5.2 (max) 52.6 68.8 $2.15
29 Claude Opus 5 (Adaptive Reasoning, Low Effort) 52.5 66.9 $10
30 GPT-5.6 Luna (max) 52.3 71.4 $0.45
31 Gemini 3.5 Flash (high) 52 70.1 $3.375
32 DeepSeek V4 Flash 0731 (Reasoning, Max Effort) 51.8 69.1 $0.175
33 Gemini 3.6 Flash (high) 51.6 69.2 $3
34 GPT-5.5 (medium) 51.4 71.5 $11.25
35 Gemini 3.7 Flash (low) 50.9 71 $1.5
36 GPT-5.6 Sol (low) 50.7 69.7 $11.25
37 GPT-5.6 Luna (xhigh) 50.1 68.6 $0.45
38 GPT-5.6 Terra (high) 50.1 67.1 $4.5
39 Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) 48.4 63 $6
40 Kimi K3 (low) 48.3 72 $6
41 Gemini 3.1 Pro Preview 47.7 68.8 $4.5
42 Motif 3 47.4 63.5 $0
43 GPT-5.6 Luna (high) 47 63.3 $0.45
44 GPT-5.6 Terra (medium) 46.8 64.7 $4.5
45 Gemini 3.5 Flash (medium) 46.7 - $3.375
46 Qwen3.7 Max 46.7 66 $3.75
47 GPT-5.3 Codex (xhigh) 45.5 - $4.813
48 MiniMax-M3 45.4 58.6 $0.525
49 DeepSeek V4 Pro (Reasoning, Max Effort) 45.3 59.4 $0.544
50 Motif 3 (Beta) 45.3 62 $0
51 Kimi K2.6 45.1 61.8 $1.712
52 Claude Opus 4.6 (Adaptive Reasoning, Max Effort) 44.9 - $10
53 GPT-5.5 (low) 44.5 60.9 $11.25
54 Muse Spark 44.3 58.6 $0
55 Claude Opus 4.7 (Non-reasoning, High Effort) 43.9 - $10
56 DeepSeek V4 Pro (Reasoning, High Effort) 43.7 58.7 $0.544
57 GPT-5.2 (xhigh) 43.3 - $4.813
58 Kimi K2.7 Code 43 60.8 $1.712
59 MiMo-V2.5-Pro 42.9 60.2 $0.544
60 Claude Sonnet 5 (Non-reasoning, High Effort) 42.6 66.4 $4
61 Inkling (xhigh) 42.3 52.1 $1.762
62 Hy3 42.2 58.8 $0.241
63 Nex-N2-Pro 42.1 59.1 $1
64 DeepSeek V4 Flash (Reasoning, Max Effort) 42.1 56.2 $0.168
65 GPT-5.6 Sol (Non-reasoning) 41.9 65.1 $11.25
66 Claude Opus 4.5 (Reasoning) 41.9 - $10
67 Solar Pro 4 41.6 52.7 $0.525
68 MiMo-V2-Pro 41.4 - $0
69 GPT-5.6 Terra (low) 41.3 58.1 $4.5
70 Inkling Small 41.2 52.9 $0.525
71 GPT-5.2 Codex (xhigh) 41.2 - $4.813
72 Qwen3.6 Max Preview 41.1 - $2.925
73 GLM-5.1 (Reasoning) 41 55.8 $2.135
74 GPT-5.4 mini (xhigh) 40.9 56.1 $1.688
75 Grok Build 0.1 0616 40.7 51.5 $1.25
76 Gemini 3 Pro Preview (high) 40.6 - $4.5
77 GLM-5 (Reasoning) 40.6 - $1.55
78 Qwen3.6 Plus 40.5 54.5 $1.125
79 GPT-5.4 (low) 40.2 - $5.625
80 JT-4.1 Flash 236B A21B 39.9 52.4 $0
81 Agnes 2.5 Pro Alpha 39.7 58.8 $0.563
82 GPT-5.4 nano (xhigh) 39.7 56.1 $0.463
83 Qwen3.7 Plus 39.4 55.9 $0.7
84 GLM-5-Turbo 39.1 - $0
85 DeepSeek V4 Flash (Reasoning, High Effort) 39 52 $0.175
86 GPT-5.6 Luna (medium) 38.9 50.7 $0.45
87 GPT-5.2 (medium) 38.9 - $4.813
88 MiniMax-M2.7 38.9 52.6 $0.525
89 Claude Opus 4.6 (Non-reasoning, High Effort) 38.8 - $10
90 Gemini 3 Flash Preview (Reasoning) 38.7 - $1.125
91 Nemotron 3 Ultra 550B A55B (Reasoning) 38.3 49.3 $1.137
92 MiMo-V2.5 38 56.8 $0.175
93 Grok 4.20 0309 v2 (Reasoning) 38 - $1.563
94 Grok 4.3 (high) 37.9 42.2 $1.563
95 Ling 3.0 Flash 37.8 50.6 $0.111
96 Qwen3.6 27B (Reasoning) 37.7 53.7 $1.35
97 GPT-5.1 (high) 37.5 49.4 $3.438
98 Gemini 3.5 Flash-Lite 37.4 49.3 $0.85
99 Solar Open2 250B 37.4 44.7 $0
100 Claude 4.5 Sonnet (Reasoning) 37.4 52.1 $6
101 Grok 4.20 0309 (Reasoning) 37.4 - $3
102 MiMo-V2-Omni-0327 37.3 - $0
103 GPT-5 Codex (high) 37 - $3.438
104 Grok 4.3 (medium) 36.9 - $1.563
105 Claude Sonnet 4.6 (Non-reasoning, High Effort) 36.8 - $6
106 Grok 4.3 (low) 36.3 - $1.563
107 GLM-5.1 (Non-reasoning) 36.3 - $2.135
108 Kimi K2.5 (Reasoning) 36 46.8 $1.2
109 MiMo-V2-Omni 35.9 - $0
110 Gemini 3.5 Flash (minimal) 35.8 - $3.375
111 GPT-5.5 (Non-reasoning) 35.8 56.5 $11.25
112 GPT-5.1 Codex (high) 35.6 - $3.438
113 Claude Opus 4.5 (Non-reasoning) 35.6 - $10
114 Kimi K2.6 (Non-reasoning) 35.4 - $1.712
115 GPT-5 (high) 35.3 37.8 $3.438
116 GLM 5V Turbo (Reasoning) 35.3 - $0
117 Muse Glimmer (high) 35.1 49 $0.581
118 Claude Sonnet 4.6 (Non-reasoning, Low Effort) 35.1 - $6
119 A.X-K2 35 38.8 $0
120 GLM-5.2 (Non-reasoning) 34.8 46.5 $2.15
121 GPT-5.6 Terra (Non-reasoning) 34.6 52.3 $4.5
122 GPT-5 (medium) 34.6 - $3.438
123 Qwen3.5 27B (Reasoning) 34.6 - $0.825
124 KAT Coder Pro V2 34.5 59.5 $0.525
125 Claude 4.1 Opus (Reasoning) 34.5 - $30
126 MiniMax-M2.5 34.5 - $0.525
127 GLM-4.7 (Reasoning) 34.5 45.3 $1
128 Hy3-preview (Reasoning) 34.4 - $0.1
129 LongCat 2.0 34.3 45.3 $1.3
130 Qwen3.5 397B A17B (Reasoning) 34.3 48.2 $1.35
131 GPT-5.5 Instant (May 2026) 34.3 - $11.25
132 Grok 4 34.1 - $6
133 MiMo-V2-Flash (Feb 2026) 34 - $0
134 GPT-5.6 Luna (low) 33.9 44.2 $0.45
135 Gemini 3 Pro Preview (low) 33.9 - $4.5
136 Kimi K2 Thinking 33.5 - $1.075
137 o3-pro 33.3 - $35
138 GLM-5 (Non-reasoning) 33.2 - $1.55
139 Qwen3.5 122B A10B (Reasoning) 32.8 45.7 $1.1
140 DeepSeek V3.2 (Reasoning) 32.8 44.2 $0.315
141 Qwen3.5 397B A17B (Non-reasoning) 32.7 - $1.35
142 Qwen3 Max Thinking 32.5 - $0
143 Qwen3.6 35B A3B (Reasoning) 32.1 41.9 $0.844
144 MiniMax-M2.1 32.1 - $0.525
145 DeepSeek V4 Pro (Non-reasoning) 31.9 - $0.544
146 GPT-5 (low) 31.9 - $3.438
147 MiMo-V2-Flash (Reasoning) 31.9 - $0.15
148 Claude 4 Opus (Reasoning) 31.7 - $30
149 G9v3-39A5B 31.6 31.7 $0
150 GPT-5 mini (medium) 31.6 - $0.688
151 Qwen3.6 27B (Non-reasoning) 31.3 46.6 $1.35
152 Qwen3.5 Omni Plus 31.3 - $1.5
153 Ring-2.6-1T 31.3 42.8 $0.85
154 GPT-5.1 Codex mini (high) 31.3 - $0.688
155 Grok 4.1 Fast (Reasoning) 31.3 - $0
156 o3 31.1 - $3.5
157 DeepSeek V3.1 Terminus (Reasoning) 31.1 43.5 $1.914
158 K-EXAONE 2.0 0803 31 40.6 $0
159 Step 3.7 Flash 30.9 39.6 $0.438
160 GPT-5.4 nano (medium) 30.8 - $0.463
161 GPT-5.4 mini (medium) 30.5 - $1.688
162 Mistral Medium 3.5 30.4 46.9 $3
163 Kimi K2.5 (Non-reasoning) 30.1 - $1.2
164 Qwen3.5 27B (Non-reasoning) 30 - $0.825
165 Claude 4.5 Haiku (Reasoning) 29.9 43.9 $2
166 Claude 4.5 Sonnet (Non-reasoning) 29.9 - $6
167 Qwen3.5 35B A3B (Reasoning) 29.9 - $0.688
168 Claude 4 Sonnet (Reasoning) 29.8 37.6 $6
169 Gemma 4 31B (Reasoning) 29.7 43.4 $0
170 DeepSeek V4 Flash (Non-reasoning) 29.3 - $0.175
171 GLM-4.6 (Reasoning) 29.3 45.8 $0.963
172 GPT-5.5 Instant (June 2026) 29.2 39.4 $11.25
173 JT-35B-Flash 29 - $0
174 MiniMax-M2 28.9 - $0.525
175 KAT-Coder-Pro V1 28.9 - $0
176 Claude 4.1 Opus (Non-reasoning) 28.8 - $30
177 MiMo-V2.5-Pro (Non-reasoning) 28.4 - $0.544
178 GPT-5.4 (Non-reasoning) 28.3 - $5.625
179 Qwen3.5 122B A10B (Non-reasoning) 28.2 43.3 $1.1
180 Gemini 3 Flash Preview (Non-reasoning) 27.9 - $1.125
181 Grok 4 Fast (Reasoning) 27.9 - $0.275
182 Claude 3.7 Sonnet (Reasoning) 27.6 36.4 $0
183 GLM-4.7 (Non-reasoning) 27.1 - $1
184 GPT-5.6 Luna (Non-reasoning) 26.8 39.3 $0.45
185 Hy3-preview (Non-reasoning) 26.6 - $0.1
186 Ling-2.6-1T 26.6 - $0.85
187 Doubao Seed Code 26.5 - $0
188 GPT-5.2 (Non-reasoning) 26.5 - $4.813
189 Step 3.5 Flash 2603 26.5 - $0.15
190 Gemma 4 26B A4B (Reasoning) 26.1 39.3 $0.198
191 o4-mini (high) 26.1 - $1.925
192 Claude 4 Opus (Non-reasoning) 26 - $30
193 Claude 4 Sonnet (Non-reasoning) 26 - $6
194 Step 3.5 Flash 26 - $0.15
195 Gemini 2.5 Pro 25.9 33.3 $3.438
196 DeepSeek V3.2 Exp (Reasoning) 25.9 - $0.315
197 GPT-5 mini (high) 25.8 15.6 $0.688
198 Nemotron 3 Super 120B A12B (Reasoning) 25.7 37.7 $0.35
199 Gemini 3.1 Flash-Lite 25.6 34.7 $0.563
200 Qwen3 Max Thinking (Preview) 25.5 - $2.4
201 MiMo-V2-Flash (Non-reasoning) 25.1 49.8 $0
202 DeepSeek V3.2 (Non-reasoning) 25.1 - $0.315
203 Grok 4.3 (Non-reasoning) 25 35.2 $1.563
204 Qwen3.6 35B A3B (Non-reasoning) 24.6 28.1 $0.844
205 Ling 3.0 Tiny 24.5 26.5 $0
206 Qwen3 Max 24.5 - $2.4
207 Qwen3.5 35B A3B (Non-reasoning) 24.3 37 $0.688
208 Gemini 2.5 Flash Preview (Sep '25) (Reasoning) 24.2 - $0
209 gpt-oss-120b (high) 24.1 30.4 $0.262
210 Claude 4.5 Haiku (Non-reasoning) 24.1 - $2
211 Kimi K2 0905 24 - $1.075
212 o1 23.9 39.7 $26.25
213 Claude 3.7 Sonnet (Non-reasoning) 23.9 - $6
214 Nemotron 3.5 Lightning 23.6 26.8 $0.088
215 Gemini 2.5 Pro Preview (Mar' 25) 23.4 46.7 $0
216 GLM-4.6 (Non-reasoning) 23.4 - $0.981
217 GLM-4.7-Flash (Reasoning) 23.3 - $0.153
218 Grok 3 mini Reasoning (high) 22.9 - $0.35
219 Grok 4.20 0309 (Non-reasoning) 22.9 - $3
220 Command A+ 22.8 27.8 $0
221 Gemini 2.5 Pro Preview (May' 25) 22.7 - $3.438
222 DeepSeek V3.2 Speciale 22.6 - $0
223 K-EXAONE (Reasoning) 22.5 32.1 $0
224 Gemma 4 31B (Non-reasoning) 22.3 33.2 $0.209
225 ERNIE 5.0 Thinking Preview 22.3 - $0
226 Gemma 4 12B (Reasoning) 22.2 31 $0.15
227 Grok 4.20 0309 v2 (Non-reasoning) 22.2 - $1.563
228 Nova 2.0 Pro Preview (medium) 22.1 34 $3.438
229 Grok Code Fast 1 22 - $0
230 Mercury 2 21.9 31.1 $0.375
231 Qwen3.5 9B (Reasoning) 21.8 28.7 $0.151
232 DeepSeek V3.1 Terminus (Non-reasoning) 21.7 - $0.453
233 DeepSeek V3.2 Exp (Non-reasoning) 21.7 - $0.315
234 Apriel-v1.5-15B-Thinker 21.6 - $0
235 DeepSeek V3.1 (Non-reasoning) 21.4 - $0.84
236 Nova 2.0 Omni (medium) 21.3 - $0.85
237 Qwen3 Coder Next 21.3 36.2 $0.563
238 DeepSeek V3.1 (Reasoning) 21 - $0.865
239 Qwen3 VL 235B A22B (Reasoning) 20.9 - $2.625
240 Nova 2.0 Lite (high) 20.8 23 $0.85
241 Apriel-v1.6-15B-Thinker 20.8 - $0
242 GPT-5.1 (Non-reasoning) 20.7 - $3.438
243 Qwen3.5 9B (Non-reasoning) 20.6 23.5 $0.19
244 EXAONE 4.5 33B 20.5 23.6 $0
245 Gemma 4 26B A4B (Non-reasoning) 20.4 - $0.198
246 Qwen3.5 4B (Reasoning) 20.4 22.6 $0.06
247 DeepSeek R1 0528 (May '25) 20.4 - $2.063
248 Gemini 2.5 Flash (Reasoning) 20.3 - $0.85
249 North Mini Code 20.2 36.5 $0
250 GPT-5 nano (high) 20.1 - $0.138
251 Qwen3 235B A22B 2507 (Reasoning) 19.9 22.1 $2.625
252 Nova 2.0 Pro Preview (low) 19.8 25.9 $3.438
253 Mistral Small 4 (Reasoning) 19.7 26.6 $0.262
254 Kimi K2 19.7 - $1.002
255 GLM-4.5 (Reasoning) 19.7 - $0
256 GPT-4.1 19.6 - $3.5
257 Qwen3 Max (Preview) 19.4 - $2.4
258 Devstral 2 19.2 31.3 $0
259 Nova 2.0 Lite (medium) 19.2 - $0.85
260 Qwen3.5 Omni Flash 19.2 - $0.275
261 o3-mini 19.2 - $1.925
262 GPT-5 nano (medium) 19.2 - $0.138
263 o1-pro 19.1 - $262.5
264 Gemini 2.5 Flash Preview (Sep '25) (Non-reasoning) 19.1 - $0
265 JT-MINI 18.8 - $0
266 DeepSeek R1 (Jan '25) 18.6 24.6 $2.431
267 Grok 3 18.6 - $8
268 Seed-OSS-36B-Instruct 18.5 - $0.3
269 Trinity Large Thinking 18.4 25.8 $0.395
270 Qwen3 235B A22B 2507 Instruct 18.4 - $1.225
271 HyperNova 60B 2605 18.3 23.2 $0
272 Qwen3 Coder 480B A35B Instruct 18.2 - $3
273 Qwen3 VL 32B (Reasoning) 18.1 - $2.625
274 Magistral Medium 1.2 18 21.3 $2.75
275 Nova 2.0 Lite (low) 18 - $0.85
276 Sonar Reasoning Pro 18 - $0
277 MiniMax M1 80k 17.9 - $0.963
278 Nemotron Cascade 2 30B A3B 17.8 25.3 $0
279 GPT-5.4 nano (Non-Reasoning) 17.8 - $0.463
280 Devstral Small 2 17.7 29.3 $0
281 Gemini 2.5 Flash Preview (Reasoning) 17.7 - $0
282 K2 Think V2 17.4 21 $0
283 LongCat Flash Lite 17.4 - $0
284 GPT-5 (minimal) 17.3 - $3.438
285 HyperCLOVA X SEED Think (32B) 17.2 - $0
286 o1-preview 17.2 34 $28.875
287 Grok 4.1 Fast (Non-reasoning) 17 - $0
288 K-EXAONE (Non-reasoning) 16.9 - $0
289 Qwen3 Next 80B A3B (Reasoning) 16.9 17.4 $1.875
290 GLM-4.6V (Reasoning) 16.9 - $0.45
291 GPT-5.4 mini (Non-Reasoning) 16.8 - $1.688
292 Nova 2.0 Omni (low) 16.7 - $0.85
293 GLM-4.5-Air 16.7 - $0.372
294 Mi:dm K 2.5 Pro 16.6 - $0
295 Grok 4 Fast (Non-reasoning) 16.6 - $0.275
296 Ring-1T 16.3 - $0
297 G9v3-3B 16.2 9.9 $0
298 Qwen3.5 4B (Non-reasoning) 16.1 20.3 $0.06
299 Mistral Large 3 15.9 20.1 $0.75
300 INTELLECT-3 15.7 - $0
301 o3-mini (high) 15.7 16.3 $1.925
302 GLM-4.7-Flash (Non-reasoning) 15.6 - $0.153
303 GPT-5 (ChatGPT) 15.4 - $0
304 gpt-oss-20b (high) 15.2 20.7 $0.092
305 Solar Open 100B (Reasoning) 15.2 - $0
306 Gemini 2.5 Flash-Lite Preview (Sep '25) (Reasoning) 15.2 - $0.175
307 DeepSeek V3 0324 15.2 21.2 $0.483
308 Grok 3 Reasoning Beta 15.2 - $0
309 Nemotron 3 Nano Omni 30B A3B Reasoning 15 13.8 $0.131
310 gpt-oss-120b (low) 14.9 21.2 $0.261
311 Mistral Small 3.1 14.9 26.3 $0.15
312 GPT-4.1 mini 14.8 20.2 $0.7
313 Mistral Medium 3.1 14.7 20.5 $0.8
314 Qwen3 30B A3B 2507 (Reasoning) 14.6 12.1 $0.75
315 Llama 4 Maverick 14.5 16.3 $0.415
316 Solar Pro 3 14.5 16.2 $0.262
317 NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) 14.5 14.4 $0.088
318 MiniMax M1 40k 14.5 - $0
319 gpt-oss-20b (low) 14.4 - $0.103
320 Nova 2.0 Pro Preview (Non-reasoning) 14.4 20.9 $3.438
321 Qwen3 VL 235B A22B Instruct 14.4 - $1.225
322 GPT-5 mini (minimal) 14.3 - $0.688
323 K2-V2 (high) 14.2 - $0
324 Gemini 2.5 Flash (Non-reasoning) 14.2 - $0.85
325 DeepSeek V3 (Dec '24) 14.2 23 $0.493
326 Ling 2.6 Flash 14.2 25.3 $0.15
327 o1-mini 14 - $0
328 Qwen3 Next 80B A3B Instruct 13.8 - $0.875
329 GPT-4.5 (Preview) 13.6 - $0
330 Tri-21B-think Preview 13.6 - $0
331 Qwen3 Coder 30B A3B Instruct 13.6 - $0.9
332 DiffusionGemma 26B A4B 13.5 19.7 $0
333 Qwen3 235B A22B (Reasoning) 13.5 - $2.625
334 QwQ 32B 13.4 - $0.745
335 Qwen3 VL 30B A3B (Reasoning) 13.4 - $0.75
336 Gemini 2.0 Flash Thinking Experimental (Jan '25) 13.3 24.1 $0
337 Gemma 4 12B (Non-reasoning) 13.2 - $0.15
338 Gemini 2.5 Flash-Lite Preview (Sep '25) (Non-reasoning) 13.1 - $0.175
339 Motif-2-12.7B-Reasoning 12.8 - $0
340 Nova Premier 12.7 - $5
341 Ling-1T 12.7 - $0
342 Mistral Medium 3 12.5 - $0.8
343 Magistral Medium 1 12.5 - $0
344 Solar Pro 2 (Preview) (Reasoning) 12.5 - $0
345 Llama Nemotron Super 49B v1.5 (Reasoning) 12.4 - $0.4
346 K2-V2 (medium) 12.4 - $0
347 Celeris-1 12.4 14.4 $0.325
348 Devstral Medium 12.4 - $0
349 Mistral Small 4 (Non-reasoning) 12.3 - $0.262
350 Tri-21B-Think 12.3 - $0
351 GPT-4o (March 2025, chatgpt-4o-latest) 12.3 - $0
352 Gemma 4 E4B (Reasoning) 12.2 9.4 $0.04
353 Gemini 2.0 Flash (Feb '25) 12.2 - $0
354 Claude 3.5 Haiku 12.2 15.9 $0
355 Llama 3.3 Nemotron Super 49B v1 (Reasoning) 12.2 - $0
356 Sarvam 105B (high) 11.9 - $0.074
357 MiniCPM5-1B (Reasoning) 11.9 - $0
358 Qwen3 4B 2507 (Reasoning) 11.9 - $0
359 Nova 2.0 Lite (Non-reasoning) 11.8 - $0.85
360 Gemini 2.0 Pro Experimental (Feb '25) 11.8 25.5 $0
361 Claude 3 Opus 11.8 19.5 $30
362 Devstral Small (May '25) 11.8 - $0
363 MiniCPM5-1B (Non-reasoning) 11.7 - $0
364 Gemini 2.5 Flash Preview (Non-reasoning) 11.6 - $0
365 Sonar Reasoning 11.6 - $0
366 Magistral Small 1.2 11.5 14.7 $0.75
367 Gemini 2.5 Flash-Lite (Reasoning) 11.4 - $0.175
368 Qwen3 32B (Reasoning) 11.4 15.3 $2.625
369 Ministral 3 14B 11.2 14.4 $0.2
370 GPT-4o (Nov '24) 11.1 - $4.375
371 Nanbeige4.1-3B 11 9.6 $0
372 DeepSeek R1 Distill Qwen 32B 11 - $0
373 Qwen3 VL 32B Instruct 11 - $1.225
374 GLM-4.6V (Non-reasoning) 10.9 - $0.45
375 Qwen3 235B A22B (Non-reasoning) 10.8 - $1.225
376 Mistral Small 3.2 10.7 12.5 $0.15
377 Gemini 2.0 Flash (experimental) 10.6 - $0
378 Magistral Small 1 10.6 - $0
379 EXAONE 4.0 32B (Reasoning) 10.5 - $0
380 Qwen3 VL 8B (Reasoning) 10.5 - $0.66
381 Nova 2.0 Omni (Non-reasoning) 10.4 - $0.85
382 Qwen3 14B (Reasoning) 10.4 13.8 $1.313
383 Llama 4 Scout 10.3 8.2 $0.3
384 DeepSeek R1 0528 Qwen3 8B 10.3 - $0
385 Qwen2.5 Max 10.1 - $0
386 Hermes 4 - Llama-3.1 70B (Reasoning) 9.9 - $0.198
387 Gemini 1.5 Pro (Sep '24) 9.9 23.6 $0
388 Solar Pro 2 (Preview) (Non-reasoning) 9.9 - $0
389 Qwen3 VL 30B A3B Instruct 9.9 - $0.35
390 Claude 3.5 Sonnet (Oct '24) 9.8 30.2 $6
391 DeepSeek R1 Distill Llama 70B 9.8 - $0.8
392 Falcon-H1R-7B 9.7 - $0
393 DeepSeek R1 Distill Qwen 14B 9.7 - $0
394 GPT-4.1 nano 9.6 11.1 $0.175
395 Ling-flash-2.0 9.6 - $0.247
396 Gemma 4 E2B (Reasoning) 9.5 7.2 $0
397 Qwen3 Omni 30B A3B (Reasoning) 9.5 - $0.43
398 GPT-4o (Aug '24) 9.4 - $4.375
399 Sonar 9.4 - $0
400 Qwen2.5 Instruct 72B 9.4 - $0.48
401 Llama 3.3 Instruct 70B 9.3 11.9 $0.671
402 Step3 VL 10B 9.3 - $0
403 Qwen3 30B A3B (Reasoning) 9.2 - $0.75
404 Devstral Small (Jul '25) 9.1 - $0
405 Sonar Pro 9.1 - $0
406 QwQ 32B-Preview 9.1 - $0
407 Ministral 3 8B 9 9.7 $0.15
408 Mistral Large 2 (Nov '24) 9 - $0
409 GLM-4.5V (Reasoning) 9 - $0.9
410 Llama 3.1 Nemotron Ultra 253B v1 (Reasoning) 8.9 - $0.9
411 ERNIE 4.5 300B A47B 8.9 - $0.485
412 Qwen3 30B A3B 2507 Instruct 8.9 - $0.35
413 NVIDIA Nemotron Nano 12B v2 VL (Reasoning) 8.8 - $0.3
414 Hermes 4 - Llama-3.1 405B (Reasoning) 8.8 - $1.5
415 Solar Pro 2 (Reasoning) 8.8 - $0
416 Gemma 4 E4B (Non-reasoning) 8.7 - $0.04
417 NVIDIA Nemotron Nano 9B V2 (Reasoning) 8.7 - $0.07
418 Granite 4.1 30B 8.7 10.4 $0
419 NVIDIA Nemotron 3 Nano 4B 8.6 8 $0
420 Hermes 4 - Llama-3.1 405B (Non-reasoning) 8.6 - $1.5
421 Gemini 2.0 Flash-Lite (Feb '25) 8.6 - $0
422 Llama Nemotron Super 49B v1.5 (Non-reasoning) 8.5 - $0.4
423 Qwen3 32B (Non-reasoning) 8.5 - $1.225
424 Kimi Linear 48B A3B Instruct 8.4 - $0
425 K2-V2 (low) 8.4 - $0
426 GPT-4o (May '24) 8.4 24.2 $7.5
427 Gemini 2.0 Flash-Lite (Preview) 8.4 - $0
428 Llama 3.1 Nemotron Nano 4B v1.1 (Reasoning) 8.4 - $0
429 Llama 3.1 Instruct 405B 8.3 - $0
430 Llama 3.3 Nemotron Super 49B v1 (Non-reasoning) 8.3 - $0
431 Qwen3 8B (Reasoning) 8.3 9 $0.66
432 Qwen3 VL 8B Instruct 8.2 - $0.31
433 Qwen3 4B (Reasoning) 8.2 - $0
434 LFM2.5-8B-A1B 8.1 - $0
435 GPT-4o (ChatGPT) 8.1 - $0
436 Claude 3.5 Sonnet (June '24) 8.1 26 $6
437 Llama 3.1 Tulu3 405B 8.1 - $0
438 Ring-flash-2.0 8 - $0.247
439 Pixtral Large 8 - $0
440 Olmo 3.1 32B Think 7.9 - $0
441 GPT-5 nano (minimal) 7.8 - $0.138
442 Gemini 1.5 Flash (Sep '24) 7.8 - $0
443 Grok 2 (Dec '24) 7.8 - $0
444 GPT-4 Turbo 7.7 21.5 $15
445 Qwen3 VL 4B (Reasoning) 7.7 - $0
446 Solar Pro 2 (Non-reasoning) 7.6 - $0
447 Command A 7.5 - $4.375
448 Nova Pro 7.5 - $1.4
449 Llama 3.1 Nemotron Instruct 70B 7.4 - $1.2
450 Qwen3.5 2B (Reasoning) 7.4 2.9 $0
451 Llama 3.1 Instruct 8B 7.4 5.4 $0.028
452 Gemma 3 27B Instruct 7.4 10.1 $0
453 Grok Beta 7.3 - $0
454 NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning) 7.2 - $0.088
455 NVIDIA Nemotron Nano 9B V2 (Non-reasoning) 7.2 - $0.086
456 Qwen2.5 Instruct 32B 7.2 - $0
457 Ministral 3 3B 7.1 4.8 $0.1
458 Mistral Large 2 (Jul '24) 7 - $3
459 Qwen2.5 Coder Instruct 32B 6.9 - $0
460 Qwen3 4B 2507 Instruct 6.9 - $0
461 GPT-4 6.8 13.1 $37.5
462 GLM-4.5V (Non-reasoning) 6.8 - $0.9
463 Qwen3 14B (Non-reasoning) 6.8 - $0.612
464 Hermes 4 - Llama-3.1 70B (Non-reasoning) 6.7 - $0.198
465 GPT-4o mini 6.7 11.4 $0.262
466 Gemini 2.5 Flash-Lite (Non-reasoning) 6.7 - $0.175
467 Mistral Small 3 6.7 - $0.15
468 Nova Lite 6.7 - $0.105
469 Qwen3 30B A3B (Non-reasoning) 6.6 - $0.35
470 Llama 3.1 Instruct 70B 6.5 - $0.56
471 DeepSeek-V2.5 (Dec '24) 6.5 - $0
472 Qwen3 4B (Non-reasoning) 6.5 - $0
473 Granite 4.1 8B 6.4 9.5 $0.063
474 Sarvam 30B (high) 6.4 - $0.047
475 Gemini 2.0 Flash Thinking Experimental (Dec '24) 6.4 - $0
476 DeepSeek-V2.5 6.4 - $0
477 Gemma 4 E2B (Non-reasoning) 6.2 - $0
478 Olmo 3.1 32B Instruct 6.2 - $0
479 Mistral Saba 6.2 - $0
480 DeepSeek R1 Distill Llama 8B 6.2 - $0
481 Gemini 1.5 Pro (May '24) 6.1 19.8 $0
482 Olmo 3 32B Think 6.1 - $0
483 Llama 3.2 Instruct 90B (Vision) 6 - $0
484 R1 1776 6 - $0
485 Solar Mini 6 - $0.15
486 Reka Flash (Sep '24) 6 - $0.35
487 Qwen2.5 Turbo 6 - $0.088
488 Grok-1 5.8 - $0
489 Phi-4 Mini Instruct 5.7 3.8 $0
490 EXAONE 4.0 32B (Non-reasoning) 5.7 - $0
491 Qwen2 Instruct 72B 5.7 - $0
492 Gemma 3 12B Instruct 5.5 5.8 $0
493 Qwen3.5 2B (Non-reasoning) 5.3 2.4 $0
494 Qwen3.5 0.8B (Reasoning) 5.2 0 $0
495 Gemini 1.5 Flash-8B 5.2 - $0
496 DeepHermes 3 - Mistral 24B Preview (Non-reasoning) 5 - $0
497 Jamba 1.7 Large 5 - $0
498 Granite 4.0 H Small 4.9 - $0.107
499 Qwen3 Omni 30B A3B Instruct 4.8 - $0.43
500 Hermes 3 - Llama-3.1 70B 4.8 - $0.7
501 Jamba 1.5 Large 4.8 - $3.5
502 Qwen3 8B (Non-reasoning) 4.8 - $0.31
503 DeepSeek-Coder-V2 4.7 - $0
504 OLMo 2 32B 4.7 - $0
505 Jamba 1.6 Large 4.7 - $0
506 Phi-4 4.6 - $0.219
507 LFM2 24B A2B 4.6 - $0
508 Gemini 1.5 Flash (May '24) 4.6 - $0
509 Nova Micro 4.4 - $0.061
510 Granite 4.1 3B 4.4 4.7 $0
511 Claude 3 Sonnet 4.4 - $0
512 Gemini 1.0 Ultra 4.3 17.6 $0
513 Mistral Small (Sep '24) 4.3 - $0.3
514 Phi-3 Mini Instruct 3.8B 4.3 - $0
515 Phi-4 Multimodal Instruct 4.2 - $0
516 NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) 4.2 - $0.3
517 Gemma 3n E4B Instruct Preview (May '25) 4.2 - $0
518 Mistral Large (Feb '24) 4.1 - $6
519 Qwen2.5 Coder Instruct 7B 4.1 - $0
520 Mixtral 8x22B Instruct 4 - $0
521 Llama 3.2 Instruct 3B 3.9 - $0
522 Llama 2 Chat 7B 3.9 - $0.1
523 MiniCPM-V 4.6 1.3B 3.8 0.7 $0
524 Jamba Reasoning 3B 3.8 - $0
525 Reka Flash 3 3.7 - $0.35
526 Qwen1.5 Chat 110B 3.7 - $0
527 Qwen3 VL 4B Instruct 3.7 - $0
528 Olmo 3 7B Think 3.6 - $0
529 Claude 3 Haiku 3.5 - $0.5
530 Claude 2.1 3.5 14 $0
531 OLMo 2 7B 3.5 - $0
532 Molmo 7B-D 3.4 - $0
533 Ling-mini-2.0 3.4 - $0
534 Claude 2.0 3.3 12.9 $0
535 DeepSeek R1 Distill Qwen 1.5B 3.3 - $0
536 DeepSeek-V2-Chat 3.3 - $0
537 GPT-3.5 Turbo 3.2 10.7 $0.75
538 Mistral Small (Feb '24) 3.2 - $0.262
539 Mistral Medium 3.2 - $3
540 Llama 3 Instruct 70B 3.1 - $1.175
541 Llama 3.2 Instruct 11B (Vision) 3 - $0.345
542 LFM 40B 3 - $0
543 Arctic Instruct 3 - $0
544 Qwen Chat 72B 3 - $0
545 Qwen3.5 0.8B (Non-reasoning) 2.9 1.2 $0
546 PALM-2 2.8 4.6 $0
547 Gemini 1.0 Pro 2.7 - $0
548 DeepSeek Coder V2 Lite Instruct 2.7 - $0
549 Llama 2 Chat 13B 2.6 - $0
550 Llama 2 Chat 70B 2.6 - $0
551 DeepSeek LLM 67B Chat (V1) 2.6 - $0
552 OpenChat 3.5 (1210) 2.6 - $0
553 DBRX Instruct 2.6 - $0
554 Sarvam M (Reasoning) 2.6 - $0
555 Command-R+ (Apr '24) 2.6 - $6
556 Exaone 4.0 1.2B (Reasoning) 2.5 - $0
557 Olmo 3 7B Instruct 2.4 - $0.125
558 Exaone 4.0 1.2B (Non-reasoning) 2.4 - $0
559 LFM2.5-1.2B-Thinking 2.3 - $0
560 LFM2 2.6B 2.3 - $0
561 LFM2.5-1.2B-Instruct 2.3 - $0
562 Jamba 1.7 Mini 2.3 - $0
563 Jamba 1.5 Mini 2.3 - $0.25
564 Granite 4.0 H 1B 2.2 - $0
565 Qwen3 1.7B (Reasoning) 2.2 - $0
566 Jamba 1.6 Mini 2.1 - $0
567 Gemma 3 270M 2 - $0
568 Granite 4.0 Micro 2 - $0
569 Apertus 70B Instruct 2 - $1.345
570 Mixtral 8x7B Instruct 2 - $0.512
571 DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning) 1.9 - $0
572 Llama 65B 1.7 - $0
573 Qwen Chat 14B 1.7 - $0
574 Claude Instant 1.7 7.8 $0
575 Mistral 7B Instruct 1.7 - $0.25
576 Command-R (Mar '24) 1.7 - $0.75
577 Molmo2-8B 1.6 - $0
578 Granite 4.0 1B 1.6 - $0
579 LFM2 8B A1B 1.3 - $0
580 Granite 3.3 8B (Non-reasoning) 1.3 - $0.085
581 Qwen3 1.7B (Non-reasoning) 1.1 - $0
582 LFM2.5-VL-1.6B 1 - $0
583 Granite 4.0 H 350M 1 - $0
584 Granite 4.0 350M 1 - $0
585 Apertus 8B Instruct 1 - $0.125
586 Tiny Aya Global 1 - $0
587 Llama 3 Instruct 8B 1 - $0.07
588 Llama 3.2 Instruct 1B 1 - $0
589 Gemma 3n E4B Instruct 1 3.2 $0.075
590 Gemma 3 4B Instruct 1 2.7 $0
591 Gemma 3 1B Instruct 1 - $0
592 Gemma 3n E2B Instruct 1 - $0
593 LFM2 1.2B 1 - $0
594 Qwen3 0.6B (Reasoning) 1 - $0
595 Qwen3 0.6B (Non-reasoning) 1 - $0
596 GPT-5.5 Pro (xhigh) - - $0
597 Gemini 3 Deep Think - - $0
598 Claude Sonnet 5 (Adaptive Reasoning, Low Effort) - - $4
599 Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort) - - $4
600 Claude Sonnet 5 (Adaptive Reasoning, Medium Effort) - - $4
601 Claude Sonnet 5 (Adaptive Reasoning, High Effort) - - $4
602 EXAONE 4.5 33B (Non-reasoning) - - $0
603 Cogito v2.1 (Reasoning) - - $1.25
604 GPT-4o mini Realtime (Dec '24) - - $0
605 GPT-5.4 Pro (xhigh) - - $67.5
606 GPT-3.5 Turbo (0613) - - $0
607 GPT-4o Realtime (Dec '24) - - $0
608 Mi:dm K 2.5 Pro Preview - - $0

榜单解读建议

参考 AI 大模型排行榜 时,应综合考虑“综合指数”与“成本价格”。如果您是开发者,编程能力 (Coding) 是更核心的指标。

值品工具箱同步的 AI 大模型排行榜 数据每 24 小时更新,确保您获取到最新的模型性能对比。

指标说明

  • 综合指数:评估通用理解与逻辑。
  • 价格 $/1M:混合 3:1 输入输出比的平均成本。
  • 编程能力:衡量代码生成的准确性。

AI 大模型排行榜 常见问题 (FAQ)

Q1: AI 大模型排行榜 的数据多久更新?

AI 大模型排行榜 数据每 24 小时自动抓取一次,确保最新模型加入列表。

Q2: 这个 AI 大模型排行榜 包含国产模型吗?

是的,只要国产模型通过了 Artificial Analysis 的全球测评,就会出现在 AI 大模型排行榜 中。

Q3: 综合指数在 AI 大模型排行榜 中代表什么?

它代表模型的全能表现。AI 大模型排行榜 通过加权算法给出这个综合评分。

Q4: 如何在 AI 大模型排行榜 中查找性价比最高的游戏?

在 AI 大模型排行榜 页面中,您可以点击“价格”标题进行排序,寻找低价高分的模型。

Q5: AI 大模型排行榜 的编程能力测试准吗?

AI 大模型排行榜 参考了 LiveCodeBench 等权威基准测试,具有极高的参考价值。

Q6: 为什么有的新模型没进入 AI 大模型排行榜?

模型进入 AI 大模型排行榜 需要经过一系列测试,通常在新模型发布后数日内会完成更新。

Q7: AI 大模型排行榜 中的价格计算标准是什么?

价格是基于百万 Token 的调用成本,由 AI 大模型排行榜 统一混合计算得出。

Q8: 手机上能查看 AI 大模型排行榜 吗?

当然可以。AI 大模型排行榜 进行了移动端响应式深度优化。

Q9: AI 大模型排行榜 这个工具免费吗?

是的,由值品工具箱免费提供 AI 大模型排行榜 信息查询服务。

Q10: 我该怎么利用 AI 大模型排行榜 做选型?

如果您需要智能客服,参考 AI 大模型排行榜 的综合指数;如果做翻译,参考编程外的语言指标。

发表评论

请友善文明留言