AI 大模型排行榜 (Artificial Analysis LLM Ranking)

信息查询
4.1k 次浏览
100% 有帮助 · 1 人反馈

值品工具箱提供的 AI 大模型排行榜聚合了来自 Artificial Analysis 的权威数据,实时追踪并排名超过 100 个主流大语言模型。

AI 大模型排行榜数据中心

重置
排名 模型名称 综合指数 ▼ 编程 价格 ($/1M)
1 Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) 57.6 - $8
2 Claude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) 56 - $8
3 Claude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback) 53.6 - $8
4 Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) 53.4 81.6 $20
5 Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) 53.2 80.7 $20
6 GPT-6 Astra (max) 52.7 76.9 $20
7 GPT-6 Astra (xhigh) 52.4 75.9 $20
8 Claude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) 51.2 - $8
9 Claude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback) 51.2 79.1 $20
10 GPT-6 Astra (high) 50.9 77.1 $20
11 Claude Opus 5 (Adaptive Reasoning, Max Effort) 50.8 78 $10
12 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) 49.7 77 $10
13 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) 49.6 76.5 $20
14 GPT-6 Astra (medium) 49.6 76.7 $20
15 Claude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback) 48.9 77.1 $20
16 Claude Opus 5 (Adaptive Reasoning, High Effort) 48.1 76.5 $10
17 Muse Spark 1.3 (max) 48.1 75.8 $2
18 GPT-6 Sol (max) 47.5 - $4
19 GPT-5.6 Sol (max) 47 77.4 $8
20 Claude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback) 46.8 75.2 $20
21 Grok 4.7 (xhigh) 46.4 - $3
22 Grok 4.7 (high) 46.3 - $3
23 MiMo-V2.6-Pro 46.3 - $0.544
24 GPT-6 Astra (low) 45.8 75.7 $20
25 Qwen3.8 Max (0902) 45.4 76.2 $3
26 Muse Spark 1.3 (xhigh) 45.1 76.5 $2
27 Claude Opus 5 (Adaptive Reasoning, Medium Effort) 44.8 74.3 $10
28 GLM-5.3 (max) 44.8 74.8 $2.15
29 Grok 4.6 (high) 44.3 76.8 $3
30 Grok 4.6 (xhigh) 44.2 75.9 $3
31 GPT-6 Sol (xhigh) 44.1 - $4
32 GPT-5.6 Sol (xhigh) 44 78.3 $8
33 Step 5 Preview 43.7 - $1.425
34 Kimi K3 (max) 43.6 76.2 $6
35 Grok 4.6 (medium) 42.8 74.4 $3
36 GPT-6 Sol (high) 42.8 - $4
37 GPT-5.6 Sol (high) 42.3 77.2 $8
38 Claude Opus 5.5 (Adaptive Reasoning, Low Effort, Default Fallback) 42.3 - $8
39 GPT-5.6 Terra (max) 42.1 76.7 $4.5
40 GLM 5.3 Flash 41.8 71.5 $0.237
41 Claude Opus 4.8 (Adaptive Reasoning, Max Effort) 41.8 74.3 $10
42 Gemini 3.8 Flash (high) 40.9 76.3 $1.5
43 Claude Opus 4.7 (Adaptive Reasoning, Max Effort) 40.7 73.6 $10
44 Qwen3.8 Max 40.2 71.8 $3
45 Qwen3.8 2.4T A95B 39.9 71.9 $3
46 Qwen3.8-Flash-Next 39.8 73.1 $0.23
47 GPT-6 Sol (medium) 39.8 - $4
48 Gemini 3.8 Flash (medium) 39.8 74.1 $1.5
49 Gemini 3.7 Flash (medium) 39.6 71.5 $1.5
50 Muse Spark 1.2 (xhigh) 39.6 72.2 $2
51 DeepSeek V4.1 Flash (Reasoning, Max Effort) 39.5 - $0.525
52 Claude Opus 5 (Adaptive Reasoning, Low Effort) 39.4 66.9 $10
53 GPT-5.6 Sol (medium) 39.2 76.3 $8
54 Gemini 3.7 Flash (high) 39.1 76.1 $1.5
55 GPT-5.4 (xhigh) 39 71.1 $5.625
56 Grok 4.5 (high) 38.8 72.4 $3
57 GPT-5.5 (xhigh) 38.4 74.9 $11.25
58 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) 38.2 71.5 $4
59 GPT-5.6 Terra (xhigh) 38 70.6 $4.5
60 MiMo-V2.6-Flash 37.9 - $0.175
61 GPT-5.6 Luna (max) 37.3 71.4 $0.45
62 GPT-6 Luna (max) 37.3 - $0.2
63 GPT-5.5 (high) 37 71.6 $11.25
64 Gemini 3.7 Flash (low) 36.9 71 $1.5
65 DeepSeek V4 Pro 0813 (Reasoning, Max Effort) 36 68.8 $1.98
66 Grok 4.6 (low) 35.1 66.3 $3
67 DeepSeek V4 Flash Vision (Reasoning, Max Effort) 34.8 65 $0.66
68 GPT-5.6 Luna (xhigh) 34.6 68.6 $0.45
69 Kimi K3 (low) 34.5 72 $6
70 Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort) 34.4 - $4
71 DeepSeek V4 Flash 0731 (Reasoning, Max Effort) 34.3 69.1 $0.66
72 GLM-5.3 (low) 34.3 - $2.15
73 GPT-5.6 Terra (high) 34.2 67.1 $4.5
74 Gemini 3.6 Flash (high) 34 69.2 $1.5
75 GPT-6 Sol (low) 33.9 - $4
76 GPT-6 Luna (xhigh) 33.9 - $0.2
77 GPT-5.5 (medium) 33.8 71.5 $11.25
78 Muse Spark 1.1 (xhigh) 33.7 71.3 $2
79 GLM-5.2 (max) 33.7 68.8 $2.15
80 Qwen3.8 27B (xhigh) 33.7 68.1 $1.125
81 Gemini 3.5 Flash (medium) 33.6 - $3.375
82 Motif 3 33.6 63.5 $0
83 GPT-5.6 Sol (low) 33.5 69.7 $8
84 Gemini 3.8 Flash (low) 33.5 73.5 $1.5
85 Gemini 3.5 Flash (high) 32.6 70.1 $3.375
86 GPT-5.3 Codex (xhigh) 32.5 - $4.813
87 Motif 3 (Beta) 32.3 62 $0
88 GPT-6 Luna (high) 32.1 - $0.2
89 GPT-5.6 Luna (high) 32.1 63.3 $0.45
90 Claude Opus 4.6 (Adaptive Reasoning, Max Effort) 31.9 - $10
91 Claude Sonnet 5 (Adaptive Reasoning, High Effort) 31.7 - $4
92 Muse Spark 31.3 58.6 $0
93 Claude Opus 4.7 (Non-reasoning, High Effort) 30.9 - $10
94 GPT-5.5 (low) 30.7 60.9 $11.25
95 K2 Horizon 375B A23B 30.5 61.5 $0
96 DeepSeek V4 Pro 0424 (Reasoning, Max Effort) 30.4 59.4 $0.544
97 GPT-5.2 (xhigh) 30.4 - $4.813
98 DeepSeek V4 Pro 0424 (Reasoning, High Effort) 30.1 58.7 $0.544
99 GPT-5.6 Terra (medium) 30.1 64.7 $4.5
100 Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) 30.1 63 $6
101 Gemini 3.1 Pro Preview 29.7 68.8 $4.5
102 GPT-6 Luna (medium) 29.5 - $0.2
103 Qwen3.7 Max 29.5 66 $3.75
104 MiniMax-M3 29.2 58.6 $0.525
105 Claude Opus 4.5 (Reasoning) 29.1 - $10
106 MiMo-V2-Pro 28.6 - $0
107 GPT-5.2 Codex (xhigh) 28.5 - $4.813
108 Qwen3.6 Max Preview 28.4 - $2.925
109 GPT-5.6 Sol (Non-reasoning) 28.3 65.1 $8
110 Nex-N2-Pro (based on Qwen3.5-397B-A17B) 28.2 59.1 $0
111 Solar Pro 4 28.2 52.7 $0.525
112 GPT-6 Sol (Non-reasoning) 28.1 - $4
113 Claude Sonnet 5 (Adaptive Reasoning, Medium Effort) 28.1 - $4
114 Gemini 3 Pro Preview (high) 28 - $4.5
115 GLM-5 (Reasoning) 27.9 - $1.55
116 Inkling Small 27.8 52.9 $0.525
117 GPT-5.4 (low) 27.6 - $5.625
118 Qwen3.8 27B (medium) 27.6 56.1 $1.125
119 GPT-5.6 Terra (low) 27.5 58.1 $4.5
120 JT-4.1 Flash 236B A21B 27.3 52.4 $0
121 Grok Build 0.1 0616 27.2 51.5 $1.25
122 Qwen3.6 Plus 27 54.5 $1.125
123 Kimi K2.6 27 61.8 $1.712
124 Quasar 438B (max, based on GLM-5.2) 26.7 61.2 $0.9
125 GLM-5-Turbo 26.6 - $0
126 GPT-5.2 (medium) 26.5 - $4.813
127 Apodex 1.1 26.4 60.8 $0.975
128 Claude Opus 4.6 (Non-reasoning, High Effort) 26.4 - $10
129 Gemini 3 Flash Preview (Reasoning) 26.3 - $1.125
130 Qwen3.8 27B (low) 26.2 58.2 $1.125
131 GLM-5.1 (Reasoning) 26.1 55.8 $1.981
132 GPT-5.5 Instant (June 2026) 26 39.4 $11.25
133 DeepSeek V4 Flash 0420 (Reasoning, High Effort) 26 52 $0.168
134 MiMo-V2.5-Pro 26 60.2 $0.544
135 Kimi K2.7 Code 25.8 60.8 $1.712
136 Grok 4.20 0309 v2 (Reasoning) 25.7 - $1.563
137 K2 Horizon MoVA 36B A4B 25.3 - $0
138 Hy3 25.3 58.8 $0.25
139 Grok 4.20 0309 (Reasoning) 25.2 - $3
140 MiMo-V2.5 25.2 56.8 $0.175
141 Qwen3.7 Plus 25.2 55.9 $0.7
142 MiMo-V2-Omni-0327 25.1 - $0
143 GPT-5.6 Luna (medium) 25 50.7 $0.45
144 Inkling (xhigh) 25 52.1 $1.762
145 Ling 3.0 Flash 24.9 50.6 $0.111
146 GPT-5 Codex (high) 24.9 - $3.438
147 Grok 4.3 (high) 24.9 42.2 $1.563
148 Grok 4.3 (medium) 24.8 - $1.563
149 Solar Open2 250B 24.7 45 $0
150 GPT-5.1 (high) 24.7 49.4 $3.438
151 Claude Sonnet 4.6 (Non-reasoning, High Effort) 24.7 - $6
152 DeepSeek V4.1 Flash (Non-Reasoning) 24.7 - $0.525
153 Ling-3.0-flash-VL 24.6 57 $0.111
154 Grok 4.3 (low) 24.3 - $1.563
155 Claude Sonnet 5 (Adaptive Reasoning, Low Effort) 24.3 - $4
156 GLM-5.1 (Non-reasoning) 24.2 - $2.135
157 DeepSeek V4 Flash 0420 (Reasoning, Max Effort) 24.2 56.2 $0.168
158 GPT-5.4 mini (xhigh) 24.1 56.1 $1.688
159 MiMo-V2-Omni 23.9 - $0
160 Gemini 3.5 Flash (minimal) 23.8 - $3.375
161 GPT-5.1 Codex (high) 23.7 - $3.438
162 Claude Opus 4.5 (Non-reasoning) 23.7 - $10
163 Kimi K2.6 (Non-reasoning) 23.6 - $1.712
164 GLM 5V Turbo (Reasoning) 23.5 - $0
165 Kimi K2.5 (Reasoning) 23.5 46.8 $1.137
166 Claude Sonnet 4.6 (Non-reasoning, Low Effort) 23.3 - $6
167 Claude Sonnet 5 (Non-reasoning, High Effort) 23.2 66.4 $4
168 GPT-5.5 (Non-reasoning) 23.2 56.5 $11.25
169 GPT-5 (high) 23 37.8 $3.438
170 Nemotron 3 Ultra 550B A55B (Reasoning) 22.9 49.3 $1.075
171 Qwen3.5 27B (Reasoning) 22.9 - $0.825
172 GPT-5 (medium) 22.9 - $3.438
173 Claude 4.1 Opus (Reasoning) 22.8 - $30
174 MiniMax-M2.5 22.8 - $0.525
175 MiniMax-M2.7 22.8 52.6 $0.525
176 Hy3-preview (Reasoning) 22.7 - $0.1
177 A.X-K2 22.7 38.8 $0
178 GPT-5.5 Instant (May 2026) 22.7 - $11.25
179 Ling-3.0-flash-Fin 22.6 55.6 $0.111
180 Grok 4 22.5 - $6
181 MiMo-V2-Flash (Feb 2026) 22.4 - $0
182 GLM-5.2 (Non-reasoning) 22.4 46.5 $2.15
183 Gemini 3 Pro Preview (low) 22.3 - $4.5
184 GLM-4.7 (Reasoning) 22.2 45.3 $1
185 Gemini 3.5 Flash-Lite 22.2 49.3 $0.85
186 Kimi K2 Thinking 22 - $1.075
187 o3-pro 21.9 - $35
188 G9v3-39A5B 21.8 33.1 $0
189 GLM-5 (Non-reasoning) 21.8 - $1.55
190 KAT Coder Pro V2 21.7 59.5 $0.525
191 DeepSeek V3.2 (Reasoning) 21.5 44.2 $0.315
192 Qwen3.5 397B A17B (Non-reasoning) 21.4 - $1.35
193 Qwen3.6 27B (Reasoning) 21.4 53.7 $1.35
194 Qwen3 Max Thinking 21.3 - $0
195 GPT-5.6 Luna (low) 21 44.2 $0.45
196 MiniMax-M2.1 20.9 - $0.525
197 GPT-6 Luna (low) 20.9 - $0.2
198 DeepSeek V4 Pro 0424 (Non-reasoning) 20.8 - $0.544
199 MiMo-V2-Flash (Reasoning) 20.8 - $0.15
200 GPT-5 (low) 20.8 - $3.438
201 GPT-5.6 Terra (Non-reasoning) 20.8 52.3 $4.5
202 GPT-5.4 nano (xhigh) 20.7 56.1 $0.463
203 Claude 4.5 Sonnet (Reasoning) 20.7 52.1 $6
204 Claude 4 Opus (Reasoning) 20.6 - $30
205 GPT-5 mini (medium) 20.6 - $0.688
206 K2 Horizon 7B 20.6 38.6 $0
207 Qwen3.5 Omni Plus 20.4 - $1.5
208 GPT-5.1 Codex mini (high) 20.4 - $0.688
209 Grok 4.1 Fast (Reasoning) 20.4 - $0
210 o3 20.2 - $3.5
211 Qwen3.8 27B (Non-reasoning) 20.2 44.6 $1.125
212 GPT-5.4 nano (medium) 20 - $0.463
213 Qwen3.6 27B (Non-reasoning) 19.8 46.6 $1.35
214 GPT-5.4 mini (medium) 19.7 - $1.688
215 K-EXAONE 2.0 0803 19.7 40.6 $0
216 Step 3.7 Flash 19.5 39.6 $0.438
217 Kimi K2.5 (Non-reasoning) 19.4 - $1.2
218 Qwen3.5 27B (Non-reasoning) 19.4 - $0.825
219 Claude 4.5 Sonnet (Non-reasoning) 19.3 - $6
220 Qwen3.5 35B A3B (Reasoning) 19.3 - $0.688
221 LongCat 2.0 19.1 45.3 $0.525
222 Gemma 4 31B (Reasoning) 19 43.4 $0
223 Claude 4 Sonnet (Reasoning) 18.9 37.6 $0
224 DeepSeek V4 Flash 0420 (Non-reasoning) 18.9 - $0.119
225 JT-35B-Flash 18.7 - $0
226 MiniMax-M2 18.6 - $0.525
227 KAT-Coder-Pro V1 18.6 - $0
228 Claude 4.1 Opus (Non-reasoning) 18.6 - $30
229 GLM-4.6 (Reasoning) 18.5 45.8 $0.963
230 Qwen3.5 397B A17B (Reasoning) 18.4 48.2 $1.35
231 MiMo-V2.5-Pro (Non-reasoning) 18.3 - $0.544
232 GPT-6 Luna (Non-reasoning) 18.3 - $0.2
233 Qwen3.6 35B A3B (Reasoning) 18.2 41.9 $0.844
234 GPT-5.4 (Non-reasoning) 18.2 - $5.625
235 Grok 4 Fast (Reasoning) 17.9 - $0.275
236 Gemini 3 Flash Preview (Non-reasoning) 17.9 - $1.125
237 Qwen3.5 122B A10B (Non-reasoning) 17.7 43.3 $1.1
238 Claude 3.7 Sonnet (Reasoning) 17.7 36.4 $0
239 Muse Glimmer (high) 17.5 49 $0.581
240 GLM-4.7 (Non-reasoning) 17.4 - $1
241 Hy3-preview (Non-reasoning) 17 - $0.1
242 Ling-2.6-1T 17 - $0.85
243 GPT-5.2 (Non-reasoning) 17 - $4.813
244 Step 3.5 Flash 2603 17 - $0.15
245 Doubao Seed Code 16.9 - $0
246 Claude 4.5 Haiku (Reasoning) 16.9 43.9 $2
247 GPT-5 mini (high) 16.8 15.6 $0.688
248 Gemma 4 26B A4B (Reasoning) 16.7 39.3 $0.168
249 o4-mini (high) 16.7 - $1.925
250 Ring-2.6-1T 16.6 42.8 $0.85
251 Step 3.5 Flash 16.6 - $0.15
252 Claude 4 Opus (Non-reasoning) 16.6 - $30
253 Claude 4 Sonnet (Non-reasoning) 16.6 - $0
254 DeepSeek V3.2 Exp (Reasoning) 16.6 - $0.315
255 Qwen3 Max Thinking (Preview) 16.3 - $2.4
256 Gemini 2.5 Pro 16.1 33.3 $3.438
257 MiMo-V2-Flash (Non-reasoning) 16 49.8 $0
258 DeepSeek V3.2 (Non-reasoning) 16 - $0.315
259 K2 Horizon 3.7B 15.6 26.1 $0
260 Qwen3 Max 15.6 - $2.4
261 Qwen3.5 122B A10B (Reasoning) 15.6 45.7 $1.1
262 Gemini 3.1 Flash-Lite 15.6 34.7 $0.563
263 GPT-5.6 Luna (Non-reasoning) 15.5 39.3 $0.45
264 Gemini 2.5 Flash Preview (Sep '25) (Reasoning) 15.5 - $0
265 Claude 4.5 Haiku (Non-reasoning) 15.4 - $2
266 Kimi K2 0905 15.3 - $1.075
267 Ling 3.0 Tiny 15.3 26.5 $0
268 Claude 3.7 Sonnet (Non-reasoning) 15.3 - $6
269 Qwen3.6 35B A3B (Non-reasoning) 15.2 28.1 $0.844
270 o1 15.2 39.7 $26.25
271 Qwen3.5 35B A3B (Non-reasoning) 15.1 37 $0.688
272 Gemini 2.5 Pro Preview (Mar' 25) 15 46.7 $0
273 GLM-4.6 (Non-reasoning) 14.9 - $0.981
274 GLM-4.7-Flash (Reasoning) 14.9 - $0.153
275 Granite 4.2 30B 14.8 29.9 $0.282
276 DeepSeek V3.1 Terminus (Reasoning) 14.8 43.5 $1.914
277 Grok 3 mini Reasoning (high) 14.6 - $0.35
278 Grok 4.20 0309 (Non-reasoning) 14.6 - $3
279 Gemini 2.5 Pro Preview (May' 25) 14.5 - $3.438
280 DeepSeek V3.2 Speciale 14.5 - $0
281 K-EXAONE (Reasoning) 14.4 32.1 $0
282 ERNIE 5.0 Thinking Preview 14.3 - $0
283 Grok 4.20 0309 v2 (Non-reasoning) 14.2 - $1.563
284 Mistral Medium 3.5 14.2 46.9 $3
285 Gemma 4 12B (Reasoning) 14.2 31 $0.15
286 Nova 2.0 Pro Preview (medium) 14.2 34 $3.438
287 Grok Code Fast 1 14.1 - $0
288 Grok 4.3 (Non-reasoning) 14 35.2 $1.563
289 DeepSeek V3.1 Terminus (Non-reasoning) 13.9 - $0.453
290 Gemma 4 31B (Non-reasoning) 13.9 33.2 $0.205
291 DeepSeek V3.2 Exp (Non-reasoning) 13.9 - $0.315
292 Apriel-v1.5-15B-Thinker 13.8 - $0
293 Mercury 2 13.8 31.1 $0.375
294 DeepSeek V3.1 (Non-reasoning) 13.7 - $0.848
295 Qwen3.5 9B (Reasoning) 13.7 28.7 $0.151
296 Nova 2.0 Omni (medium) 13.6 - $0.85
297 DeepSeek V3.1 (Reasoning) 13.5 - $0.865
298 Qwen3 VL 235B A22B (Reasoning) 13.4 - $1.3
299 Apriel-v1.6-15B-Thinker 13.4 - $0
300 Nova 2.0 Lite (high) 13.4 23 $0.85
301 GPT-5.1 (Non-reasoning) 13.3 - $3.438
302 Qwen3.5 9B (Non-reasoning) 13.3 23.5 $0.19
303 EXAONE 4.5 33B 13.2 23.6 $0
304 Command A+ 13.1 27.8 $0
305 Gemma 4 26B A4B (Non-reasoning) 13.1 - $0.198
306 Qwen3.5 4B (Reasoning) 13.1 22.6 $0.06
307 DeepSeek R1 0528 (May '25) 13.1 - $1.763
308 Gemini 2.5 Flash (Reasoning) 13.1 - $0.85
309 GPT-5 nano (high) 13 - $0.138
310 Nemotron 3.5 Lightning 12.9 26.8 $0.108
311 Nemotron 3 Super 120B A12B (Reasoning) 12.8 37.7 $0.45
312 Nova 2.0 Pro Preview (low) 12.8 25.9 $3.438
313 GLM-4.5 (Reasoning) 12.8 - $0
314 Kimi K2 12.7 - $1.002
315 Qwen3 235B A22B 2507 (Reasoning) 12.7 22.1 $0.747
316 GPT-4.1 12.7 - $3.5
317 Qwen3 Max (Preview) 12.6 - $2.4
318 Nova 2.0 Lite (medium) 12.5 - $0.85
319 GPT-5 nano (medium) 12.5 - $0.138
320 Qwen3.5 Omni Flash 12.5 - $0.275
321 o3-mini 12.5 - $1.925
322 MiniCPM5-2B 12.5 14.5 $0
323 o1-pro 12.4 - $262.5
324 Gemini 2.5 Flash Preview (Sep '25) (Non-reasoning) 12.4 - $0
325 Mercury 2.5 12.3 - $0.375
326 JT-MINI 12.2 - $0
327 Grok 3 12.1 - $8
328 Seed-OSS-36B-Instruct 12.1 - $0.3
329 Qwen3 235B A22B 2507 Instruct 12 - $0.403
330 Qwen3 Coder 480B A35B Instruct 11.9 - $3
331 Qwen3 VL 32B (Reasoning) 11.9 - $0.28
332 Magistral Medium 1.2 11.8 21.3 $0
333 Sonar Reasoning Pro 11.8 - $0
334 Nova 2.0 Lite (low) 11.8 - $0.85
335 HyperNova 60B 2605 (high, based on gpt-oss-120b) 11.7 23.2 $0
336 MiniMax M1 80k 11.7 - $0.963
337 GPT-5.4 nano (Non-Reasoning) 11.7 - $0.463
338 Nemotron Cascade 2 30B A3B 11.7 25.3 $0
339 Gemini 2.5 Flash Preview (Reasoning) 11.7 - $0
340 gpt-oss-120b (high) 11.6 30.4 $0.261
341 K2 Think V2 11.5 21 $0
342 LongCat Flash Lite 11.5 - $0
343 GPT-5 (minimal) 11.4 - $3.438
344 DeepSeek R1 (Jan '25) 11.4 24.6 $2.5
345 o1-preview 11.4 34 $28.875
346 HyperCLOVA X SEED Think (32B) 11.4 - $0
347 Grok 4.1 Fast (Non-reasoning) 11.3 - $0
348 Mistral Small 4 (Reasoning) 11.3 26.6 $0.262
349 GLM-4.6V (Reasoning) 11.2 - $0.45
350 K-EXAONE (Non-reasoning) 11.2 - $0
351 Qwen3 Next 80B A3B (Reasoning) 11.2 17.4 $0.412
352 GPT-5.4 mini (Non-Reasoning) 11.1 - $1.688
353 Nova 2.0 Omni (low) 11.1 - $0.85
354 GLM-4.5-Air 11.1 - $0.372
355 Granite 4.2 8B 11.1 22.4 $0.107
356 Grok 4 Fast (Non-reasoning) 11.1 - $0.275
357 Mi:dm K 2.5 Pro 11 - $0
358 o3-mini (high) 11 16.3 $1.925
359 Ring-1T 10.9 - $0
360 G9v3-3B 10.8 9.9 $0
361 Trinity Large Thinking 10.8 25.8 $0.388
362 Qwen3.5 4B (Non-reasoning) 10.8 20.3 $0.06
363 INTELLECT-3 (based on GLM-4.5-Air) 10.6 - $0
364 GLM-4.7-Flash (Non-reasoning) 10.6 - $0.153
365 GPT-5 (ChatGPT) 10.4 - $0
366 Solar Open 100B (Reasoning) 10.4 - $0
367 Grok 3 Reasoning Beta 10.4 - $0
368 Gemini 2.5 Flash-Lite Preview (Sep '25) (Reasoning) 10.4 - $0.175
369 Nemotron 3 Nano Omni 30B A3B Reasoning 10.3 13.8 $0.42
370 gpt-oss-120b (low) 10.2 21.2 $0.249
371 GPT-4.1 mini 10.2 20.2 $0.7
372 Llama 4 Maverick 10 16.3 $0.422
373 MiniMax M1 40k 10 - $0
374 Nova 2.0 Pro Preview (Non-reasoning) 10 20.9 $3.438
375 gpt-oss-20b (low) 10 - $0.106
376 Qwen3 VL 235B A22B Instruct 9.9 - $0.7
377 North Mini Code 9.9 36.5 $0
378 GPT-5 mini (minimal) 9.9 - $0.688
379 K2-V2 (high) 9.9 - $0
380 Gemini 2.5 Flash (Non-reasoning) 9.9 - $0.85
381 Qwen3 30B A3B 2507 (Reasoning) 9.8 12.1 $0.75
382 o1-mini 9.8 - $0
383 DeepSeek V3 0324 9.7 21.2 $0.927
384 Ling 2.6 Flash 9.7 25.3 $0
385 Qwen3 Next 80B A3B Instruct 9.6 - $0.412
386 Tri-21B-think Preview 9.6 - $0
387 Qwen3 Coder 30B A3B Instruct 9.6 - $0.9
388 GPT-4.5 (Preview) 9.6 - $0
389 DiffusionGemma 26B A4B 9.5 19.7 $0
390 Qwen3 235B A22B (Reasoning) 9.5 - $2.625
391 QwQ 32B 9.5 - $0.745
392 Qwen3 VL 30B A3B (Reasoning) 9.5 - $0.75
393 Gemini 2.0 Flash Thinking Experimental (Jan '25) 9.4 24.1 $0
394 Gemma 4 12B (Non-reasoning) 9.4 - $0.15
395 Gemini 2.5 Flash-Lite Preview (Sep '25) (Non-reasoning) 9.3 - $0.175
396 Mistral Large 3 9.3 20.1 $0.75
397 Qwen3 Coder Next 9.2 36.2 $0.563
398 Motif-2-12.7B-Reasoning 9.2 - $0
399 Mistral Medium 3.1 9.2 20.5 $0
400 Ling-1T 9.2 - $0
401 Nova Premier 9.2 - $0
402 Solar Pro 2 (Preview) (Reasoning) 9.1 - $0
403 Granite 4.2 3B 9.1 17.5 $0.052
404 Magistral Medium 1 9.1 - $0
405 Mistral Medium 3 9 - $0.8
406 K2-V2 (medium) 9 - $0
407 Llama Nemotron Super 49B v1.5 (Reasoning) 9 - $0
408 Devstral Medium 9 - $0
409 Mistral Small 4 (Non-reasoning) 9 - $0.262
410 Tri-21B-Think 9 - $0
411 gpt-oss-20b (high) 9 20.7 $0.098
412 GPT-4o (March 2025, chatgpt-4o-latest) 9 - $0
413 Gemini 2.0 Flash (Feb '25) 8.9 - $0
414 Claude 3.5 Haiku 8.9 15.9 $0
415 Llama 3.3 Nemotron Super 49B v1 (Reasoning) 8.9 - $0
416 Gemma 4 E4B (Reasoning) 8.9 9.4 $0.04
417 NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) 8.9 14.4 $0.088
418 Qwen3 4B 2507 (Reasoning) 8.8 - $0
419 MiniCPM5-1B (Reasoning) 8.8 - $0
420 Sarvam 105B (high) 8.8 - $0.074
421 Gemini 2.0 Pro Experimental (Feb '25) 8.7 25.5 $0
422 Nova 2.0 Lite (Non-reasoning) 8.7 - $0.85
423 Devstral Small (May '25) 8.7 - $0
424 Claude 3 Opus 8.7 19.5 $30
425 MiniCPM5-1B (Non-reasoning) 8.7 - $0
426 Sonar Reasoning 8.7 - $0
427 Gemini 2.5 Flash Preview (Non-reasoning) 8.7 - $0
428 Devstral 2 8.6 31.3 $0
429 Magistral Small 1.2 8.6 14.7 $0.75
430 Qwen3 32B (Reasoning) 8.6 15.3 $0.28
431 Gemini 2.5 Flash-Lite (Reasoning) 8.5 - $0.175
432 DeepSeek V3 (Dec '24) 8.5 23 $0.463
433 GPT-4o (Nov '24) 8.4 - $4.375
434 Nanbeige4.1-3B 8.4 9.6 $0
435 LFM2.5-2.6B 8.4 7.7 $0
436 Qwen3 VL 32B Instruct 8.4 - $0.28
437 DeepSeek R1 Distill Qwen 32B 8.4 - $0
438 GLM-4.6V (Non-reasoning) 8.4 - $0.45
439 Qwen3 235B A22B (Non-reasoning) 8.3 - $1.225
440 Mistral Small 3.2 8.2 12.5 $0.106
441 Magistral Small 1 8.2 - $0
442 Gemini 2.0 Flash (experimental) 8.2 - $0
443 EXAONE 4.0 32B (Reasoning) 8.2 - $0
444 Qwen3 VL 8B (Reasoning) 8.2 - $0.66
445 Qwen3 14B (Reasoning) 8.2 13.8 $1.313
446 Nova 2.0 Omni (Non-reasoning) 8.2 - $0.85
447 DeepSeek R1 0528 Qwen3 8B 8.1 - $0
448 Llama 4 Scout 8.1 8.2 $0.313
449 Qwen2.5 Max 8 - $0
450 Qwen3 VL 30B A3B Instruct 7.9 - $0.35
451 Hermes 4 - Llama-3.1 70B (Reasoning) 7.9 - $0
452 Gemini 1.5 Pro (Sep '24) 7.9 23.6 $0
453 Solar Pro 2 (Preview) (Non-reasoning) 7.9 - $0
454 DeepSeek R1 Distill Llama 70B 7.9 - $0.8
455 Claude 3.5 Sonnet (Oct '24) 7.9 30.2 $6
456 DeepSeek R1 Distill Qwen 14B 7.8 - $0
457 Falcon-H1R-7B 7.8 - $0
458 GPT-4.1 nano 7.8 11.1 $0.175
459 Solar Pro 3 7.8 16.2 $0.262
460 Ling-flash-2.0 7.8 - $0.247
461 Gemma 4 E2B (Reasoning) 7.8 7.2 $0
462 Qwen3 Omni 30B A3B (Reasoning) 7.8 - $0.43
463 GPT-4o (Aug '24) 7.7 - $4.375
464 Qwen2.5 Instruct 72B 7.7 - $0.48
465 Sonar 7.7 - $0
466 Step3 VL 10B 7.7 - $0
467 Llama 3.3 Instruct 70B 7.7 11.9 $0.712
468 Qwen3 30B A3B (Reasoning) 7.6 - $0.75
469 Sonar Pro 7.6 - $0
470 Devstral Small (Jul '25) 7.6 - $0
471 QwQ 32B-Preview 7.6 - $0
472 GLM-4.5V (Reasoning) 7.6 - $0.9
473 Mistral Large 2 (Nov '24) 7.6 - $0
474 Devstral Small 2 7.5 29.3 $0
475 Llama 3.1 Nemotron Ultra 253B v1 (Reasoning) 7.5 - $0
476 Qwen3 30B A3B 2507 Instruct 7.5 - $0.35
477 ERNIE 4.5 300B A47B 7.5 - $0.485
478 Hermes 4 - Llama-3.1 405B (Reasoning) 7.5 - $1.5
479 Solar Pro 2 (Reasoning) 7.5 - $0
480 NVIDIA Nemotron Nano 12B v2 VL (Reasoning) 7.5 - $0
481 Gemma 4 E4B (Non-reasoning) 7.5 - $0.04
482 Granite 4.1 30B 7.4 10.4 $0
483 NVIDIA Nemotron Nano 9B V2 (Reasoning) 7.4 - $0.07
484 Hermes 4 - Llama-3.1 405B (Non-reasoning) 7.4 - $1.5
485 Gemini 2.0 Flash-Lite (Feb '25) 7.4 - $0
486 NVIDIA Nemotron 3 Nano 4B 7.4 8 $0
487 Llama Nemotron Super 49B v1.5 (Non-reasoning) 7.4 - $0
488 Qwen3 32B (Non-reasoning) 7.3 - $0.28
489 GPT-4o (May '24) 7.3 24.2 $7.5
490 Gemini 2.0 Flash-Lite (Preview) 7.3 - $0
491 K2-V2 (low) 7.3 - $0
492 Llama 3.1 Nemotron Nano 4B v1.1 (Reasoning) 7.3 - $0
493 Kimi Linear 48B A3B Instruct 7.3 - $0
494 Llama 3.1 Instruct 405B 7.3 - $0
495 Qwen3 8B (Reasoning) 7.3 9 $0.66
496 Llama 3.3 Nemotron Super 49B v1 (Non-reasoning) 7.3 - $0
497 Qwen3 VL 8B Instruct 7.3 - $0.31
498 Qwen3 4B (Reasoning) 7.2 - $0
499 LFM2.5-8B-A1B 7.2 - $0
500 Claude 3.5 Sonnet (June '24) 7.2 26 $6
501 Llama 3.1 Tulu3 405B 7.2 - $0
502 GPT-4o (ChatGPT) 7.2 - $0
503 Ring-flash-2.0 7.2 - $0.247
504 Pixtral Large 7.1 - $0
505 Olmo 3.1 32B Think 7.1 - $0
506 Mistral Small 3.1 7.1 26.3 $0.125
507 Grok 2 (Dec '24) 7.1 - $0
508 GPT-5 nano (minimal) 7.1 - $0.138
509 Gemini 1.5 Flash (Sep '24) 7.1 - $0
510 Qwen3 VL 4B (Reasoning) 7 - $0
511 GPT-4 Turbo 7 21.5 $15
512 Solar Pro 2 (Non-reasoning) 7 - $0
513 Nova Pro 7 - $1.4
514 Command A 7 - $4.375
515 Qwen3.5 2B (Reasoning) 6.9 2.9 $0
516 Llama 3.1 Nemotron Instruct 70B 6.9 - $0
517 Llama 3.1 Instruct 8B 6.9 5.4 $0.028
518 Grok Beta 6.9 - $0
519 Qwen2.5 Instruct 32B 6.9 - $0
520 NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning) 6.8 - $0.088
521 NVIDIA Nemotron Nano 9B V2 (Non-reasoning) 6.8 - $0.086
522 Mistral Large 2 (Jul '24) 6.8 - $3
523 Qwen3 4B 2507 Instruct 6.7 - $0
524 Qwen2.5 Coder Instruct 32B 6.7 - $0
525 Qwen3 14B (Non-reasoning) 6.7 - $0.612
526 GPT-4 6.7 13.1 $37.5
527 GLM-4.5V (Non-reasoning) 6.7 - $0.9
528 Mistral Small 3 6.7 - $0.058
529 Gemini 2.5 Flash-Lite (Non-reasoning) 6.7 - $0.175
530 Nova Lite 6.7 - $0.105
531 GPT-4o mini 6.7 11.4 $0.262
532 Hermes 4 - Llama-3.1 70B (Non-reasoning) 6.7 - $0
533 Qwen3 30B A3B (Non-reasoning) 6.6 - $0.35
534 DeepSeek-V2.5 (Dec '24) 6.6 - $0
535 Qwen3 4B (Non-reasoning) 6.6 - $0
536 Llama 3.1 Instruct 70B 6.6 - $0.56
537 Granite 4.1 8B 6.6 9.5 $0.063
538 Sarvam 30B (high) 6.6 - $0.047
539 Gemini 2.0 Flash Thinking Experimental (Dec '24) 6.6 - $0
540 DeepSeek-V2.5 6.6 - $0
541 Olmo 3.1 32B Instruct 6.5 - $0
542 Mistral Saba 6.5 - $0
543 DeepSeek R1 Distill Llama 8B 6.5 - $0
544 Gemma 4 E2B (Non-reasoning) 6.5 - $0
545 Olmo 3 32B Think 6.5 - $0
546 Gemini 1.5 Pro (May '24) 6.4 19.8 $0
547 R1 1776 6.4 - $0
548 Qwen2.5 Turbo 6.4 - $0.088
549 Reka Flash (Sep '24) 6.4 - $0.35
550 Llama 3.2 Instruct 90B (Vision) 6.4 - $0
551 Solar Mini 6.4 - $0.15
552 Celeris-1 6.3 14.4 $0.325
553 Grok-1 6.3 - $0
554 Qwen2 Instruct 72B 6.3 - $0
555 Phi-4 Mini Instruct 6.3 3.8 $0
556 EXAONE 4.0 32B (Non-reasoning) 6.3 - $0
557 Qwen3.5 2B (Non-reasoning) 6.2 2.4 $0
558 Gemini 1.5 Flash-8B 6.2 - $0
559 Qwen3.5 0.8B (Reasoning) 6.1 0 $0
560 DeepHermes 3 - Mistral 24B Preview (Non-reasoning) 6.1 - $0
561 Jamba 1.7 Large 6.1 - $0
562 Granite 4.0 H Small 6 - $0.107
563 Ministral 3 14B 6 14.4 $0.2
564 Jamba 1.5 Large 6 - $3.5
565 Qwen3 Omni 30B A3B Instruct 6 - $0.43
566 Hermes 3 - Llama-3.1 70B 6 - $0.7
567 Qwen3 8B (Non-reasoning) 6 - $0.31
568 DeepSeek-Coder-V2 6 - $0
569 OLMo 2 32B 6 - $0
570 Jamba 1.6 Large 6 - $0
571 LFM2 24B A2B 5.9 - $0
572 Gemini 1.5 Flash (May '24) 5.9 - $0
573 Phi-4 5.9 - $0.219
574 Claude 3 Sonnet 5.9 - $0
575 Nova Micro 5.9 - $0.061
576 Granite 4.1 3B 5.9 4.7 $0
577 Mistral Small (Sep '24) 5.8 - $0
578 Gemini 1.0 Ultra 5.8 17.6 $0
579 Phi-3 Mini Instruct 3.8B 5.8 - $0
580 NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) 5.8 - $0.3
581 Gemma 3n E4B Instruct Preview (May '25) 5.8 - $0
582 Phi-4 Multimodal Instruct 5.8 - $0
583 Qwen2.5 Coder Instruct 7B 5.8 - $0
584 Mistral Large (Feb '24) 5.8 - $6
585 Mixtral 8x22B Instruct 5.7 - $0
586 Llama 2 Chat 7B 5.7 - $0.1
587 Llama 3.2 Instruct 3B 5.7 - $0
588 MiniCPM-V 4.6 1.3B 5.7 0.7 $0
589 Jamba Reasoning 3B 5.7 - $0
590 Qwen3 VL 4B Instruct 5.7 - $0
591 Qwen1.5 Chat 110B 5.7 - $0
592 Reka Flash 3 5.6 - $0.35
593 Olmo 3 7B Think 5.6 - $0
594 Claude 2.1 5.6 14 $0
595 Claude 3 Haiku 5.6 - $0.5
596 OLMo 2 7B 5.6 - $0
597 Molmo 7B-D 5.6 - $0
598 Ling-mini-2.0 5.5 - $0
599 DeepSeek R1 Distill Qwen 1.5B 5.5 - $0
600 Claude 2.0 5.5 12.9 $0
601 DeepSeek-V2-Chat 5.5 - $0
602 Mistral Small (Feb '24) 5.5 - $0
603 Mistral Medium 5.5 - $0
604 GPT-3.5 Turbo 5.5 10.7 $0.75
605 Ministral 3 8B 5.5 9.7 $0.15
606 Llama 3 Instruct 70B 5.5 - $1.175
607 Arctic Instruct 5.4 - $0
608 Qwen Chat 72B 5.4 - $0
609 LFM 40B 5.4 - $0
610 Llama 3.2 Instruct 11B (Vision) 5.4 - $0.345
611 Qwen3.5 0.8B (Non-reasoning) 5.4 1.2 $0
612 PALM-2 5.4 4.6 $0
613 Gemini 1.0 Pro 5.3 - $0
614 DeepSeek Coder V2 Lite Instruct 5.3 - $0
615 Sarvam M (Reasoning, based on Mistral Small 3.1) 5.3 - $0
616 DeepSeek LLM 67B Chat (V1) 5.3 - $0
617 Llama 2 Chat 70B 5.3 - $0
618 Llama 2 Chat 13B 5.3 - $0
619 Command-R+ (Apr '24) 5.3 - $0
620 OpenChat 3.5 (1210) 5.3 - $0
621 DBRX Instruct 5.3 - $0
622 Exaone 4.0 1.2B (Reasoning) 5.3 - $0
623 Olmo 3 7B Instruct 5.2 - $0.125
624 Exaone 4.0 1.2B (Non-reasoning) 5.2 - $0
625 LFM2.5-1.2B-Thinking 5.2 - $0
626 Jamba 1.7 Mini 5.2 - $0
627 LFM2 2.6B 5.2 - $0
628 LFM2.5-1.2B-Instruct 5.2 - $0
629 Jamba 1.5 Mini 5.2 - $0.25
630 Granite 4.0 H 1B 5.2 - $0
631 Qwen3 1.7B (Reasoning) 5.2 - $0
632 Jamba 1.6 Mini 5.2 - $0
633 Mixtral 8x7B Instruct 5.1 - $0.512
634 Gemma 3 270M 5.1 - $0
635 Apertus 70B Instruct 5.1 - $1.345
636 Granite 4.0 Micro 5.1 - $0
637 DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning) 5.1 - $0
638 Claude Instant 5 7.8 $0
639 Command-R (Mar '24) 5 - $0
640 Llama 65B 5 - $0
641 Mistral 7B Instruct 5 - $0.162
642 Qwen Chat 14B 5 - $0
643 Granite 4.0 1B 5 - $0
644 Molmo2-8B 5 - $0
645 LFM2 8B A1B 4.9 - $0
646 Granite 3.3 8B (Non-reasoning) 4.9 - $0.085
647 Qwen3 1.7B (Non-reasoning) 4.9 - $0
648 Gemma 3 27B Instruct 4.9 10.1 $0
649 Ministral 3 3B 4.8 4.8 $0.1
650 Apertus 8B Instruct 4.8 - $0.125
651 Gemma 3 1B Instruct 4.8 - $0
652 Gemma 3 4B Instruct 4.8 2.7 $0
653 Gemma 3n E2B Instruct 4.8 - $0
654 Gemma 3n E4B Instruct 4.8 3.2 $0
655 Granite 4.0 350M 4.8 - $0
656 Granite 4.0 H 350M 4.8 - $0
657 LFM2 1.2B 4.8 - $0
658 LFM2.5-VL-1.6B 4.8 - $0
659 Llama 3 Instruct 8B 4.8 - $0.07
660 Llama 3.2 Instruct 1B 4.8 - $0
661 Qwen3 0.6B (Non-reasoning) 4.8 - $0
662 Qwen3 0.6B (Reasoning) 4.8 - $0
663 Tiny Aya Global 4.8 - $0
664 Gemma 3 12B Instruct 3.8 5.8 $0
665 K2 Horizon 0.9B 3 3.4 $0
666 Cogito v2.1 (Reasoning) - - $1.25
667 EXAONE 4.5 33B (Non-reasoning) - - $0
668 Gemini 3 Deep Think - - $0
669 GPT-3.5 Turbo (0613) - - $0
670 GPT-4o mini Realtime (Dec '24) - - $0
671 GPT-4o Realtime (Dec '24) - - $0
672 GPT-5.4 Pro (xhigh) - - $67.5
673 GPT-5.5 Pro (xhigh) - - $0
674 Mi:dm K 2.5 Pro Preview - - $0

榜单解读建议

参考 AI 大模型排行榜 时,应综合考虑“综合指数”与“成本价格”。如果您是开发者,编程能力 (Coding) 是更核心的指标。

值品工具箱同步的 AI 大模型排行榜 数据每 24 小时更新,确保您获取到最新的模型性能对比。

指标说明

  • ● 综合指数:评估通用理解与逻辑。
  • ● 价格 $/1M:混合 3:1 输入输出比的平均成本。
  • ● 编程能力:衡量代码生成的准确性。

AI 大模型排行榜 常见问题 (FAQ)

Q1: AI 大模型排行榜 的数据多久更新?

AI 大模型排行榜 数据每 24 小时自动抓取一次,确保最新模型加入列表。

Q2: 这个 AI 大模型排行榜 包含国产模型吗?

是的,只要国产模型通过了 Artificial Analysis 的全球测评,就会出现在 AI 大模型排行榜 中。

Q3: 综合指数在 AI 大模型排行榜 中代表什么?

它代表模型的全能表现。AI 大模型排行榜 通过加权算法给出这个综合评分。

Q4: 如何在 AI 大模型排行榜 中查找性价比最高的游戏?

在 AI 大模型排行榜 页面中,您可以点击“价格”标题进行排序,寻找低价高分的模型。

Q5: AI 大模型排行榜 的编程能力测试准吗?

AI 大模型排行榜 参考了 LiveCodeBench 等权威基准测试,具有极高的参考价值。

Q6: 为什么有的新模型没进入 AI 大模型排行榜?

模型进入 AI 大模型排行榜 需要经过一系列测试,通常在新模型发布后数日内会完成更新。

Q7: AI 大模型排行榜 中的价格计算标准是什么?

价格是基于百万 Token 的调用成本,由 AI 大模型排行榜 统一混合计算得出。

Q8: 手机上能查看 AI 大模型排行榜 吗?

当然可以。AI 大模型排行榜 进行了移动端响应式深度优化。

Q9: AI 大模型排行榜 这个工具免费吗?

是的,由值品工具箱免费提供 AI 大模型排行榜 信息查询服务。

Q10: 我该怎么利用 AI 大模型排行榜 做选型?

如果您需要智能客服,参考 AI 大模型排行榜 的综合指数;如果做翻译,参考编程外的语言指标。

发表评论

请友善文明留言