Overview

12,500 道竞赛数学题(含 7 个学科:代数、几何、数论、微积分、概率等,5 个难度级别),评估模型的复杂数学推理能力。

Metrics

MetricUnitDirection
accuracy%↑ Higher is better

Sources

Model Score Ranking

#ModelVendorScore
1Qwen2.5 72Balibaba83.1
2Grok-2xai76.8
3GPT-4oopenai76.6
4Claude 3.5 Sonnet (2024-10-22)anthropic75.1
5Llama 3.1 405Bmeta73.8
6Llama 3.3 70Bmeta73.8
7Qwen2 72Balibaba71.9
8o1 Previewopenai71.5
9Gemini 2.0 Flashgoogle71.3
10Claude 3.5 Sonnetanthropic71.1
11Mistral Large 2mistral71
12GPT-4o (2024-08-06)openai68.2
13GPT-4o (2024-05-13)openai67.1
14Microsoft WizardMath 7B v1other65.8
15Hermes 3 Llama 3.1 405Bother65.7
16DeepSeek V3deepseek61.6
17Mathstral 7Bmistral61
18Jamba 1.5 Largeother60
19Yi Visionother58.7
20Gemini 1.5 Progoogle58.5
21Claude 3 Opus (2024-02-29)anthropic58.2
22Gemini 1.0 Progoogle58.2
23Gemini 1.0 Ultragoogle57.1
24Gemini 2.0 Flash Thinkinggoogle56.6
25Gemini 1.5 Pro 002google55.9
26Yi Largeother55.9
27o1openai55.7
28Mistral Mediummistral54.9
29Qwen2 57Balibaba54.8
30Hermes 3 Llama 3.1 70Bother54.4
31Qwen1.5 110Balibaba54.1
32DeepSeek Math 7Bdeepseek53.8
33GLM-4 Plusother53.3
34Gemini 1.5 Flashgoogle53.2
35Claude 3.5 Haikuanthropic53
36Llama 3.2 90B Visionmeta52.5
37Claude 3 Haikuanthropic52.2
38GPT-4 1106 Previewopenai52.1
39GPT-4o miniopenai51.8
40Sonar Largeother51.6
41GPT-4 Turboopenai50.9
42Sonar Reasoningother50.7
43Yi 1.5 34Bother50.4
44DBRX Instructother50.1
45Llama 3 70Bmeta50.1
46GPT-4 Vision Previewopenai49.7
47DeepSeek V2 Chatdeepseek49.3
48Mistral Largemistral49.3
49Command Nightlycohere49.2
50Jamba Instructother49
51GPT-4 Visionopenai48.3
52Jamba 1.5other48.1
53Llemma 7Bother47.5
54Code Llama 7Bmeta47.1
55Command R (08-2024)cohere47.1
56Sonar Hugeother46.8
57Claude 3 Sonnet (2024-02-29)anthropic46.4
58Qwen1.5 32Balibaba46.3
59Qwen1.5 14Balibaba46.2
60StarCoder2 15Bother46.2
61Mixtral 8x22Bmistral46
62Mistral Smallmistral45.8
63StarChat2 15B v0.1other45.3
64Grok-2 Minixai45.1
65Nous Hermes 2 Yi 34Bother44.9
66Falcon 180Bother44.8
67Claude 3 Haiku (2024-03-07)anthropic44.7
68Code Llama 34Bmeta44.7
69Yi 1.5 6Bother44.5
70o1 miniopenai44.2
71Command Rcohere43.9
72GPT-4openai43.9
73Mistral Nemomistral43.6
74Phi-3 Smallother43.5
75Command R+ (08-2024)cohere43.3
76Sonar Smallother43.3
77OLMo 7B SFTother43.1
78Gemini 1.5 Flash-8B 002google43
79Claude 3 Opusanthropic42.7
80Nous Hermes 2 Mixtral 8x7Bother42.6
81Jamba 1.5 Miniother42.5
82Claude 3 Sonnetanthropic42.4
83GPT-4 0125 Previewopenai42.3
84Orca 2 13Bother42.3
85Yi Large Turboother42.2
86Qwen2.5 7Balibaba41.9
87Grok-2 Visionxai41.3
88Qwen2 7Balibaba40.7
89Command R7Bcohere40.6
90GLM-4 Airother40.6
91Microsoft WizardCoder Python 34Bother40.6
92Gemini 1.5 Flash 002google40.5
93NVIDIA Llama 3.1 Nemotron 70Bother40.3
94GPT-4 32Kopenai40.1
95Llama 3.1 Nemotron 70Bmeta40
96Llama 3 8Bmeta40
97Gemini 1.0 Flashgoogle39.9
98Ministral 8Bmistral39.9
99Microsoft WizardLM 2 8x22Bother39.8
100Phi-4other39.6
101Qwen2.5 32Balibaba39.5
102GLM-4 Flashother39.4
103DeepSeek LLM 67Bdeepseek39.3
104OLMo 2 1124 7Bother39.3
105Mistral Small 3mistral39.2
106Qwen2.5 14Balibaba38.8
107Llama 3.1 70Bmeta38.5
108DeepSeek Coder V2deepseek38.2
109Phi-3.5 MoEother38.2
110OLMo 1.7 7Bother38.1
111Mixtral 8x7Bmistral38
112StableLM 2 12Bother37.9
113Code Llama 13Bmeta37.8
114Nous Hermes 2 Solar 10.7Bother37.7
115Qwen1.5 72Balibaba37.6
116StarCoder2 3Bother37.4
117WizardLM Team WizardLM 2 8x22Bother37.4
118OLMo 7Bother37.1
119DBRX Baseother36
120Zephyr ORPO 141B Alphaother36
121Command R+cohere35.6
122Gemma 2 27Bgoogle35.5
123DeepSeek V2deepseek35.2
124Gemini 1.5 Flash-8Bgoogle35
125OLMo 7B Instructother34.7
126Code Bisongoogle34.4
127Phi-3 Mediumother34.3
128Yi 1.5 9Bother34.1
129Code Llama 70Bmeta34
130OpenChat 3.6 8Bother33.7
131GLM-4V 9Bother33.3
132DeepSeek Coder 7Bdeepseek32.2
133DeepSeek Coder 33Bdeepseek32.1
134ChatGLM3 6Bother31.8
135Zephyr 7B Betaother31.3
136Llama 3.1 8Bmeta31.1
137Argilla Notus 7B v1other31
138StableCode 3Bother31
139Microsoft WizardLM 2 7Bother30.5
140StarCoder2 7Bother30.5
141GPT-3.5 Turbo 16Kopenai30.4
142Llama 3.2 11B Visionmeta29.9
143PaLM 2google29.9
144Yi 6Bother29.9
145Codestralmistral29.4
146GPT-3.5openai29.4
147Gemma 2 9Bgoogle28.9
148Codestral Mambamistral28.1
149Flan-T5 XLother27.8
150Gemma 2Bgoogle27.5
151Jurassic-2 Ultraother27.5
152Phi-1.5other27.2
153Mistral 7B v0.3mistral26.7
154GLM-4 9B Chatother26.6
155Mistral 7B v0.1mistral26.5
156StableLM 2 1.6Bother26.4
157Llama Guard 2 8Bmeta26.3
158Llama 2 13Bmeta25.9
159Mistral 7B v0.2mistral25.9
160Gemma 7Bgoogle25.7
161Capybara 1.5Bother25.6
162Hermes 3 Llama 3.1 8Bother25.5
163Qwen2 1.5Balibaba25.4
164Phi-1other25.2
165Phi-3 Visionother24.3
166Command Lightcohere24.1
167Grok Vision Betaxai24
168Baichuan2 13B Chatother23.4
169Zephyr 7B Alphaother23.2
170Nous Capybara 34Bother23.1
171Claude Instant 1anthropic22.5
172Llama 2 70Bmeta22.4
173Flan-UL2other22
174Claude 2anthropic21.9
175Flan-T5 XXLother21.4
176Falcon 40Bother21.3
177MPT 30Bother21.2
178Open-Platypusother21
179Phi-2other20.5
180Yi 34Bother20.5
181Jurassic-2 Midother20.3
182Chat Bisongoogle19.8
183Nous Capybara 7Bother19.7
184Llama 2 7Bmeta19.4
185Grok Betaxai19.1
186Baichuan2 7B Chatother18.8
187MPT 7Bother18.4
188Phi-3 Miniother17.9
189GPT-3.5 Turboopenai17.8
190Qwen2.5 3Balibaba17
191Mistral Tinymistral16
192Claude 2.1anthropic15.5
193Qwen2.5 0.5Balibaba15.4
194TigerBot 70B Chatother15.4
195OpenChat 3.5 1210other15.2
196StableLM Zephyr 3Bother14.9
197Text Bisongoogle12.1
198StableLM 3 4Bother11.5
199Llama 3.2 1Bmeta11.4
200Ministral 3Bmistral10.1
201Llama 3.2 3Bmeta10
202Phi-3.5 Miniother9.6
203Qwen2.5 1.5Balibaba9.5
204Llama Guard 3 8Bmeta8.4
205Embed English v3cohere0

MATH

Açıklama

12,500 道竞赛数学题(含 7 个学科:代数、几何、数论、微积分、概率等,5 个难度级别),评估模型的复杂数学推理能力。

Temel Özellikler

KategoriLisansSon Güncelleme
mathMIT2021-01-01

Benchmark

Birim
%

Kaynaklar

Resmi URL