Matches Regex

Test: Language Comprehension

Avg. Score
73.1%
Scenarios
2

Overall Performance

Rank ▲ Model Score Avg. Cost Avg. Time Stability
1Ministral 3 3B100.0%$0.0000811ms100%
2Mistral Large 3100.0%$0.00014.2s100%
3DeepSeek V3 (2024-12-26)100.0%$0.00024.5s100%
4Qwen 3.5 Plus (2026-02-15)100.0%$0.00035.4s100%
5DeepSeek V3.1100.0%$0.00025.6s100%
6GPT-4o, May 13th (temp=0)100.0%$0.00222.9s100%
7Mistral Large 2100.0%$0.00124.9s100%
8DeepSeek-V2 Chat100.0%$0.00017.5s100%
9Claude Sonnet 4100.0%$0.00253.4s100%
10Claude Sonnet 4.6100.0%$0.00263.3s100%
11Claude Sonnet 4.5100.0%$0.00283.7s100%
12DeepSeek V3 (2025-03-24)100.0%$0.00019.8s100%
13Claude Opus 4.5100.0%$0.00433.7s100%
14Hermes 3 405B100.0%$0.000011.9s100%
15DeepSeek V3.2100.0%$0.000212.8s100%
16Claude Opus 4.6100.0%$0.00474.9s100%
17ByteDance Seed 1.6100.0%$0.001214.4s100%
18Gemini 2.5 Pro100.0%$0.00848.1s100%
19Z.AI GLM 4.7 Flash100.0%$0.000523.4s100%
20Claude Opus 4100.0%$0.010011.3s100%
21Gemini 3 Pro (Preview)100.0%$0.0128.8s100%
22Gemini 3.1 Pro (Preview)100.0%$0.01112.6s100%
23Z.AI GLM 4.6100.0%$0.002132.1s100%
24Qwen 3.5 397B A17B100.0%$0.004837.0s100%
25MoonshotAI: Kimi K2.5100.0%$0.005843.3s100%
26Z.AI GLM 4.7100.0%$0.002454.2s100%
27Z.AI GLM 5100.0%$0.004052.7s100%
28Mistral NeMO90.0%$0.0000380ms40%
29Gemini 2.5 Flash Lite90.0%$0.0001728ms40%
30Gemini 3 Flash (Preview)90.0%$0.00031.2s40%
31Z.AI GLM 4.590.0%$0.00013.1s40%
32Claude Haiku 4.590.0%$0.00092.1s40%
33GPT-4.190.0%$0.00132.7s40%
34WizardLM 2 8x22b90.0%$0.00025.0s40%
35GPT-4o, May 13th (temp=1)90.0%$0.00202.8s40%
36GPT-5.290.0%$0.00265.0s40%
37GPT-5 Mini90.0%$0.001512.5s40%
38GPT-5.190.0%$0.00459.8s40%
39Minimax M2.590.0%$0.001932.8s40%
40GPT-590.0%$0.01628.4s40%
41Gemini 2.5 Flash80.0%$0.0003987ms20%
42Gemma 3 4B80.0%$0.00002.2s20%
43Claude 3.5 Haiku80.0%$0.00051.8s20%
44Mistral Large80.0%$0.00123.7s20%
45Claude 3.7 Sonnet80.0%$0.00243.4s20%
46Grok 4.1 Fast80.0%$0.001323.5s20%
47Claude 3 Haiku70.0%$0.00011.3s8%
48Grok 4 Fast70.0%$0.00045.7s8%
49Cohere Command R+ (Aug. 2024)70.0%$0.00285.5s8%
50o4 Mini70.0%$0.00267.4s8%
51o4 Mini High70.0%$0.004613.2s8%
52Arcee AI: Trinity Large (Preview)60.0%$0.00001.7s2%
53GPT-4.1 Mini60.0%$0.00042.8s2%
54Llama 3.1 70B60.0%$0.00023.4s2%
55Claude 3.5 Sonnet60.0%$0.00193.6s2%
56GPT-5 Nano60.0%$0.000617.5s2%
57Ministral 8B50.0%$0.0000347ms0%
58GPT-4o Mini (temp=1)50.0%$0.0000862ms0%
59GPT-4o Mini (temp=0)50.0%$0.0000877ms0%
60Mistral Small 3.2 24B50.0%$0.00002.4s0%
61Hermes 3 70B50.0%$0.00013.2s0%
62Gemma 3 12B50.0%$0.00004.9s0%
63Rocinante 12B50.0%$0.00015.5s0%
64Writer: Palmyra X550.0%$0.00177.8s0%
65Qwen 2.5 72B40.0%$0.00013.9s0%
66ByteDance Seed 1.6 Flash40.0%$0.00025.1s0%
67Stealth: Aurora Alpha70.0%1.4s8%
68Llama 3.1 Nemotron 70B40.0%$0.000213.5s0%
69Ministral 3B30.0%$0.0000357ms0%
70Llama 3.1 8B30.0%$0.00001.4s0%
71GPT-4.1 Nano30.0%$0.00001.6s0%
72GPT-4o, Aug. 6th (temp=1)30.0%$0.00071.3s0%
73Gemma 3 27B30.0%$0.00018.2s0%
74Arcee AI: Trinity Mini40.0%$0.000132.5s0%
75Ministral 3 8B0.0%$0.0000711ms0%
76Ministral 3 14B0.0%$0.0000912ms0%
77Mistral Small Creative0.0%$0.00011.1s0%
78GPT-4o, Aug. 6th (temp=0)0.0%$0.00071.9s0%
79Mistral Medium 3.10.0%$0.00043.6s0%
80Grok 470.0%$0.1103.2m8%
73.13%

Individual Scenarios

Model # 1 # 2 # 3 # 4 # 5 Avg ▼
Gemini 3.1 Pro (Preview)100100100100100100.0%
Qwen 3.5 397B A17B100100100100100100.0%
Claude Opus 4.6100100100100100100.0%
GPT-5.1100100100100100100.0%
MoonshotAI: Kimi K2.5100100100100100100.0%
Claude Opus 4.5100100100100100100.0%
GPT-5100100100100100100.0%
Z.AI GLM 5100100100100100100.0%
Gemini 2.5 Pro100100100100100100.0%
GPT-5.2100100100100100100.0%
Z.AI GLM 4.7100100100100100100.0%
Gemini 3 Pro (Preview)100100100100100100.0%
Claude Opus 4100100100100100100.0%
Claude Sonnet 4100100100100100100.0%
Claude Sonnet 4.6100100100100100100.0%
Claude Sonnet 4.5100100100100100100.0%
Grok 4.1 Fast100100100100100100.0%
ByteDance Seed 1.6100100100100100100.0%
Z.AI GLM 4.6100100100100100100.0%
GPT-4.1100100100100100100.0%
Z.AI GLM 4.7 Flash100100100100100100.0%
Qwen 3.5 Plus (2026-02-15)100100100100100100.0%
DeepSeek V3 (2025-03-24)100100100100100100.0%
Grok 4 Fast100100100100100100.0%
DeepSeek V3 (2024-12-26)100100100100100100.0%
Mistral Large 3100100100100100100.0%
GPT-4o, May 13th (temp=0)100100100100100100.0%
DeepSeek-V2 Chat100100100100100100.0%
DeepSeek V3.2100100100100100100.0%
Z.AI GLM 4.5100100100100100100.0%
GPT-4o, May 13th (temp=1)100100100100100100.0%
DeepSeek V3.1100100100100100100.0%
Mistral Large 2100100100100100100.0%
Hermes 3 405B100100100100100100.0%
Ministral 3 3B100100100100100100.0%
GPT-5 Mini100100100100080.0%
o4 Mini High100100100100080.0%
Minimax M2.5100100100100080.0%
Gemini 3 Flash (Preview)100100100100080.0%
Stealth: Aurora Alpha100100100100080.0%
GPT-5 Nano100100100100080.0%
Claude Haiku 4.5100100100100080.0%
GPT-4.1 Mini100100100100080.0%
Writer: Palmyra X5100100100100080.0%
Gemini 2.5 Flash100100100100080.0%
Gemini 2.5 Flash Lite100100100100080.0%
Mistral NeMO100100100100080.0%
WizardLM 2 8x22b100100100100080.0%
Grok 41001001000060.0%
Claude 3.7 Sonnet1001001000060.0%
Claude 3.5 Haiku1001001000060.0%
Mistral Large1001001000060.0%
Llama 3.1 8B1001001000060.0%
Gemma 3 4B1001001000060.0%
o4 Mini10010000040.0%
ByteDance Seed 1.6 Flash10010000040.0%
Arcee AI: Trinity Large (Preview)10010000040.0%
Claude 3 Haiku10010000040.0%
Arcee AI: Trinity Mini10010000040.0%
Cohere Command R+ (Aug. 2024)10010000040.0%
GPT-4.1 Nano10010000040.0%
Ministral 8B10010000040.0%
Ministral 3B10010000040.0%
Rocinante 12B10010000040.0%
Claude 3.5 Sonnet100000020.0%
GPT-4o, Aug. 6th (temp=1)100000020.0%
Llama 3.1 70B100000020.0%
Llama 3.1 Nemotron 70B100000020.0%
Hermes 3 70B100000020.0%
GPT-4o, Aug. 6th (temp=0)000000.0%
Mistral Medium 3.1000000.0%
GPT-4o Mini (temp=1)000000.0%
GPT-4o Mini (temp=0)000000.0%
Gemma 3 12B000000.0%
Gemma 3 27B000000.0%
Mistral Small Creative000000.0%
Ministral 3 14B000000.0%
Qwen 2.5 72B000000.0%
Mistral Small 3.2 24B000000.0%
Ministral 3 8B000000.0%
Model # 1 # 2 # 3 # 4 # 5 Avg ▼
Gemini 3.1 Pro (Preview)100100100100100100.0%
Qwen 3.5 397B A17B100100100100100100.0%
GPT-5 Mini100100100100100100.0%
Claude Opus 4.6100100100100100100.0%
MoonshotAI: Kimi K2.5100100100100100100.0%
Claude Opus 4.5100100100100100100.0%
o4 Mini100100100100100100.0%
Z.AI GLM 5100100100100100100.0%
Gemini 2.5 Pro100100100100100100.0%
Z.AI GLM 4.7100100100100100100.0%
Gemini 3 Pro (Preview)100100100100100100.0%
Claude Opus 4100100100100100100.0%
Minimax M2.5100100100100100100.0%
Claude Sonnet 4100100100100100100.0%
Claude Sonnet 4.6100100100100100100.0%
Claude Sonnet 4.5100100100100100100.0%
ByteDance Seed 1.6100100100100100100.0%
Z.AI GLM 4.6100100100100100100.0%
Gemini 3 Flash (Preview)100100100100100100.0%
Z.AI GLM 4.7 Flash100100100100100100.0%
Qwen 3.5 Plus (2026-02-15)100100100100100100.0%
DeepSeek V3 (2025-03-24)100100100100100100.0%
Claude 3.5 Sonnet100100100100100100.0%
DeepSeek V3 (2024-12-26)100100100100100100.0%
Mistral Large 3100100100100100100.0%
GPT-4o, May 13th (temp=0)100100100100100100.0%
DeepSeek-V2 Chat100100100100100100.0%
Claude 3.7 Sonnet100100100100100100.0%
Claude Haiku 4.5100100100100100100.0%
DeepSeek V3.2100100100100100100.0%
Claude 3.5 Haiku100100100100100100.0%
DeepSeek V3.1100100100100100100.0%
Mistral Large 2100100100100100100.0%
Hermes 3 405B100100100100100100.0%
GPT-4o Mini (temp=1)100100100100100100.0%
GPT-4o Mini (temp=0)100100100100100100.0%
Gemma 3 12B100100100100100100.0%
Llama 3.1 70B100100100100100100.0%
Gemini 2.5 Flash Lite100100100100100100.0%
Mistral Large100100100100100100.0%
Mistral Small 3.2 24B100100100100100100.0%
Claude 3 Haiku100100100100100100.0%
Ministral 3 3B100100100100100100.0%
Cohere Command R+ (Aug. 2024)100100100100100100.0%
Mistral NeMO100100100100100100.0%
Gemma 3 4B100100100100100100.0%
WizardLM 2 8x22b100100100100100100.0%
GPT-5.1100100100100080.0%
GPT-5100100100100080.0%
GPT-5.2100100100100080.0%
Grok 4100100100100080.0%
GPT-4.1100100100100080.0%
Z.AI GLM 4.5100100100100080.0%
GPT-4o, May 13th (temp=1)100100100100080.0%
Gemini 2.5 Flash100100100100080.0%
Qwen 2.5 72B100100100100080.0%
Arcee AI: Trinity Large (Preview)100100100100080.0%
Hermes 3 70B100100100100080.0%
o4 Mini High1001001000060.0%
Grok 4.1 Fast1001001000060.0%
Stealth: Aurora Alpha1001001000060.0%
Llama 3.1 Nemotron 70B1001001000060.0%
Gemma 3 27B1001001000060.0%
Ministral 8B1001001000060.0%
Rocinante 12B1001001000060.0%
GPT-5 Nano10010000040.0%
Grok 4 Fast10010000040.0%
GPT-4o, Aug. 6th (temp=1)10010000040.0%
GPT-4.1 Mini10010000040.0%
ByteDance Seed 1.6 Flash10010000040.0%
Arcee AI: Trinity Mini10010000040.0%
Writer: Palmyra X5100000020.0%
GPT-4.1 Nano100000020.0%
Ministral 3B100000020.0%
GPT-4o, Aug. 6th (temp=0)000000.0%
Mistral Medium 3.1000000.0%
Mistral Small Creative000000.0%
Ministral 3 14B000000.0%
Ministral 3 8B000000.0%
Llama 3.1 8B000000.0%