Language Comprehension
Does the model understand more than just English?
Performance Score Distribution (Top 20)
Click a model name to view its detail page.
Price-Performance Score Distribution (Top 20)
Click a model name to view its detail page.
| Score | Cost | Time | ||
|---|---|---|---|---|
| Mistral NeMO | 95% | $0.0000 | 625ms | |
| Ministral 3 3B | 100% | $0.0000 | 833ms | |
| Ministral 8B | 55% | $0.0000 | 540ms | |
| Gemini 2.5 Flash Lite | 80% | $0.0001 | 648ms | |
| GPT-4o Mini (temp=1) | 55% | $0.0000 | 942ms | |
| GPT-6 Luna | 90% | $0.0000 | 1.6s | |
| GPT-5.4 Nano | 70% | $0.0001 | 990ms | |
| Gemini 3.1 Flash Lite (Preview) | 95% | $0.0001 | 975ms | |
| Gemini 3.1 Flash Lite (Reasoning) | 95% | $0.0002 | 1.6s | |
| Gemma 3 4B | 70% | $0.0000 | 1.7s | |
| Gemini 3.1 Flash Lite | 85% | $0.0001 | 905ms | |
| GPT-5.4 Mini | 80% | $0.0002 | 747ms | |
| Mistral Small 3.2 24B | 75% | $0.0000 | 2.4s | |
| Cydonia 24B V4.1 | 95% | $0.0001 | 2.2s | |
| GPT-4.1 Nano | 65% | $0.0000 | 1.7s | |
| Mistral Small 4 | 55% | $0.0001 | 1.1s | |
| DeepSeek V4 Flash | 90% | $0.0000 | 21.6s | |
| GPT-5.6 Luna | 100% | $0.0003 | 911ms | |
| Inception Mercury 2 | 75% | $0.0003 | 771ms | |
| Gemini 2.5 Flash | 75% | $0.0002 | 817ms | |
Cost vs Performance
Compares total cost for this test against the test score. Quadrant lines are drawn at the median values. Only models with available cost data are shown.
10 low-scoring outliers hidden: Gemma 4 31B (50.0%), GPT-4o, Aug. 6th (temp=0) (50.0%), Mistral Small 4 (Reasoning) (50.0%), GPT-4o Mini (temp=0) (50.0%), Mistral Medium 3.1 (50.0%), Ministral 3 14B (50.0%), Ministral 3 8B (50.0%), ByteDance Seed 1.6 Flash (40.0%), Laguna XS 2.1 (35.0%), Ministral 3B (25.0%).
Most Stable Models (Top 20)
Ranked by stability (median × consistency). Click a model name to view its detail page.
| Score | Consistency | Stability | ||
|---|---|---|---|---|
| Claude Opus 5.5 (Reasoning) | 100% | 100% | 100% | |
| Qwen 3.8 Max (Reasoning, XHigh) | 100% | 100% | 100% | |
| Gemini 3.8 Flash (Reasoning, Medium) | 100% | 100% | 100% | |
| Z.AI GLM 5.3 (Reasoning, Max) | 100% | 100% | 100% | |
| GPT-5.6 Sol (Reasoning, Medium) | 100% | 100% | 100% | |
| Claude Opus 4.6 (Reasoning) | 100% | 100% | 100% | |
| Gemini 3.6 Flash (Reasoning) | 100% | 100% | 100% | |
| Gemini 3.7 Flash (Reasoning, Medium) | 100% | 100% | 100% | |
| Hy4 Preview (Reasoning, High) | 100% | 100% | 100% | |
| Muse Spark 1.2 (Reasoning, Medium) | 100% | 100% | 100% | |
| Qwen 3.7 Max | 100% | 100% | 100% | |
| Z.AI GLM 5.3 Flash (Reasoning, Max) | 100% | 100% | 100% | |
| Qwen 3.8 Flash (Reasoning) | 100% | 100% | 100% | |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | |
| GPT-6 Luna (Reasoning, High) | 100% | 100% | 100% | |
| Qwen 3.6 Max Preview | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning, Medium) | 100% | 100% | 100% | |
| Claude Sonnet 4.6 (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5.2 (Reasoning, High) | 100% | 100% | 100% | |
| Z.AI GLM 5 Turbo | 100% | 100% | 100% | |
Top Overall Models (Top 20)
Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.
| Score | Cost | Speed | Stability | ||
|---|---|---|---|---|---|
| Ministral 3 3B | 100% | $0.0000 | 833ms | 100% | |
| GPT-5.6 Luna | 100% | $0.0003 | 911ms | 100% | |
| GPT-6 Sol | 100% | $0.0004 | 2.1s | 100% | |
| Mistral Large 3 | 100% | $0.0002 | 3.7s | 100% | |
| GPT-6 Luna (Reasoning, Medium) | 100% | $0.0001 | 4.3s | 100% | |
| DeepSeek V3 (2024-12-26) | 100% | $0.0002 | 4.8s | 100% | |
| GPT-6 Luna (Reasoning, High) | 100% | $0.0001 | 5.0s | 100% | |
| Qwen 3.5 Plus (2026-02-15) | 100% | $0.0003 | 4.8s | 100% | |
| DeepSeek-V2 Chat | 100% | $0.0000 | 6.6s | 100% | |
| DeepSeek V3 (2025-03-24) | 100% | $0.0001 | 7.5s | 100% | |
| Mistral Large 2 | 100% | $0.0009 | 3.8s | 100% | |
| GPT-6 Sol (Reasoning, Medium) | 100% | $0.0011 | 3.8s | 100% | |
| Hermes 3 405B | 100% | $0.0000 | 12.1s | 100% | |
| GPT-5.4 Mini (Reasoning) | 100% | $0.0012 | 4.8s | 100% | |
| Claude Sonnet 4.6 | 100% | $0.0021 | 3.1s | 100% | |
| Gemini 3.8 Flash (Reasoning, Medium) | 100% | $0.0022 | 3.0s | 100% | |
| Gemini 3.7 Flash (Reasoning, Medium) | 100% | $0.0018 | 6.7s | 100% | |
| Muse Glimmer 30B (Reasoning, Medium) | 100% | $0.0010 | 13.2s | 100% | |
| Aion 2.0 | 100% | $0.0011 | 15.9s | 100% | |
| ByteDance Seed 1.6 | 100% | $0.0013 | 15.8s | 100% | |
| Model | Total â–¼ | Friend got new kittens (Tagalog) | Friend got new kittens (German) | Asking for directions (German) | Asking for directions (Dutch) |
|---|---|---|---|---|---|
| Claude Opus 5.5 (Reasoning) | 100% | 100% | 100% | 100% | 100% |
| Qwen 3.8 Max (Reasoning, XHigh) | 100% | 100% | 100% | 100% | 100% |
| Gemini 3.8 Flash (Reasoning, Medium) | 100% | 100% | 100% | 100% | 100% |
| Z.AI GLM 5.3 (Reasoning, Max) | 100% | 100% | 100% | 100% | 100% |
| GPT-5.6 Sol (Reasoning, Medium) | 100% | 100% | 100% | 100% | 100% |
| Claude Opus 4.6 (Reasoning) | 100% | 100% | 100% | 100% | 100% |
| Gemini 3.6 Flash (Reasoning) | 100% | 100% | 100% | 100% | 100% |
| Gemini 3.7 Flash (Reasoning, Medium) | 100% | 100% | 100% | 100% | 100% |
| Hy4 Preview (Reasoning, High) | 100% | 100% | 100% | 100% | 100% |
| Muse Spark 1.2 (Reasoning, Medium) | 100% | 100% | 100% | 100% | 100% |
| Qwen 3.7 Max | 100% | 100% | 100% | 100% | 100% |
| Z.AI GLM 5.3 Flash (Reasoning, Max) | 100% | 100% | 100% | 100% | 100% |
| Qwen 3.8 Flash (Reasoning) | 100% | 100% | 100% | 100% | 100% |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | 100% | 100% |
| GPT-6 Luna (Reasoning, High) | 100% | 100% | 100% | 100% | 100% |
Friend got new kittens (Tagalog)
Performance Score Distribution (Top 20)
Click a model name to view its detail page.
Price-Performance Score Distribution (Top 20)
Click a model name to view its detail page.
| Score | Cost | Time | ||
|---|---|---|---|---|
| Mistral NeMO | 100% | $0.0000 | 515ms | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 505ms | |
| Gemini 2.5 Flash | 100% | $0.0001 | 518ms | |
| Ministral 3 3B | 100% | $0.0000 | 715ms | |
| Gemma 3 4B | 100% | $0.0000 | 903ms | |
| GPT-4o Mini (temp=1) | 100% | $0.0000 | 883ms | |
| GPT-5.4 Nano | 100% | $0.0001 | 759ms | |
| GPT-5.4 Nano (Reasoning, Low) | 100% | $0.0001 | 955ms | |
| Ministral 3 8B | 100% | $0.0000 | 1.2s | |
| Gemma 4 26B | 100% | $0.0000 | 4.0s | |
| Mistral Small 4 | 100% | $0.0001 | 997ms | |
| GPT-4.1 Nano | 100% | $0.0000 | 1.4s | |
| GPT-4o Mini (temp=0) | 100% | $0.0000 | 1.2s | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0002 | 560ms | |
| GPT-6 Luna | 100% | $0.0000 | 1.2s | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0001 | 642ms | |
| Ministral 3 14B | 100% | $0.0000 | 1.3s | |
| Gemini 3.1 Flash Lite | 100% | $0.0001 | 940ms | |
| GPT-5.4 Mini | 100% | $0.0002 | 597ms | |
| Inception Mercury 2 | 100% | $0.0002 | 614ms | |
Most Stable Models (Top 20)
Ranked by stability (median × consistency). Click a model name to view its detail page.
| Score | Consistency | Stability | ||
|---|---|---|---|---|
| Claude Opus 5.5 (Reasoning) | 100% | 100% | 100% | |
| Qwen 3.8 Max (Reasoning, XHigh) | 100% | 100% | 100% | |
| Gemini 3.8 Flash (Reasoning, Medium) | 100% | 100% | 100% | |
| Z.AI GLM 5.3 (Reasoning, Max) | 100% | 100% | 100% | |
| GPT-5.6 Sol (Reasoning, Medium) | 100% | 100% | 100% | |
| Claude Opus 4.6 (Reasoning) | 100% | 100% | 100% | |
| Gemini 3.6 Flash (Reasoning) | 100% | 100% | 100% | |
| Gemini 3.7 Flash (Reasoning, Medium) | 100% | 100% | 100% | |
| Hy4 Preview (Reasoning, High) | 100% | 100% | 100% | |
| Muse Spark 1.2 (Reasoning, Medium) | 100% | 100% | 100% | |
| Qwen 3.7 Max | 100% | 100% | 100% | |
| DeepSeek V4.1 Flash (Reasoning, High) | 100% | 100% | 100% | |
| Z.AI GLM 5.3 Flash (Reasoning, Max) | 100% | 100% | 100% | |
| Qwen 3.8 Flash (Reasoning) | 100% | 100% | 100% | |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.1 Pro (Preview) | 100% | 100% | 100% | |
| Muse Spark 1.3 (Reasoning, Medium) | 100% | 100% | 100% | |
| GPT-6 Luna (Reasoning, High) | 100% | 100% | 100% | |
| GPT-5.4 (Reasoning) | 100% | 100% | 100% | |
| Muse Spark 1.1 (Reasoning, Medium) | 100% | 100% | 100% | |
Top Overall Models (Top 20)
Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.
| Score | Cost | Speed | Stability | ||
|---|---|---|---|---|---|
| Mistral NeMO | 100% | $0.0000 | 515ms | 100% | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 505ms | 100% | |
| Ministral 3 3B | 100% | $0.0000 | 715ms | 100% | |
| Gemini 2.5 Flash | 100% | $0.0001 | 518ms | 100% | |
| Gemma 3 4B | 100% | $0.0000 | 903ms | 100% | |
| GPT-5.4 Nano | 100% | $0.0001 | 759ms | 100% | |
| GPT-4o Mini (temp=1) | 100% | $0.0000 | 883ms | 100% | |
| GPT-5.4 Nano (Reasoning, Low) | 100% | $0.0001 | 955ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0001 | 642ms | 100% | |
| GPT-6 Luna | 100% | $0.0000 | 1.2s | 100% | |
| Mistral Small 4 | 100% | $0.0001 | 997ms | 100% | |
| Ministral 3 8B | 100% | $0.0000 | 1.2s | 100% | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0002 | 560ms | 100% | |
| GPT-4o Mini (temp=0) | 100% | $0.0000 | 1.2s | 100% | |
| Ministral 3 14B | 100% | $0.0000 | 1.3s | 100% | |
| GPT-4.1 Nano | 100% | $0.0000 | 1.4s | 100% | |
| Inception Mercury 2 | 100% | $0.0002 | 614ms | 100% | |
| Gemini 3.1 Flash Lite | 100% | $0.0001 | 940ms | 100% | |
| GPT-5.4 Mini | 100% | $0.0002 | 597ms | 100% | |
| Gemini 3.1 Flash Lite (Reasoning) | 100% | $0.0002 | 981ms | 100% | |
| Median | Evaluator | Top 3 | Flop 3 |
|---|---|---|---|
| 100.0% | Contains a count of nouns |
Friend got new kittens (German)
Performance Score Distribution (Top 20)
Click a model name to view its detail page.
Price-Performance Score Distribution (Top 20)
Click a model name to view its detail page.
| Score | Cost | Time | ||
|---|---|---|---|---|
| Ministral 8B | 80% | $0.0000 | 733ms | |
| Ministral 3 3B | 100% | $0.0000 | 995ms | |
| Ministral 3 8B | 100% | $0.0000 | 1.0s | |
| GPT-6 Luna | 100% | $0.0000 | 1.6s | |
| Mistral NeMO | 100% | $0.0000 | 1.2s | |
| Ministral 3 14B | 100% | $0.0000 | 1.6s | |
| Mistral Small 4 | 100% | $0.0001 | 1.2s | |
| Gemini 3.1 Flash Lite (Reasoning) | 80% | $0.0001 | 827ms | |
| Gemini 3.1 Flash Lite | 60% | $0.0001 | 725ms | |
| Gemini 3.1 Flash Lite (Preview) | 80% | $0.0001 | 841ms | |
| GPT-5.4 Nano | 100% | $0.0001 | 989ms | |
| Mistral Small 3.2 24B | 100% | $0.0001 | 2.3s | |
| GPT-4.1 Nano | 100% | $0.0000 | 2.4s | |
| Cydonia 24B V4.1 | 100% | $0.0001 | 2.0s | |
| Gemma 3 12B | 80% | $0.0000 | 3.6s | |
| Hermes 3 70B | 100% | $0.0001 | 3.1s | |
| GPT-4.1 Mini | 100% | $0.0001 | 2.0s | |
| GPT-6 Luna (Reasoning, Medium) | 100% | $0.0001 | 3.3s | |
| GPT-5.4 Mini | 80% | $0.0002 | 896ms | |
| GPT-6 Luna (Reasoning, High) | 100% | $0.0001 | 3.8s | |
Most Stable Models (Top 20)
Ranked by stability (median × consistency). Click a model name to view its detail page.
| Score | Consistency | Stability | ||
|---|---|---|---|---|
| Claude Opus 5.5 (Reasoning) | 100% | 100% | 100% | |
| Qwen 3.8 Max (Reasoning, XHigh) | 100% | 100% | 100% | |
| Gemini 3.8 Flash (Reasoning, Medium) | 100% | 100% | 100% | |
| Z.AI GLM 5.3 (Reasoning, Max) | 100% | 100% | 100% | |
| GPT-5.6 Sol (Reasoning, Medium) | 100% | 100% | 100% | |
| Claude Opus 4.6 (Reasoning) | 100% | 100% | 100% | |
| Gemini 3.6 Flash (Reasoning) | 100% | 100% | 100% | |
| Gemini 3.7 Flash (Reasoning, Medium) | 100% | 100% | 100% | |
| Hy4 Preview (Reasoning, High) | 100% | 100% | 100% | |
| Muse Spark 1.2 (Reasoning, Medium) | 100% | 100% | 100% | |
| Qwen 3.7 Max | 100% | 100% | 100% | |
| DeepSeek V4.1 Flash (Reasoning, High) | 100% | 100% | 100% | |
| Z.AI GLM 5.3 Flash (Reasoning, Max) | 100% | 100% | 100% | |
| Qwen 3.8 Flash (Reasoning) | 100% | 100% | 100% | |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | |
| GPT-6 Luna (Reasoning, High) | 100% | 100% | 100% | |
| Qwen 3.6 Max Preview | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning, Medium) | 100% | 100% | 100% | |
| Claude Sonnet 4.6 (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5.2 (Reasoning, High) | 100% | 100% | 100% | |
Top Overall Models (Top 20)
Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.
| Score | Cost | Speed | Stability | ||
|---|---|---|---|---|---|
| Ministral 3 3B | 100% | $0.0000 | 995ms | 100% | |
| Ministral 3 8B | 100% | $0.0000 | 1.0s | 100% | |
| Mistral NeMO | 100% | $0.0000 | 1.2s | 100% | |
| GPT-5.4 Nano | 100% | $0.0001 | 989ms | 100% | |
| Mistral Small 4 | 100% | $0.0001 | 1.2s | 100% | |
| GPT-6 Luna | 100% | $0.0000 | 1.6s | 100% | |
| Ministral 3 14B | 100% | $0.0000 | 1.6s | 100% | |
| Cydonia 24B V4.1 | 100% | $0.0001 | 2.0s | 100% | |
| Mistral Small 3.2 24B | 100% | $0.0001 | 2.3s | 100% | |
| GPT-4.1 Nano | 100% | $0.0000 | 2.4s | 100% | |
| GPT-4.1 Mini | 100% | $0.0001 | 2.0s | 100% | |
| Mistral Medium 3.1 | 100% | $0.0003 | 1.6s | 100% | |
| GPT-5.6 Luna | 100% | $0.0003 | 1.3s | 100% | |
| Hermes 3 70B | 100% | $0.0001 | 3.1s | 100% | |
| GPT-6 Luna (Reasoning, Medium) | 100% | $0.0001 | 3.3s | 100% | |
| GPT-6 Sol | 100% | $0.0004 | 1.8s | 100% | |
| Mistral Large 3 | 100% | $0.0002 | 2.7s | 100% | |
| GPT-6 Luna (Reasoning, High) | 100% | $0.0001 | 3.8s | 100% | |
| Qwen 2.5 72B | 100% | $0.0001 | 3.9s | 100% | |
| GPT-4o, Aug. 6th (temp=1) | 100% | $0.0006 | 1.3s | 100% | |
| Median | Evaluator | Top 3 | Flop 3 |
|---|---|---|---|
| 100.0% | Contains a count of nouns |
Asking for directions (German)
Performance Score Distribution (Top 20)
Click a model name to view its detail page.
Price-Performance Score Distribution (Top 20)
Click a model name to view its detail page.
| Score | Cost | Time | ||
|---|---|---|---|---|
| Mistral NeMO | 100% | $0.0000 | 321ms | |
| Ministral 8B | 60% | $0.0000 | 341ms | |
| Ministral 3 3B | 100% | $0.0000 | 691ms | |
| GPT-4o Mini (temp=0) | 100% | $0.0000 | 728ms | |
| GPT-4o Mini (temp=1) | 100% | $0.0000 | 800ms | |
| Mistral Large 3 | 100% | $0.0001 | 7.5s | |
| Gemini 2.5 Flash Lite | 100% | $0.0001 | 722ms | |
| Mistral Small 3.2 24B | 100% | $0.0000 | 1.6s | |
| Gemma 4 26B | 100% | $0.0000 | 11.1s | |
| GPT-5.4 Mini | 80% | $0.0002 | 812ms | |
| GPT-6 Luna | 80% | $0.0000 | 2.3s | |
| Gemma 3 4B | 100% | $0.0000 | 2.4s | |
| Grok 4.3 | 60% | $0.0002 | 740ms | |
| Gemini 3.1 Flash Lite | 80% | $0.0002 | 922ms | |
| Gemini 3.1 Flash Lite (Reasoning) | 100% | $0.0002 | 3.5s | |
| Cydonia 24B V4.1 | 100% | $0.0001 | 2.1s | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0003 | 693ms | |
| DeepSeek V3 (2025-03-24) | 100% | $0.0001 | 7.5s | |
| Gemini 3.1 Flash Lite (Preview) | 100% | $0.0002 | 995ms | |
| Hermes 3 70B | 80% | $0.0001 | 3.2s | |
Most Stable Models (Top 20)
Ranked by stability (median × consistency). Click a model name to view its detail page.
| Score | Consistency | Stability | ||
|---|---|---|---|---|
| Claude Opus 5.5 (Reasoning) | 100% | 100% | 100% | |
| Qwen 3.8 Max (Reasoning, XHigh) | 100% | 100% | 100% | |
| Gemini 3.8 Flash (Reasoning, Medium) | 100% | 100% | 100% | |
| Z.AI GLM 5.3 (Reasoning, Max) | 100% | 100% | 100% | |
| GPT-5.6 Sol (Reasoning, Medium) | 100% | 100% | 100% | |
| Claude Opus 4.6 (Reasoning) | 100% | 100% | 100% | |
| Gemini 3.6 Flash (Reasoning) | 100% | 100% | 100% | |
| Gemini 3.7 Flash (Reasoning, Medium) | 100% | 100% | 100% | |
| Hy4 Preview (Reasoning, High) | 100% | 100% | 100% | |
| Muse Spark 1.2 (Reasoning, Medium) | 100% | 100% | 100% | |
| Qwen 3.7 Max | 100% | 100% | 100% | |
| DeepSeek V4.1 Flash (Reasoning, High) | 100% | 100% | 100% | |
| Z.AI GLM 5.3 Flash (Reasoning, Max) | 100% | 100% | 100% | |
| Qwen 3.8 Flash (Reasoning) | 100% | 100% | 100% | |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.1 Pro (Preview) | 100% | 100% | 100% | |
| GPT-6 Luna (Reasoning, High) | 100% | 100% | 100% | |
| GPT-5.4 (Reasoning) | 100% | 100% | 100% | |
| Muse Spark 1.1 (Reasoning, Medium) | 100% | 100% | 100% | |
| Z.AI GLM 5.1 | 100% | 100% | 100% | |
Top Overall Models (Top 20)
Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.
| Score | Cost | Speed | Stability | ||
|---|---|---|---|---|---|
| Mistral NeMO | 100% | $0.0000 | 321ms | 100% | |
| Ministral 3 3B | 100% | $0.0000 | 691ms | 100% | |
| GPT-4o Mini (temp=0) | 100% | $0.0000 | 728ms | 100% | |
| GPT-4o Mini (temp=1) | 100% | $0.0000 | 800ms | 100% | |
| Gemini 2.5 Flash Lite | 100% | $0.0001 | 722ms | 100% | |
| Mistral Small 3.2 24B | 100% | $0.0000 | 1.6s | 100% | |
| Gemini 3.1 Flash Lite (Preview) | 100% | $0.0002 | 995ms | 100% | |
| Gemma 3 4B | 100% | $0.0000 | 2.4s | 100% | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0003 | 693ms | 100% | |
| Cydonia 24B V4.1 | 100% | $0.0001 | 2.1s | 100% | |
| GPT-5.6 Luna | 100% | $0.0003 | 709ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0004 | 740ms | 100% | |
| Gemini 3 Flash (Preview) | 100% | $0.0003 | 1.3s | 100% | |
| DeepSeek V3.1 | 100% | $0.0002 | 2.4s | 100% | |
| DeepSeek V3 (2024-12-26) | 100% | $0.0001 | 3.2s | 100% | |
| Llama 3.1 70B | 100% | $0.0002 | 3.3s | 100% | |
| GPT-6 Sol | 100% | $0.0004 | 1.9s | 100% | |
| Gemini 3.1 Flash Lite (Reasoning) | 100% | $0.0002 | 3.5s | 100% | |
| DeepSeek-V2 Chat | 100% | $0.0000 | 4.6s | 100% | |
| WizardLM 2 8x22b | 100% | $0.0002 | 3.8s | 100% | |
| Median | Evaluator | Top 3 | Flop 3 |
|---|---|---|---|
| 100.0% | Matches Regex |
Asking for directions (Dutch)
Performance Score Distribution (Top 20)
Click a model name to view its detail page.
Price-Performance Score Distribution (Top 20)
Click a model name to view its detail page.
| Score | Cost | Time | ||
|---|---|---|---|---|
| Mistral NeMO | 80% | $0.0000 | 440ms | |
| Ministral 3 3B | 100% | $0.0000 | 930ms | |
| Gemini 2.5 Flash Lite | 80% | $0.0001 | 735ms | |
| GPT-6 Luna | 80% | $0.0000 | 1.3s | |
| Mistral Large 3 | 100% | $0.0001 | 984ms | |
| Gemini 3.1 Flash Lite (Preview) | 100% | $0.0001 | 954ms | |
| GPT-5.4 Mini | 60% | $0.0001 | 685ms | |
| Gemma 4 31B | 100% | $0.0000 | 2.2s | |
| Gemma 3 4B | 60% | $0.0000 | 2.0s | |
| Gemini 3.1 Flash Lite | 100% | $0.0002 | 1.0s | |
| Gemini 3.1 Flash Lite (Reasoning) | 100% | $0.0002 | 1.1s | |
| Thinking Machines Inkling | 80% | $0.0003 | 1.7s | |
| Gemma 4 26B | 100% | $0.0000 | 7.2s | |
| Grok 4.3 | 80% | $0.0003 | 1.1s | |
| Z.AI GLM 4.5 | 100% | $0.0001 | 2.4s | |
| GPT-5.6 Luna | 100% | $0.0003 | 1.0s | |
| Gemini 2.5 Flash | 80% | $0.0002 | 796ms | |
| Gemini 3 Flash (Preview) | 80% | $0.0003 | 1.2s | |
| Z.AI GLM 5.3 Flash (Reasoning, Low) | 80% | $0.0005 | 48.9s | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0004 | 780ms | |
Most Stable Models (Top 20)
Ranked by stability (median × consistency). Click a model name to view its detail page.
| Score | Consistency | Stability | ||
|---|---|---|---|---|
| Claude Opus 5.5 (Reasoning) | 100% | 100% | 100% | |
| Qwen 3.8 Max (Reasoning, XHigh) | 100% | 100% | 100% | |
| Gemini 3.8 Flash (Reasoning, Medium) | 100% | 100% | 100% | |
| Z.AI GLM 5.3 (Reasoning, Max) | 100% | 100% | 100% | |
| GPT-5.6 Sol (Reasoning, Medium) | 100% | 100% | 100% | |
| Claude Opus 4.6 (Reasoning) | 100% | 100% | 100% | |
| Gemini 3.6 Flash (Reasoning) | 100% | 100% | 100% | |
| Gemini 3.7 Flash (Reasoning, Medium) | 100% | 100% | 100% | |
| Hy4 Preview (Reasoning, High) | 100% | 100% | 100% | |
| Muse Spark 1.2 (Reasoning, Medium) | 100% | 100% | 100% | |
| Qwen 3.7 Max | 100% | 100% | 100% | |
| Z.AI GLM 5.3 Flash (Reasoning, Max) | 100% | 100% | 100% | |
| Qwen 3.8 Flash (Reasoning) | 100% | 100% | 100% | |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.1 Pro (Preview) | 100% | 100% | 100% | |
| Muse Spark 1.3 (Reasoning, Medium) | 100% | 100% | 100% | |
| GPT-6 Luna (Reasoning, High) | 100% | 100% | 100% | |
| GPT-5.4 (Reasoning) | 100% | 100% | 100% | |
| Muse Spark 1.1 (Reasoning, Medium) | 100% | 100% | 100% | |
| Qwen 3.6 Max Preview | 100% | 100% | 100% | |
Top Overall Models (Top 20)
Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.
| Score | Cost | Speed | Stability | ||
|---|---|---|---|---|---|
| Ministral 3 3B | 100% | $0.0000 | 930ms | 100% | |
| Mistral Large 3 | 100% | $0.0001 | 984ms | 100% | |
| Gemini 3.1 Flash Lite (Preview) | 100% | $0.0001 | 954ms | 100% | |
| Gemma 4 31B | 100% | $0.0000 | 2.2s | 100% | |
| Gemini 3.1 Flash Lite (Reasoning) | 100% | $0.0002 | 1.1s | 100% | |
| Gemini 3.1 Flash Lite | 100% | $0.0002 | 1.0s | 100% | |
| GPT-5.6 Luna | 100% | $0.0003 | 1.0s | 100% | |
| Z.AI GLM 4.5 | 100% | $0.0001 | 2.4s | 100% | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0004 | 780ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0005 | 887ms | 100% | |
| GPT-6 Sol | 100% | $0.0004 | 1.7s | 100% | |
| Laguna S 2.1 | 100% | $0.0001 | 4.9s | 100% | |
| GPT-5.6 Terra | 100% | $0.0007 | 1.1s | 100% | |
| DeepSeek V3 (2024-12-26) | 100% | $0.0002 | 5.9s | 100% | |
| GPT-6 Luna (Reasoning, Medium) | 100% | $0.0002 | 6.0s | 100% | |
| Gemma 4 26B | 100% | $0.0000 | 7.2s | 100% | |
| Qwen 3.5 Plus (2026-02-15) | 100% | $0.0003 | 5.5s | 100% | |
| GPT-6 Luna (Reasoning, High) | 100% | $0.0003 | 7.2s | 100% | |
| Gemini 3.6 Flash (Reasoning, Minimal) | 100% | $0.0010 | 1.1s | 100% | |
| GPT-5.4 Nano (Reasoning) | 100% | $0.0005 | 6.1s | 100% | |
| Median | Evaluator | Top 3 | Flop 3 |
|---|---|---|---|
| 100.0% | Matches Regex |