Data extraction
Extract key details from a given block of text.
Performance Score Distribution (Top 20)
Click a model name to view its detail page.
Price-Performance Score Distribution (Top 20)
Click a model name to view its detail page.
| Score | Cost | Time | ||
|---|---|---|---|---|
| Gemma 3 4B | 92% | $0.0000 | 303ms | |
| Ministral 3B | 73% | $0.0000 | 308ms | |
| Ministral 8B | 75% | $0.0000 | 331ms | |
| Gemini 2.5 Flash Lite | 92% | $0.0000 | 357ms | |
| Ministral 3 3B | 78% | $0.0000 | 416ms | |
| Ministral 3 14B | 88% | $0.0000 | 448ms | |
| Gemma 3 12B | 92% | $0.0000 | 542ms | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0000 | 395ms | |
| Ministral 3 8B | 71% | $0.0000 | 382ms | |
| Mistral Small 3.2 24B | 83% | $0.0000 | 691ms | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0000 | 462ms | |
| Mistral Small 4 | 88% | $0.0000 | 539ms | |
| Gemini 2.5 Flash | 83% | $0.0000 | 473ms | |
| Gemma 3 27B | 92% | $0.0000 | 780ms | |
| Cydonia 24B V4.1 | 81% | $0.0000 | 593ms | |
| GPT-5.4 Nano | 93% | $0.0000 | 768ms | |
| Mistral Medium 3.1 | 88% | $0.0000 | 655ms | |
| Gemini 3.1 Flash Lite (Reasoning) | 91% | $0.0000 | 1.2s | |
| Mistral Large 3 | 88% | $0.0000 | 945ms | |
| Gemma 4 31B | 97% | $0.0000 | 5.1s | |
Cost vs Performance
Compares total cost for this test against the test score. Quadrant lines are drawn at the median values. Only models with available cost data are shown.
12 low-scoring outliers hidden: Grok 4.20 (88.6%), Arcee AI: Trinity Mini (88.6%), Cydonia 24B V4.1 (87.3%), Ministral 3 3B (85.5%), Ministral 8B (80.9%), WizardLM 2 8x22b (79.1%), DeepSeek V4 Flash (77.7%), Ministral 3B (77.7%), Ministral 3 8B (77.3%), Thinking Machines Inkling (75.0%), Grok 4.20 (Reasoning) (72.3%), Cohere Command R+ (Aug. 2024) (68.2%).
Most Stable Models (Top 20)
Ranked by stability (median × consistency). Click a model name to view its detail page.
| Score | Consistency | Stability | ||
|---|---|---|---|---|
| GPT-5.6 Sol (Reasoning) | 100% | 100% | 100% | |
| Gemini 3.6 Flash (Reasoning) | 100% | 100% | 100% | |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | |
| Z.AI GLM 5.2 (Reasoning, High) | 100% | 100% | 100% | |
| MoonshotAI: Kimi K3 (Reasoning, High) | 100% | 100% | 100% | |
| GPT-5.6 Terra (Reasoning) | 100% | 100% | 100% | |
| MoonshotAI: Kimi K3 (Reasoning, Low) | 100% | 100% | 100% | |
| GPT-5.6 Sol | 100% | 100% | 100% | |
| GPT-5.6 Luna (Reasoning) | 100% | 100% | 100% | |
| GPT-5.6 Terra | 100% | 100% | 100% | |
| Gemini 3.6 Flash (Reasoning, Minimal) | 100% | 100% | 100% | |
| GPT-5.6 Luna | 100% | 100% | 100% | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | 100% | 100% | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | 100% | 100% | |
| Grok 4.5 (Reasoning, Low) | 100% | 91% | 91% | |
| Gemini 3.5 Flash (Reasoning) | 99% | 82% | 82% | |
| Gemini 3 Flash (Preview, Reasoning) | 99% | 82% | 82% | |
| Aion 3.0 Mini | 98% | 75% | 75% | |
| Gemma 4 26B (Reasoning) | 98% | 74% | 74% | |
| Claude Sonnet 4 | 96% | 72% | 72% | |
Top Overall Models (Top 20)
Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.
| Score | Cost | Speed | Stability | ||
|---|---|---|---|---|---|
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0000 | 395ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0000 | 462ms | 100% | |
| GPT-5.6 Luna | 100% | $0.0001 | 698ms | 100% | |
| Gemini 3.6 Flash (Reasoning, Minimal) | 100% | $0.0002 | 693ms | 100% | |
| GPT-5.6 Terra | 100% | $0.0003 | 800ms | 100% | |
| GPT-5.6 Terra (Reasoning) | 100% | $0.0004 | 1.1s | 100% | |
| GPT-5.6 Luna (Reasoning) | 100% | $0.0003 | 1.7s | 100% | |
| GPT-5.6 Sol | 100% | $0.0006 | 1.1s | 100% | |
| GPT-5.6 Sol (Reasoning) | 100% | $0.0007 | 1.2s | 100% | |
| MoonshotAI: Kimi K3 (Reasoning, Low) | 100% | $0.0013 | 4.2s | 100% | |
| Gemini 3.6 Flash (Reasoning) | 100% | $0.0023 | 1.8s | 100% | |
| Z.AI GLM 5.2 (Reasoning, High) | 100% | $0.0012 | 5.7s | 100% | |
| MoonshotAI: Kimi K3 (Reasoning, High) | 100% | $0.0017 | 4.9s | 100% | |
| Grok 4.5 (Reasoning, High) | 100% | $0.0018 | 5.6s | 100% | |
| Grok 4.5 (Reasoning, Low) | 100% | $0.0014 | 6.2s | 91% | |
| Gemini 3 Flash (Preview, Reasoning) | 99% | $0.0026 | 7.0s | 82% | |
| Claude Sonnet 4 | 96% | $0.0004 | 1.6s | 72% | |
| Laguna XS 2.1 | 95% | $0.0000 | 2.4s | 71% | |
| Aion 3.0 Mini | 98% | $0.0005 | 8.5s | 75% | |
| Claude Sonnet 5 (Reasoning, Low) | 95% | $0.0005 | 3.4s | 71% | |
| Model | Total â–¼ | Who's the tallest? | What's the color of the car? | What instrument does Lucy play? | Guess the pet | Who's the sister? | Contextual pronoun | Indirect birth year | Fruits excluding citrus | Future event time | Highest-rated movie | All valid emails |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-5.6 Sol (Reasoning) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| Gemini 3.6 Flash (Reasoning) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| Qwen3.7 Max | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| Gemini 3.1 Pro (Preview) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| GPT-5.4 (Reasoning) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| Z.AI GLM 5.1 | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| Qwen3.6 Max Preview | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| GPT-5.5 (Reasoning) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| Z.AI GLM 5.2 (Reasoning, High) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| Gemini 3.5 Flash (Reasoning) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| MoonshotAI: Kimi K3 (Reasoning, High) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| MoonshotAI: Kimi K2.6 | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| Claude Opus 4.7 (Reasoning) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| GPT-5.5 (Reasoning, Low) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
Who's the tallest?
Performance Score Distribution (Top 20)
Click a model name to view its detail page.
Price-Performance Score Distribution (Top 20)
Click a model name to view its detail page.
| Score | Cost | Time | ||
|---|---|---|---|---|
| Gemma 3 4B | 100% | $0.0000 | 217ms | |
| Ministral 3B | 100% | $0.0000 | 269ms | |
| Gemma 3 12B | 100% | $0.0000 | 322ms | |
| Ministral 8B | 90% | $0.0000 | 261ms | |
| Ministral 3 3B | 100% | $0.0000 | 272ms | |
| Ministral 3 8B | 100% | $0.0000 | 368ms | |
| Mistral Small 3.2 24B | 100% | $0.0000 | 363ms | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 389ms | |
| Ministral 3 14B | 100% | $0.0000 | 392ms | |
| Mistral NeMO | 80% | $0.0000 | 318ms | |
| Gemma 3 27B | 100% | $0.0000 | 467ms | |
| Mistral Small 4 | 100% | $0.0000 | 463ms | |
| DeepSeek V3 (2024-12-26) | 100% | $0.0000 | 594ms | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0000 | 377ms | |
| Gemma 4 26B | 100% | $0.0000 | 1.8s | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0000 | 407ms | |
| Gemini 2.5 Flash | 100% | $0.0000 | 399ms | |
| DeepSeek V4 Flash | 100% | $0.0000 | 2.4s | |
| GPT-5.4 Nano | 100% | $0.0000 | 579ms | |
| Gemini 3.1 Flash Lite | 100% | $0.0000 | 742ms | |
Most Stable Models (Top 20)
Ranked by stability (median × consistency). Click a model name to view its detail page.
| Score | Consistency | Stability | ||
|---|---|---|---|---|
| GPT-5.6 Sol (Reasoning) | 100% | 100% | 100% | |
| Claude Opus 4.6 (Reasoning) | 100% | 100% | 100% | |
| Gemini 3.6 Flash (Reasoning) | 100% | 100% | 100% | |
| Qwen3.7 Max | 100% | 100% | 100% | |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.1 Pro (Preview) | 100% | 100% | 100% | |
| GPT-5.4 (Reasoning) | 100% | 100% | 100% | |
| Muse Spark 1.1 (Reasoning, Medium) | 100% | 100% | 100% | |
| Z.AI GLM 5.1 | 100% | 100% | 100% | |
| Qwen3.6 Max Preview | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning) | 100% | 100% | 100% | |
| Claude Sonnet 4.6 (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5.2 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.5 Flash (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5 Turbo | 100% | 100% | 100% | |
| MoonshotAI: Kimi K3 (Reasoning, High) | 100% | 100% | 100% | |
| MoonshotAI: Kimi K2.6 | 100% | 100% | 100% | |
| Claude Opus 4.7 (Reasoning) | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning, Low) | 100% | 100% | 100% | |
| GPT-5.6 Terra (Reasoning) | 100% | 100% | 100% | |
Top Overall Models (Top 20)
Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.
| Score | Cost | Speed | Stability | ||
|---|---|---|---|---|---|
| Gemma 3 4B | 100% | $0.0000 | 217ms | 100% | |
| Ministral 3B | 100% | $0.0000 | 269ms | 100% | |
| Ministral 3 3B | 100% | $0.0000 | 272ms | 100% | |
| Gemma 3 12B | 100% | $0.0000 | 322ms | 100% | |
| Mistral Small 3.2 24B | 100% | $0.0000 | 363ms | 100% | |
| Ministral 3 8B | 100% | $0.0000 | 368ms | 100% | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 389ms | 100% | |
| Ministral 3 14B | 100% | $0.0000 | 392ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0000 | 377ms | 100% | |
| Gemini 2.5 Flash | 100% | $0.0000 | 399ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0000 | 407ms | 100% | |
| Gemma 3 27B | 100% | $0.0000 | 467ms | 100% | |
| Mistral Small 4 | 100% | $0.0000 | 463ms | 100% | |
| GPT-5.4 Nano | 100% | $0.0000 | 579ms | 100% | |
| DeepSeek V3 (2024-12-26) | 100% | $0.0000 | 594ms | 100% | |
| Llama 3.1 70B | 100% | $0.0000 | 419ms | 100% | |
| Qwen3 235B A22B Instruct 2507 | 100% | $0.0000 | 658ms | 100% | |
| Qwen 2.5 72B | 100% | $0.0000 | 596ms | 100% | |
| Gemini 3.1 Flash Lite | 100% | $0.0000 | 742ms | 100% | |
| Gemini 3.1 Flash Lite (Preview) | 100% | $0.0000 | 746ms | 100% | |
| Median | Evaluator | Top 3 | Flop 3 |
|---|---|---|---|
| 100.0% | Matches Regex | ||
| 100.0% | Matches text |
What's the color of the car?
Performance Score Distribution (Top 20)
Click a model name to view its detail page.
Price-Performance Score Distribution (Top 20)
Click a model name to view its detail page.
| Score | Cost | Time | ||
|---|---|---|---|---|
| Gemma 3 4B | 100% | $0.0000 | 221ms | |
| Ministral 3B | 100% | $0.0000 | 273ms | |
| Mistral NeMO | 100% | $0.0000 | 536ms | |
| Ministral 8B | 100% | $0.0000 | 256ms | |
| Ministral 3 3B | 100% | $0.0000 | 293ms | |
| Qwen3 235B A22B Instruct 2507 | 100% | $0.0000 | 758ms | |
| Ministral 3 8B | 100% | $0.0000 | 361ms | |
| Mistral Small 3.2 24B | 100% | $0.0000 | 406ms | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 375ms | |
| Ministral 3 14B | 100% | $0.0000 | 362ms | |
| Gemma 4 26B | 100% | $0.0000 | 1.5s | |
| Gemma 3 12B | 100% | $0.0000 | 403ms | |
| Gemma 4 31B | 100% | $0.0000 | 2.0s | |
| Gemma 3 27B | 100% | $0.0000 | 518ms | |
| Mistral Small 4 | 100% | $0.0000 | 488ms | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0000 | 414ms | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0000 | 397ms | |
| Mistral Medium 3.1 | 100% | $0.0000 | 441ms | |
| Cydonia 24B V4.1 | 100% | $0.0000 | 300ms | |
| Gemini 2.5 Flash | 100% | $0.0000 | 549ms | |
Most Stable Models (Top 20)
Ranked by stability (median × consistency). Click a model name to view its detail page.
| Score | Consistency | Stability | ||
|---|---|---|---|---|
| GPT-5.6 Sol (Reasoning) | 100% | 100% | 100% | |
| Claude Opus 4.6 (Reasoning) | 100% | 100% | 100% | |
| Gemini 3.6 Flash (Reasoning) | 100% | 100% | 100% | |
| Qwen3.7 Max | 100% | 100% | 100% | |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.1 Pro (Preview) | 100% | 100% | 100% | |
| GPT-5.4 (Reasoning) | 100% | 100% | 100% | |
| Muse Spark 1.1 (Reasoning, Medium) | 100% | 100% | 100% | |
| Z.AI GLM 5.1 | 100% | 100% | 100% | |
| Qwen3.6 Max Preview | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning) | 100% | 100% | 100% | |
| Claude Sonnet 4.6 (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5.2 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.5 Flash (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5 Turbo | 100% | 100% | 100% | |
| MoonshotAI: Kimi K3 (Reasoning, High) | 100% | 100% | 100% | |
| MoonshotAI: Kimi K2.6 | 100% | 100% | 100% | |
| Claude Opus 4.7 (Reasoning) | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning, Low) | 100% | 100% | 100% | |
| GPT-5.6 Terra (Reasoning) | 100% | 100% | 100% | |
Top Overall Models (Top 20)
Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.
| Score | Cost | Speed | Stability | ||
|---|---|---|---|---|---|
| Gemma 3 4B | 100% | $0.0000 | 221ms | 100% | |
| Ministral 3B | 100% | $0.0000 | 273ms | 100% | |
| Ministral 8B | 100% | $0.0000 | 256ms | 100% | |
| Ministral 3 3B | 100% | $0.0000 | 293ms | 100% | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 375ms | 100% | |
| Ministral 3 8B | 100% | $0.0000 | 361ms | 100% | |
| Ministral 3 14B | 100% | $0.0000 | 362ms | 100% | |
| Gemma 3 12B | 100% | $0.0000 | 403ms | 100% | |
| Mistral Small 3.2 24B | 100% | $0.0000 | 406ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0000 | 397ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0000 | 414ms | 100% | |
| Cydonia 24B V4.1 | 100% | $0.0000 | 300ms | 100% | |
| Mistral Small 4 | 100% | $0.0000 | 488ms | 100% | |
| Gemma 3 27B | 100% | $0.0000 | 518ms | 100% | |
| Mistral NeMO | 100% | $0.0000 | 536ms | 100% | |
| Mistral Medium 3.1 | 100% | $0.0000 | 441ms | 100% | |
| Gemini 2.5 Flash | 100% | $0.0000 | 549ms | 100% | |
| GPT-5.4 Nano | 100% | $0.0000 | 578ms | 100% | |
| Gemini 3.1 Flash Lite (Reasoning) | 100% | $0.0000 | 610ms | 100% | |
| Gemini 3.1 Flash Lite (Preview) | 100% | $0.0000 | 638ms | 100% | |
| Median | Evaluator | Top 3 | Flop 3 |
|---|---|---|---|
| 100.0% | Matches Regex | ||
| 100.0% | Matches text |
What instrument does Lucy play?
Performance Score Distribution (Top 20)
Click a model name to view its detail page.
Price-Performance Score Distribution (Top 20)
Click a model name to view its detail page.
| Score | Cost | Time | ||
|---|---|---|---|---|
| Gemma 3 4B | 100% | $0.0000 | 217ms | |
| Mistral NeMO | 100% | $0.0000 | 577ms | |
| Mistral Small 3.2 24B | 100% | $0.0000 | 335ms | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 335ms | |
| Gemma 3 12B | 100% | $0.0000 | 362ms | |
| Ministral 3 14B | 100% | $0.0000 | 515ms | |
| Gemma 4 26B | 100% | $0.0000 | 566ms | |
| Mistral Small 4 | 100% | $0.0000 | 543ms | |
| Gemma 3 27B | 100% | $0.0000 | 621ms | |
| Gemma 4 31B | 100% | $0.0000 | 12.3s | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0000 | 377ms | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0000 | 429ms | |
| DeepSeek V3 (2024-12-26) | 100% | $0.0000 | 547ms | |
| Mistral Medium 3.1 | 100% | $0.0000 | 350ms | |
| Gemini 2.5 Flash | 100% | $0.0000 | 499ms | |
| GPT-5.4 Nano | 90% | $0.0000 | 550ms | |
| Gemini 3.1 Flash Lite | 100% | $0.0000 | 709ms | |
| Gemini 3.1 Flash Lite (Preview) | 100% | $0.0000 | 680ms | |
| Gemini 3.1 Flash Lite (Reasoning) | 100% | $0.0000 | 2.3s | |
| GPT-4.1 Nano | 100% | $0.0000 | 928ms | |
Most Stable Models (Top 20)
Ranked by stability (median × consistency). Click a model name to view its detail page.
| Score | Consistency | Stability | ||
|---|---|---|---|---|
| GPT-5.6 Sol (Reasoning) | 100% | 100% | 100% | |
| Claude Opus 4.6 (Reasoning) | 100% | 100% | 100% | |
| Gemini 3.6 Flash (Reasoning) | 100% | 100% | 100% | |
| Qwen3.7 Max | 100% | 100% | 100% | |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.1 Pro (Preview) | 100% | 100% | 100% | |
| GPT-5.4 (Reasoning) | 100% | 100% | 100% | |
| Muse Spark 1.1 (Reasoning, Medium) | 100% | 100% | 100% | |
| Z.AI GLM 5.1 | 100% | 100% | 100% | |
| Qwen3.6 Max Preview | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning) | 100% | 100% | 100% | |
| Claude Sonnet 4.6 (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5.2 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.5 Flash (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5 Turbo | 100% | 100% | 100% | |
| MoonshotAI: Kimi K3 (Reasoning, High) | 100% | 100% | 100% | |
| MoonshotAI: Kimi K2.6 | 100% | 100% | 100% | |
| Claude Opus 4.7 (Reasoning) | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning, Low) | 100% | 100% | 100% | |
| GPT-5.6 Terra (Reasoning) | 100% | 100% | 100% | |
Top Overall Models (Top 20)
Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.
| Score | Cost | Speed | Stability | ||
|---|---|---|---|---|---|
| Gemma 3 4B | 100% | $0.0000 | 217ms | 100% | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 335ms | 100% | |
| Mistral Small 3.2 24B | 100% | $0.0000 | 335ms | 100% | |
| Gemma 3 12B | 100% | $0.0000 | 362ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0000 | 377ms | 100% | |
| Mistral Medium 3.1 | 100% | $0.0000 | 350ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0000 | 429ms | 100% | |
| Ministral 3 14B | 100% | $0.0000 | 515ms | 100% | |
| Mistral Small 4 | 100% | $0.0000 | 543ms | 100% | |
| Gemma 4 26B | 100% | $0.0000 | 566ms | 100% | |
| Gemini 2.5 Flash | 100% | $0.0000 | 499ms | 100% | |
| Mistral NeMO | 100% | $0.0000 | 577ms | 100% | |
| Gemma 3 27B | 100% | $0.0000 | 621ms | 100% | |
| DeepSeek V3 (2024-12-26) | 100% | $0.0000 | 547ms | 100% | |
| Hermes 3 70B | 100% | $0.0000 | 507ms | 100% | |
| Mistral Large 3 | 100% | $0.0000 | 547ms | 100% | |
| Gemini 3.1 Flash Lite (Preview) | 100% | $0.0000 | 680ms | 100% | |
| Gemini 3.1 Flash Lite | 100% | $0.0000 | 709ms | 100% | |
| Qwen 2.5 72B | 100% | $0.0000 | 729ms | 100% | |
| Inception Mercury 2 | 100% | $0.0001 | 390ms | 100% | |
| Median | Evaluator | Top 3 | Flop 3 |
|---|---|---|---|
| 100.0% | Matches Regex | ||
| 100.0% | Matches text |
Guess the pet
Performance Score Distribution (Top 20)
Click a model name to view its detail page.
Price-Performance Score Distribution (Top 20)
Click a model name to view its detail page.
| Score | Cost | Time | ||
|---|---|---|---|---|
| Gemma 3 4B | 100% | $0.0000 | 261ms | |
| Ministral 3B | 100% | $0.0000 | 267ms | |
| Ministral 8B | 100% | $0.0000 | 282ms | |
| Ministral 3 3B | 100% | $0.0000 | 656ms | |
| Mistral NeMO | 100% | $0.0000 | 647ms | |
| Gemma 3 12B | 100% | $0.0000 | 327ms | |
| Mistral Small 3.2 24B | 100% | $0.0000 | 361ms | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 345ms | |
| Ministral 3 8B | 100% | $0.0000 | 377ms | |
| Ministral 3 14B | 100% | $0.0000 | 438ms | |
| Gemma 3 27B | 100% | $0.0000 | 467ms | |
| Mistral Small 4 | 100% | $0.0000 | 441ms | |
| Gemma 4 31B | 100% | $0.0000 | 1.2s | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0000 | 355ms | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0000 | 413ms | |
| DeepSeek V4 Flash | 100% | $0.0000 | 3.7s | |
| Gemini 2.5 Flash | 100% | $0.0000 | 466ms | |
| Qwen3 235B A22B Instruct 2507 | 100% | $0.0000 | 1.7s | |
| GPT-4.1 Nano | 100% | $0.0000 | 846ms | |
| DeepSeek V3 (2024-12-26) | 100% | $0.0000 | 975ms | |
Most Stable Models (Top 20)
Ranked by stability (median × consistency). Click a model name to view its detail page.
| Score | Consistency | Stability | ||
|---|---|---|---|---|
| GPT-5.6 Sol (Reasoning) | 100% | 100% | 100% | |
| Claude Opus 4.6 (Reasoning) | 100% | 100% | 100% | |
| Gemini 3.6 Flash (Reasoning) | 100% | 100% | 100% | |
| Qwen3.7 Max | 100% | 100% | 100% | |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.1 Pro (Preview) | 100% | 100% | 100% | |
| GPT-5.4 (Reasoning) | 100% | 100% | 100% | |
| Muse Spark 1.1 (Reasoning, Medium) | 100% | 100% | 100% | |
| Z.AI GLM 5.1 | 100% | 100% | 100% | |
| Qwen3.6 Max Preview | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning) | 100% | 100% | 100% | |
| Claude Sonnet 4.6 (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5.2 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.5 Flash (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5 Turbo | 100% | 100% | 100% | |
| MoonshotAI: Kimi K3 (Reasoning, High) | 100% | 100% | 100% | |
| MoonshotAI: Kimi K2.6 | 100% | 100% | 100% | |
| Claude Opus 4.7 (Reasoning) | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning, Low) | 100% | 100% | 100% | |
| GPT-5.6 Terra (Reasoning) | 100% | 100% | 100% | |
Top Overall Models (Top 20)
Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.
| Score | Cost | Speed | Stability | ||
|---|---|---|---|---|---|
| Gemma 3 4B | 100% | $0.0000 | 261ms | 100% | |
| Ministral 3B | 100% | $0.0000 | 267ms | 100% | |
| Ministral 8B | 100% | $0.0000 | 282ms | 100% | |
| Gemma 3 12B | 100% | $0.0000 | 327ms | 100% | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 345ms | 100% | |
| Mistral Small 3.2 24B | 100% | $0.0000 | 361ms | 100% | |
| Ministral 3 8B | 100% | $0.0000 | 377ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0000 | 355ms | 100% | |
| Mistral Small 4 | 100% | $0.0000 | 441ms | 100% | |
| Cydonia 24B V4.1 | 100% | $0.0000 | 277ms | 100% | |
| Ministral 3 14B | 100% | $0.0000 | 438ms | 100% | |
| Gemma 3 27B | 100% | $0.0000 | 467ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0000 | 413ms | 100% | |
| Gemini 2.5 Flash | 100% | $0.0000 | 466ms | 100% | |
| Mistral Large 3 | 100% | $0.0000 | 480ms | 100% | |
| GPT-5.4 Nano | 100% | $0.0000 | 548ms | 100% | |
| Ministral 3 3B | 100% | $0.0000 | 656ms | 100% | |
| Mistral NeMO | 100% | $0.0000 | 647ms | 100% | |
| Inception Mercury 2 | 100% | $0.0001 | 350ms | 100% | |
| Gemini 3.1 Flash Lite (Reasoning) | 100% | $0.0000 | 661ms | 100% | |
| Median | Evaluator | Top 3 | Flop 3 |
|---|---|---|---|
| 100.0% | Matches Regex | — | |
| 100.0% | Matches text |
Who's the sister?
Performance Score Distribution (Top 20)
Click a model name to view its detail page.
Price-Performance Score Distribution (Top 20)
Click a model name to view its detail page.
| Score | Cost | Time | ||
|---|---|---|---|---|
| Nemotron 3 Super | 100% | $0.0000 | 1.4s | |
| Ministral 3B | 100% | $0.0000 | 289ms | |
| Gemma 3 4B | 100% | $0.0000 | 252ms | |
| Gemma 3 12B | 100% | $0.0000 | 392ms | |
| Ministral 8B | 100% | $0.0000 | 276ms | |
| Ministral 3 3B | 100% | $0.0000 | 363ms | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 375ms | |
| Mistral Small 3.2 24B | 100% | $0.0000 | 461ms | |
| Gemma 3 27B | 100% | $0.0000 | 533ms | |
| Gemma 4 26B | 100% | $0.0000 | 701ms | |
| Mistral NeMO | 100% | $0.0000 | 13.7s | |
| GPT-4.1 Nano | 100% | $0.0000 | 1.2s | |
| Ministral 3 8B | 100% | $0.0000 | 333ms | |
| Qwen3 235B A22B Instruct 2507 | 90% | $0.0000 | 1.5s | |
| DeepSeek V4 Flash | 100% | $0.0000 | 2.1s | |
| DeepSeek-V2 Chat | 100% | $0.0000 | 1.3s | |
| Ministral 3 14B | 100% | $0.0000 | 360ms | |
| Mistral Small 4 | 100% | $0.0000 | 451ms | |
| Gemma 4 31B | 100% | $0.0000 | 1.1s | |
| DeepSeek V3.1 | 100% | $0.0000 | 1.7s | |
Most Stable Models (Top 20)
Ranked by stability (median × consistency). Click a model name to view its detail page.
| Score | Consistency | Stability | ||
|---|---|---|---|---|
| GPT-5.6 Sol (Reasoning) | 100% | 100% | 100% | |
| Claude Opus 4.6 (Reasoning) | 100% | 100% | 100% | |
| Gemini 3.6 Flash (Reasoning) | 100% | 100% | 100% | |
| Qwen3.7 Max | 100% | 100% | 100% | |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.1 Pro (Preview) | 100% | 100% | 100% | |
| GPT-5.4 (Reasoning) | 100% | 100% | 100% | |
| Muse Spark 1.1 (Reasoning, Medium) | 100% | 100% | 100% | |
| Z.AI GLM 5.1 | 100% | 100% | 100% | |
| Qwen3.6 Max Preview | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning) | 100% | 100% | 100% | |
| Claude Sonnet 4.6 (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5.2 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.5 Flash (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5 Turbo | 100% | 100% | 100% | |
| MoonshotAI: Kimi K3 (Reasoning, High) | 100% | 100% | 100% | |
| MoonshotAI: Kimi K2.6 | 100% | 100% | 100% | |
| Claude Opus 4.7 (Reasoning) | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning, Low) | 100% | 100% | 100% | |
| GPT-5.6 Terra (Reasoning) | 100% | 100% | 100% | |
Top Overall Models (Top 20)
Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.
| Score | Cost | Speed | Stability | ||
|---|---|---|---|---|---|
| Gemma 3 4B | 100% | $0.0000 | 252ms | 100% | |
| Ministral 3B | 100% | $0.0000 | 289ms | 100% | |
| Ministral 8B | 100% | $0.0000 | 276ms | 100% | |
| Gemma 3 12B | 100% | $0.0000 | 392ms | 100% | |
| Ministral 3 3B | 100% | $0.0000 | 363ms | 100% | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 375ms | 100% | |
| Mistral Small 3.2 24B | 100% | $0.0000 | 461ms | 100% | |
| Ministral 3 8B | 100% | $0.0000 | 333ms | 100% | |
| Gemma 3 27B | 100% | $0.0000 | 533ms | 100% | |
| Ministral 3 14B | 100% | $0.0000 | 360ms | 100% | |
| Mistral Small 4 | 100% | $0.0000 | 451ms | 100% | |
| Gemma 4 26B | 100% | $0.0000 | 701ms | 100% | |
| Nemotron 3 Super | 100% | $0.0000 | 1.4s | 100% | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0000 | 362ms | 100% | |
| Gemini 2.5 Flash | 100% | $0.0000 | 410ms | 100% | |
| GPT-4.1 Nano | 100% | $0.0000 | 1.2s | 100% | |
| Gemini 3.1 Flash Lite (Preview) | 100% | $0.0000 | 714ms | 100% | |
| GPT-5.4 Nano (Reasoning, Low) | 100% | $0.0000 | 671ms | 100% | |
| Gemma 4 31B | 100% | $0.0000 | 1.1s | 100% | |
| DeepSeek-V2 Chat | 100% | $0.0000 | 1.3s | 100% | |
| Median | Evaluator | Top 3 | Flop 3 |
|---|---|---|---|
| 100.0% | Matches Regex | ||
| 100.0% | Matches text |
Contextual pronoun
Performance Score Distribution (Top 20)
Click a model name to view its detail page.
Price-Performance Score Distribution (Top 20)
Click a model name to view its detail page.
| Score | Cost | Time | ||
|---|---|---|---|---|
| Gemma 3 4B | 100% | $0.0000 | 238ms | |
| Ministral 3B | 80% | $0.0000 | 262ms | |
| Ministral 3 3B | 100% | $0.0000 | 277ms | |
| Ministral 8B | 90% | $0.0000 | 276ms | |
| Mistral NeMO | 100% | $0.0000 | 3.9s | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 328ms | |
| Ministral 3 8B | 100% | $0.0000 | 393ms | |
| Gemma 3 12B | 100% | $0.0000 | 509ms | |
| Ministral 3 14B | 100% | $0.0000 | 414ms | |
| Mistral Small 3.2 24B | 100% | $0.0000 | 634ms | |
| Mistral Small 4 | 100% | $0.0000 | 413ms | |
| Gemma 4 26B | 100% | $0.0000 | 920ms | |
| Gemma 3 27B | 100% | $0.0000 | 1.5s | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0000 | 402ms | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0000 | 608ms | |
| Gemini 2.5 Flash | 100% | $0.0000 | 430ms | |
| Qwen3 235B A22B Instruct 2507 | 100% | $0.0000 | 1.1s | |
| Gemini 3.1 Flash Lite | 100% | $0.0000 | 631ms | |
| GPT-5.4 Nano | 100% | $0.0000 | 774ms | |
| Gemini 3.1 Flash Lite (Reasoning) | 100% | $0.0000 | 691ms | |
Most Stable Models (Top 20)
Ranked by stability (median × consistency). Click a model name to view its detail page.
| Score | Consistency | Stability | ||
|---|---|---|---|---|
| GPT-5.6 Sol (Reasoning) | 100% | 100% | 100% | |
| Claude Opus 4.6 (Reasoning) | 100% | 100% | 100% | |
| Gemini 3.6 Flash (Reasoning) | 100% | 100% | 100% | |
| Qwen3.7 Max | 100% | 100% | 100% | |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.1 Pro (Preview) | 100% | 100% | 100% | |
| GPT-5.4 (Reasoning) | 100% | 100% | 100% | |
| Muse Spark 1.1 (Reasoning, Medium) | 100% | 100% | 100% | |
| Z.AI GLM 5.1 | 100% | 100% | 100% | |
| Qwen3.6 Max Preview | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning) | 100% | 100% | 100% | |
| Claude Sonnet 4.6 (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5.2 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.5 Flash (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5 Turbo | 100% | 100% | 100% | |
| MoonshotAI: Kimi K3 (Reasoning, High) | 100% | 100% | 100% | |
| MoonshotAI: Kimi K2.6 | 100% | 100% | 100% | |
| Claude Opus 4.7 (Reasoning) | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning, Low) | 100% | 100% | 100% | |
| GPT-5.6 Terra (Reasoning) | 100% | 100% | 100% | |
Top Overall Models (Top 20)
Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.
| Score | Cost | Speed | Stability | ||
|---|---|---|---|---|---|
| Gemma 3 4B | 100% | $0.0000 | 238ms | 100% | |
| Ministral 3 3B | 100% | $0.0000 | 277ms | 100% | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 328ms | 100% | |
| Ministral 3 8B | 100% | $0.0000 | 393ms | 100% | |
| Ministral 3 14B | 100% | $0.0000 | 414ms | 100% | |
| Mistral Small 4 | 100% | $0.0000 | 413ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0000 | 402ms | 100% | |
| Gemini 2.5 Flash | 100% | $0.0000 | 430ms | 100% | |
| Gemma 3 12B | 100% | $0.0000 | 509ms | 100% | |
| Mistral Large 3 | 100% | $0.0000 | 499ms | 100% | |
| Mistral Small 3.2 24B | 100% | $0.0000 | 634ms | 100% | |
| Mistral Medium 3.1 | 100% | $0.0000 | 551ms | 100% | |
| Llama 3.1 70B | 100% | $0.0001 | 413ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0000 | 608ms | 100% | |
| Gemini 3.1 Flash Lite | 100% | $0.0000 | 631ms | 100% | |
| Hermes 3 70B | 100% | $0.0000 | 590ms | 100% | |
| Gemini 3.1 Flash Lite (Preview) | 100% | $0.0000 | 682ms | 100% | |
| Gemini 3.1 Flash Lite (Reasoning) | 100% | $0.0000 | 691ms | 100% | |
| Inception Mercury 2 | 100% | $0.0001 | 335ms | 100% | |
| Qwen 2.5 72B | 100% | $0.0000 | 680ms | 100% | |
| Median | Evaluator | Top 3 | Flop 3 |
|---|---|---|---|
| 100.0% | Matches Regex | ||
| 100.0% | Matches text |
Indirect birth year
Performance Score Distribution (Top 20)
Click a model name to view its detail page.
Price-Performance Score Distribution (Top 20)
Click a model name to view its detail page.
| Score | Cost | Time | ||
|---|---|---|---|---|
| Gemma 3 4B | 100% | $0.0000 | 309ms | |
| Gemma 3 12B | 100% | $0.0000 | 427ms | |
| Ministral 3 3B | 90% | $0.0000 | 361ms | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 350ms | |
| Ministral 3 14B | 100% | $0.0000 | 434ms | |
| Mistral Small 4 | 90% | $0.0000 | 749ms | |
| DeepSeek V3 (2024-12-26) | 100% | $0.0000 | 640ms | |
| Gemma 3 27B | 100% | $0.0000 | 790ms | |
| GPT-4.1 Nano | 100% | $0.0000 | 937ms | |
| Qwen3 235B A22B Instruct 2507 | 100% | $0.0000 | 1.9s | |
| GPT-5.4 Nano | 100% | $0.0000 | 704ms | |
| Gemini 3.1 Flash Lite | 100% | $0.0000 | 595ms | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0000 | 375ms | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0000 | 400ms | |
| Gemini 3.1 Flash Lite (Reasoning) | 100% | $0.0000 | 702ms | |
| Gemma 4 31B | 100% | $0.0000 | 5.7s | |
| Hermes 3 70B | 100% | $0.0000 | 447ms | |
| Gemini 3.1 Flash Lite (Preview) | 100% | $0.0000 | 702ms | |
| Gemini 2.5 Flash | 100% | $0.0000 | 501ms | |
| Mistral Medium 3.1 | 100% | $0.0000 | 560ms | |
Most Stable Models (Top 20)
Ranked by stability (median × consistency). Click a model name to view its detail page.
| Score | Consistency | Stability | ||
|---|---|---|---|---|
| GPT-5.6 Sol (Reasoning) | 100% | 100% | 100% | |
| Claude Opus 4.6 (Reasoning) | 100% | 100% | 100% | |
| Gemini 3.6 Flash (Reasoning) | 100% | 100% | 100% | |
| Qwen3.7 Max | 100% | 100% | 100% | |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.1 Pro (Preview) | 100% | 100% | 100% | |
| GPT-5.4 (Reasoning) | 100% | 100% | 100% | |
| Muse Spark 1.1 (Reasoning, Medium) | 100% | 100% | 100% | |
| Z.AI GLM 5.1 | 100% | 100% | 100% | |
| Qwen3.6 Max Preview | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning) | 100% | 100% | 100% | |
| Claude Sonnet 4.6 (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5.2 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.5 Flash (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5 Turbo | 100% | 100% | 100% | |
| MoonshotAI: Kimi K3 (Reasoning, High) | 100% | 100% | 100% | |
| MoonshotAI: Kimi K2.6 | 100% | 100% | 100% | |
| Claude Opus 4.7 (Reasoning) | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning, Low) | 100% | 100% | 100% | |
| GPT-5.6 Terra (Reasoning) | 100% | 100% | 100% | |
Top Overall Models (Top 20)
Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.
| Score | Cost | Speed | Stability | ||
|---|---|---|---|---|---|
| Gemma 3 4B | 100% | $0.0000 | 309ms | 100% | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 350ms | 100% | |
| Gemma 3 12B | 100% | $0.0000 | 427ms | 100% | |
| Ministral 3 14B | 100% | $0.0000 | 434ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0000 | 375ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0000 | 400ms | 100% | |
| Hermes 3 70B | 100% | $0.0000 | 447ms | 100% | |
| Gemini 2.5 Flash | 100% | $0.0000 | 501ms | 100% | |
| Gemini 3.1 Flash Lite | 100% | $0.0000 | 595ms | 100% | |
| Cydonia 24B V4.1 | 100% | $0.0000 | 425ms | 100% | |
| DeepSeek V3 (2024-12-26) | 100% | $0.0000 | 640ms | 100% | |
| Gemma 3 27B | 100% | $0.0000 | 790ms | 100% | |
| Mistral Medium 3.1 | 100% | $0.0000 | 560ms | 100% | |
| Mistral Large 3 | 100% | $0.0000 | 533ms | 100% | |
| GPT-5.4 Nano | 100% | $0.0000 | 704ms | 100% | |
| Gemini 3.1 Flash Lite (Reasoning) | 100% | $0.0000 | 702ms | 100% | |
| Gemini 3.1 Flash Lite (Preview) | 100% | $0.0000 | 702ms | 100% | |
| Qwen 2.5 72B | 100% | $0.0000 | 700ms | 100% | |
| GPT-4.1 Nano | 100% | $0.0000 | 937ms | 100% | |
| Llama 3.1 70B | 100% | $0.0001 | 646ms | 100% | |
| Median | Evaluator | Top 3 | Flop 3 |
|---|---|---|---|
| 100.0% | Matches Regex | ||
| 100.0% | Matches text |
Fruits excluding citrus
Performance Score Distribution (Top 20)
Click a model name to view its detail page.
Price-Performance Score Distribution (Top 20)
Click a model name to view its detail page.
| Score | Cost | Time | ||
|---|---|---|---|---|
| Ministral 3B | 90% | $0.0000 | 344ms | |
| Gemma 3 4B | 100% | $0.0000 | 419ms | |
| Ministral 3 3B | 100% | $0.0000 | 387ms | |
| Ministral 8B | 90% | $0.0000 | 404ms | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 368ms | |
| Mistral NeMO | 100% | $0.0000 | 1.7s | |
| Ministral 3 8B | 100% | $0.0000 | 436ms | |
| Ministral 3 14B | 100% | $0.0000 | 608ms | |
| Mistral Small 3.2 24B | 100% | $0.0000 | 712ms | |
| Mistral Small 4 | 100% | $0.0000 | 532ms | |
| Gemma 3 27B | 100% | $0.0000 | 1.1s | |
| Gemma 4 31B | 100% | $0.0000 | 3.5s | |
| Gemma 3 12B | 100% | $0.0000 | 987ms | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0001 | 395ms | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0001 | 509ms | |
| Gemini 3.1 Flash Lite (Reasoning) | 100% | $0.0000 | 1.4s | |
| Gemini 3.1 Flash Lite | 100% | $0.0000 | 647ms | |
| DeepSeek V4 Flash | 100% | $0.0000 | 2.9s | |
| Gemini 3.1 Flash Lite (Preview) | 100% | $0.0000 | 742ms | |
| Gemma 4 26B | 100% | $0.0000 | 1.9s | |
Most Stable Models (Top 20)
Ranked by stability (median × consistency). Click a model name to view its detail page.
| Score | Consistency | Stability | ||
|---|---|---|---|---|
| GPT-5.6 Sol (Reasoning) | 100% | 100% | 100% | |
| Claude Opus 4.6 (Reasoning) | 100% | 100% | 100% | |
| Gemini 3.6 Flash (Reasoning) | 100% | 100% | 100% | |
| Qwen3.7 Max | 100% | 100% | 100% | |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.1 Pro (Preview) | 100% | 100% | 100% | |
| GPT-5.4 (Reasoning) | 100% | 100% | 100% | |
| Muse Spark 1.1 (Reasoning, Medium) | 100% | 100% | 100% | |
| Z.AI GLM 5.1 | 100% | 100% | 100% | |
| Qwen3.6 Max Preview | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning) | 100% | 100% | 100% | |
| Claude Sonnet 4.6 (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5.2 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.5 Flash (Reasoning) | 100% | 100% | 100% | |
| MoonshotAI: Kimi K3 (Reasoning, High) | 100% | 100% | 100% | |
| MoonshotAI: Kimi K2.6 | 100% | 100% | 100% | |
| Claude Opus 4.7 (Reasoning) | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning, Low) | 100% | 100% | 100% | |
| GPT-5.6 Terra (Reasoning) | 100% | 100% | 100% | |
| Claude Opus 4.8 (Reasoning) | 100% | 100% | 100% | |
Top Overall Models (Top 20)
Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.
| Score | Cost | Speed | Stability | ||
|---|---|---|---|---|---|
| Gemma 3 4B | 100% | $0.0000 | 419ms | 100% | |
| Ministral 3 3B | 100% | $0.0000 | 387ms | 100% | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 368ms | 100% | |
| Ministral 3 8B | 100% | $0.0000 | 436ms | 100% | |
| Mistral Small 4 | 100% | $0.0000 | 532ms | 100% | |
| Ministral 3 14B | 100% | $0.0000 | 608ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0001 | 395ms | 100% | |
| Mistral Small 3.2 24B | 100% | $0.0000 | 712ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0001 | 509ms | 100% | |
| Gemini 3.1 Flash Lite | 100% | $0.0000 | 647ms | 100% | |
| Gemini 3.1 Flash Lite (Preview) | 100% | $0.0000 | 742ms | 100% | |
| Mistral Medium 3.1 | 100% | $0.0001 | 559ms | 100% | |
| Gemma 3 12B | 100% | $0.0000 | 987ms | 100% | |
| Mistral Large 3 | 100% | $0.0001 | 643ms | 100% | |
| Gemma 3 27B | 100% | $0.0000 | 1.1s | 100% | |
| Gemini 3 Flash (Preview) | 100% | $0.0001 | 832ms | 100% | |
| Qwen 2.5 72B | 100% | $0.0000 | 1.1s | 100% | |
| Llama 3.1 70B | 100% | $0.0001 | 683ms | 100% | |
| GPT-5.4 Nano | 100% | $0.0000 | 1.1s | 100% | |
| Inception Mercury 2 | 100% | $0.0002 | 489ms | 100% | |
| Median | Evaluator | Top 3 | Flop 3 |
|---|---|---|---|
| 100.0% | Contains a list of texts |
Future event time
Performance Score Distribution (Top 20)
Click a model name to view its detail page.
Price-Performance Score Distribution (Top 20)
Click a model name to view its detail page.
| Score | Cost | Time | ||
|---|---|---|---|---|
| Gemma 3 4B | 100% | $0.0000 | 257ms | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 350ms | |
| Gemma 3 12B | 100% | $0.0000 | 503ms | |
| Gemma 3 27B | 100% | $0.0000 | 696ms | |
| Gemma 4 26B | 100% | $0.0000 | 1.2s | |
| Gemma 4 31B | 100% | $0.0000 | 2.4s | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0000 | 351ms | |
| Gemini 2.5 Flash | 100% | $0.0000 | 380ms | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0000 | 471ms | |
| Gemini 3.1 Flash Lite (Reasoning) | 95% | $0.0000 | 608ms | |
| Gemini 3.1 Flash Lite | 100% | $0.0000 | 670ms | |
| GPT-4.1 Nano | 95% | $0.0000 | 958ms | |
| Gemini 3.1 Flash Lite (Preview) | 95% | $0.0000 | 728ms | |
| Inception Mercury 2 | 100% | $0.0001 | 337ms | |
| Gemini 3 Flash (Preview) | 100% | $0.0001 | 789ms | |
| GPT-5.4 Mini | 100% | $0.0001 | 653ms | |
| Qwen 3.5 Plus (2026-02-15) | 100% | $0.0001 | 1.4s | |
| DeepSeek V4 Flash (Reasoning) | 95% | $0.0000 | 2.6s | |
| GPT-5.6 Luna | 100% | $0.0001 | 851ms | |
| GPT-5.6 Luna (Reasoning) | 100% | $0.0002 | 1.1s | |
Most Stable Models (Top 20)
Ranked by stability (median × consistency). Click a model name to view its detail page.
| Score | Consistency | Stability | ||
|---|---|---|---|---|
| GPT-5.6 Sol (Reasoning) | 100% | 100% | 100% | |
| Gemini 3.6 Flash (Reasoning) | 100% | 100% | 100% | |
| Qwen3.7 Max | 100% | 100% | 100% | |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.1 Pro (Preview) | 100% | 100% | 100% | |
| GPT-5.4 (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5.1 | 100% | 100% | 100% | |
| Qwen3.6 Max Preview | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5.2 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.5 Flash (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5 Turbo | 100% | 100% | 100% | |
| MoonshotAI: Kimi K3 (Reasoning, High) | 100% | 100% | 100% | |
| MoonshotAI: Kimi K2.6 | 100% | 100% | 100% | |
| Claude Opus 4.7 (Reasoning) | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning, Low) | 100% | 100% | 100% | |
| GPT-5.6 Terra (Reasoning) | 100% | 100% | 100% | |
| MoonshotAI: Kimi K3 (Reasoning, Low) | 100% | 100% | 100% | |
| GPT-5 | 100% | 100% | 100% | |
| GPT-5 Mini | 100% | 100% | 100% | |
Top Overall Models (Top 20)
Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.
| Score | Cost | Speed | Stability | ||
|---|---|---|---|---|---|
| Gemma 3 4B | 100% | $0.0000 | 257ms | 100% | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 350ms | 100% | |
| Gemma 3 12B | 100% | $0.0000 | 503ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0000 | 351ms | 100% | |
| Gemini 2.5 Flash | 100% | $0.0000 | 380ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0000 | 471ms | 100% | |
| Gemma 3 27B | 100% | $0.0000 | 696ms | 100% | |
| Gemini 3.1 Flash Lite | 100% | $0.0000 | 670ms | 100% | |
| Inception Mercury 2 | 100% | $0.0001 | 337ms | 100% | |
| Gemini 3 Flash (Preview) | 100% | $0.0001 | 789ms | 100% | |
| GPT-5.4 Mini | 100% | $0.0001 | 653ms | 100% | |
| Gemma 4 26B | 100% | $0.0000 | 1.2s | 100% | |
| GPT-5.6 Luna | 100% | $0.0001 | 851ms | 100% | |
| Qwen 3.5 Plus (2026-02-15) | 100% | $0.0001 | 1.4s | 100% | |
| Gemini 3.6 Flash (Reasoning, Minimal) | 100% | $0.0002 | 818ms | 100% | |
| GPT-5.6 Luna (Reasoning) | 100% | $0.0002 | 1.1s | 100% | |
| Gemma 4 31B | 100% | $0.0000 | 2.4s | 100% | |
| GPT-5.4 | 100% | $0.0003 | 681ms | 100% | |
| Nemotron 3 Super | 100% | $0.0000 | 2.8s | 100% | |
| GPT-5.6 Terra | 100% | $0.0003 | 742ms | 100% | |
| Median | Evaluator | Top 3 | Flop 3 |
|---|---|---|---|
| 100.0% | Matches Regex | ||
| 100.0% | Matches text |
Highest-rated movie
Performance Score Distribution (Top 20)
Click a model name to view its detail page.
Price-Performance Score Distribution (Top 20)
Click a model name to view its detail page.
| Score | Cost | Time | ||
|---|---|---|---|---|
| Gemma 3 4B | 100% | $0.0000 | 267ms | |
| Ministral 3B | 100% | $0.0000 | 273ms | |
| Ministral 8B | 100% | $0.0000 | 311ms | |
| Ministral 3 8B | 100% | $0.0000 | 353ms | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 351ms | |
| Ministral 3 14B | 100% | $0.0000 | 350ms | |
| Ministral 3 3B | 100% | $0.0000 | 388ms | |
| Mistral Small 3.2 24B | 100% | $0.0000 | 414ms | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0000 | 383ms | |
| Cydonia 24B V4.1 | 100% | $0.0001 | 752ms | |
| Gemma 3 12B | 100% | $0.0000 | 486ms | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0000 | 425ms | |
| Inception Mercury 2 | 100% | $0.0001 | 373ms | |
| Mistral Small 4 | 100% | $0.0000 | 523ms | |
| Gemini 2.5 Flash | 100% | $0.0000 | 465ms | |
| Mistral Medium 3.1 | 100% | $0.0001 | 556ms | |
| Qwen3 235B A22B Instruct 2507 | 100% | $0.0000 | 1.0s | |
| Gemma 3 27B | 100% | $0.0000 | 621ms | |
| Gemma 4 31B | 100% | $0.0000 | 886ms | |
| Mistral Large 3 | 100% | $0.0001 | 547ms | |
Most Stable Models (Top 20)
Ranked by stability (median × consistency). Click a model name to view its detail page.
| Score | Consistency | Stability | ||
|---|---|---|---|---|
| GPT-5.6 Sol (Reasoning) | 100% | 100% | 100% | |
| Claude Opus 4.6 (Reasoning) | 100% | 100% | 100% | |
| Gemini 3.6 Flash (Reasoning) | 100% | 100% | 100% | |
| Qwen3.7 Max | 100% | 100% | 100% | |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.1 Pro (Preview) | 100% | 100% | 100% | |
| GPT-5.4 (Reasoning) | 100% | 100% | 100% | |
| Muse Spark 1.1 (Reasoning, Medium) | 100% | 100% | 100% | |
| Z.AI GLM 5.1 | 100% | 100% | 100% | |
| Qwen3.6 Max Preview | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning) | 100% | 100% | 100% | |
| Claude Sonnet 4.6 (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5.2 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.5 Flash (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5 Turbo | 100% | 100% | 100% | |
| MoonshotAI: Kimi K3 (Reasoning, High) | 100% | 100% | 100% | |
| MoonshotAI: Kimi K2.6 | 100% | 100% | 100% | |
| Claude Opus 4.7 (Reasoning) | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning, Low) | 100% | 100% | 100% | |
| GPT-5.6 Terra (Reasoning) | 100% | 100% | 100% | |
Top Overall Models (Top 20)
Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.
| Score | Cost | Speed | Stability | ||
|---|---|---|---|---|---|
| Gemma 3 4B | 100% | $0.0000 | 267ms | 100% | |
| Ministral 3B | 100% | $0.0000 | 273ms | 100% | |
| Ministral 8B | 100% | $0.0000 | 311ms | 100% | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 351ms | 100% | |
| Ministral 3 8B | 100% | $0.0000 | 353ms | 100% | |
| Ministral 3 14B | 100% | $0.0000 | 350ms | 100% | |
| Ministral 3 3B | 100% | $0.0000 | 388ms | 100% | |
| Mistral Small 3.2 24B | 100% | $0.0000 | 414ms | 100% | |
| Gemma 3 12B | 100% | $0.0000 | 486ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0000 | 383ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0000 | 425ms | 100% | |
| Mistral Small 4 | 100% | $0.0000 | 523ms | 100% | |
| Gemini 2.5 Flash | 100% | $0.0000 | 465ms | 100% | |
| Inception Mercury 2 | 100% | $0.0001 | 373ms | 100% | |
| Gemma 3 27B | 100% | $0.0000 | 621ms | 100% | |
| Mistral Medium 3.1 | 100% | $0.0001 | 556ms | 100% | |
| Mistral Large 3 | 100% | $0.0001 | 547ms | 100% | |
| Gemini 3.1 Flash Lite (Reasoning) | 100% | $0.0000 | 648ms | 100% | |
| Gemini 3.1 Flash Lite (Preview) | 100% | $0.0000 | 677ms | 100% | |
| Gemini 3.1 Flash Lite | 100% | $0.0000 | 720ms | 100% | |
| Median | Evaluator | Top 3 | Flop 3 |
|---|---|---|---|
| 100.0% | Matches Regex | ||
| 100.0% | Matches text |
All valid emails
Performance Score Distribution (Top 20)
Click a model name to view its detail page.
Price-Performance Score Distribution (Top 20)
Click a model name to view its detail page.
| Score | Cost | Time | ||
|---|---|---|---|---|
| Ministral 3B | 80% | $0.0000 | 380ms | |
| Ministral 3 3B | 100% | $0.0000 | 1.0s | |
| Gemma 3 4B | 100% | $0.0000 | 656ms | |
| Ministral 3 8B | 100% | $0.0000 | 469ms | |
| Ministral 8B | 80% | $0.0000 | 521ms | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 398ms | |
| Ministral 3 14B | 100% | $0.0000 | 610ms | |
| Gemma 4 26B | 100% | $0.0000 | 981ms | |
| Mistral Small 4 | 100% | $0.0000 | 670ms | |
| Gemma 3 12B | 100% | $0.0000 | 1.1s | |
| DeepSeek V4 Flash | 100% | $0.0000 | 2.0s | |
| Mistral Small 3.2 24B | 100% | $0.0000 | 1.9s | |
| GPT-4.1 Nano | 100% | $0.0000 | 1.4s | |
| GPT-5.4 Nano | 100% | $0.0001 | 902ms | |
| GPT-5.4 Nano (Reasoning, Low) | 100% | $0.0001 | 984ms | |
| Gemma 3 27B | 100% | $0.0000 | 1.5s | |
| GPT-5.4 Nano (Reasoning) | 100% | $0.0001 | 1.3s | |
| Gemini 3.1 Flash Lite | 100% | $0.0001 | 764ms | |
| Gemini 3.1 Flash Lite (Preview) | 100% | $0.0001 | 722ms | |
| Gemini 3.1 Flash Lite (Reasoning) | 100% | $0.0001 | 824ms | |
Most Stable Models (Top 20)
Ranked by stability (median × consistency). Click a model name to view its detail page.
| Score | Consistency | Stability | ||
|---|---|---|---|---|
| GPT-5.6 Sol (Reasoning) | 100% | 100% | 100% | |
| Claude Opus 4.6 (Reasoning) | 100% | 100% | 100% | |
| Gemini 3.6 Flash (Reasoning) | 100% | 100% | 100% | |
| Qwen3.7 Max | 100% | 100% | 100% | |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.1 Pro (Preview) | 100% | 100% | 100% | |
| GPT-5.4 (Reasoning) | 100% | 100% | 100% | |
| Muse Spark 1.1 (Reasoning, Medium) | 100% | 100% | 100% | |
| Z.AI GLM 5.1 | 100% | 100% | 100% | |
| Qwen3.6 Max Preview | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning) | 100% | 100% | 100% | |
| Claude Sonnet 4.6 (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5.2 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.5 Flash (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5 Turbo | 100% | 100% | 100% | |
| MoonshotAI: Kimi K3 (Reasoning, High) | 100% | 100% | 100% | |
| MoonshotAI: Kimi K2.6 | 100% | 100% | 100% | |
| Claude Opus 4.7 (Reasoning) | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning, Low) | 100% | 100% | 100% | |
| GPT-5.6 Terra (Reasoning) | 100% | 100% | 100% | |
Top Overall Models (Top 20)
Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.
| Score | Cost | Speed | Stability | ||
|---|---|---|---|---|---|
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 398ms | 100% | |
| Ministral 3 8B | 100% | $0.0000 | 469ms | 100% | |
| Gemma 3 4B | 100% | $0.0000 | 656ms | 100% | |
| Ministral 3 14B | 100% | $0.0000 | 610ms | 100% | |
| Mistral Small 4 | 100% | $0.0000 | 670ms | 100% | |
| Ministral 3 3B | 100% | $0.0000 | 1.0s | 100% | |
| Gemma 4 26B | 100% | $0.0000 | 981ms | 100% | |
| Gemma 3 12B | 100% | $0.0000 | 1.1s | 100% | |
| Gemini 3.1 Flash Lite (Preview) | 100% | $0.0001 | 722ms | 100% | |
| GPT-5.4 Nano | 100% | $0.0001 | 902ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning) | 100% | $0.0001 | 438ms | 100% | |
| Gemini 3.1 Flash Lite | 100% | $0.0001 | 764ms | 100% | |
| Gemini 2.5 Flash | 100% | $0.0001 | 467ms | 100% | |
| Inception Mercury 2 | 100% | $0.0001 | 354ms | 100% | |
| GPT-5.4 Nano (Reasoning, Low) | 100% | $0.0001 | 984ms | 100% | |
| Gemini 3.1 Flash Lite (Reasoning) | 100% | $0.0001 | 824ms | 100% | |
| Gemini 3.5 Flash Lite (Reasoning, Minimal) | 100% | $0.0001 | 553ms | 100% | |
| Mistral Large 3 | 100% | $0.0001 | 803ms | 100% | |
| GPT-4.1 Nano | 100% | $0.0000 | 1.4s | 100% | |
| Gemma 3 27B | 100% | $0.0000 | 1.5s | 100% | |
| Median | Evaluator | Top 3 | Flop 3 |
|---|---|---|---|
| 100.0% | Contains a list of texts |