Data extraction
Extract key details from a given block of text.
Performance Score Distribution (Top 20)
Click a model name to view its detail page.
Price-Performance Score Distribution (Top 20)
Click a model name to view its detail page.
| Score | Cost | Time | ||
|---|---|---|---|---|
| Gemma 3 4B | 92% | $0.0000 | 303ms | |
| Ministral 3B | 73% | $0.0000 | 308ms | |
| Ministral 8B | 75% | $0.0000 | 331ms | |
| Gemini 2.5 Flash Lite | 92% | $0.0000 | 357ms | |
| Ministral 3 3B | 78% | $0.0000 | 416ms | |
| Ministral 3 14B | 88% | $0.0000 | 448ms | |
| Gemma 3 12B | 92% | $0.0000 | 542ms | |
| Ministral 3 8B | 71% | $0.0000 | 382ms | |
| Mistral Small 3.2 24B | 83% | $0.0000 | 691ms | |
| Mistral Small 4 | 88% | $0.0000 | 539ms | |
| Gemini 2.5 Flash | 83% | $0.0000 | 473ms | |
| Gemma 3 27B | 92% | $0.0000 | 780ms | |
| Cydonia 24B V4.1 | 81% | $0.0000 | 593ms | |
| GPT-5.4 Nano | 93% | $0.0000 | 768ms | |
| Mistral Medium 3.1 | 88% | $0.0000 | 655ms | |
| Gemini 3.1 Flash Lite (Reasoning) | 91% | $0.0000 | 1.2s | |
| Mistral Large 3 | 88% | $0.0000 | 945ms | |
| Gemma 4 31B | 97% | $0.0000 | 5.1s | |
| Gemini 3.1 Flash Lite | 92% | $0.0000 | 757ms | |
| Gemma 4 26B | 95% | $0.0000 | 1.8s | |
Cost vs Performance
Compares total cost for this test against the test score. Quadrant lines are drawn at the median values. Only models with available cost data are shown.
11 low-scoring outliers hidden: Grok 4.20 (88.6%), Arcee AI: Trinity Mini (88.6%), Cydonia 24B V4.1 (87.3%), Ministral 3 3B (85.5%), Ministral 8B (80.9%), WizardLM 2 8x22b (79.1%), DeepSeek V4 Flash (77.7%), Ministral 3B (77.7%), Ministral 3 8B (77.3%), Grok 4.20 (Reasoning) (72.3%), Cohere Command R+ (Aug. 2024) (68.2%).
Most Stable Models (Top 20)
Ranked by stability (median × consistency). Click a model name to view its detail page.
| Score | Consistency | Stability | ||
|---|---|---|---|---|
| GPT-5.6 Sol (Reasoning) | 100% | 100% | 100% | |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | |
| Z.AI GLM 5.2 (Reasoning, High) | 100% | 100% | 100% | |
| GPT-5.6 Terra (Reasoning) | 100% | 100% | 100% | |
| GPT-5.6 Sol | 100% | 100% | 100% | |
| GPT-5.6 Luna (Reasoning) | 100% | 100% | 100% | |
| GPT-5.6 Terra | 100% | 100% | 100% | |
| GPT-5.6 Luna | 100% | 100% | 100% | |
| Grok 4.5 (Reasoning, Low) | 100% | 91% | 91% | |
| Gemini 3.5 Flash (Reasoning) | 99% | 82% | 82% | |
| Gemini 3 Flash (Preview, Reasoning) | 99% | 82% | 82% | |
| Aion 3.0 Mini | 98% | 75% | 75% | |
| Gemma 4 26B (Reasoning) | 98% | 74% | 74% | |
| Claude Sonnet 4 | 96% | 72% | 72% | |
| GPT-4o Mini (temp=0) | 96% | 72% | 72% | |
| Claude Sonnet 5 (Reasoning, Low) | 95% | 71% | 71% | |
| Gemma 4 31B (Reasoning) | 98% | 69% | 69% | |
| Aion 3.0 | 96% | 68% | 68% | |
| Claude Sonnet 5 (Reasoning) | 94% | 68% | 68% | |
| Gemma 4 31B | 97% | 64% | 64% | |
Top Overall Models (Top 20)
Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.
| Score | Cost | Speed | Stability | ||
|---|---|---|---|---|---|
| GPT-5.6 Luna | 100% | $0.0001 | 698ms | 100% | |
| GPT-5.6 Terra | 100% | $0.0003 | 800ms | 100% | |
| GPT-5.6 Terra (Reasoning) | 100% | $0.0004 | 1.1s | 100% | |
| GPT-5.6 Luna (Reasoning) | 100% | $0.0003 | 1.7s | 100% | |
| GPT-5.6 Sol | 100% | $0.0006 | 1.1s | 100% | |
| GPT-5.6 Sol (Reasoning) | 100% | $0.0007 | 1.2s | 100% | |
| Z.AI GLM 5.2 (Reasoning, High) | 100% | $0.0012 | 5.7s | 100% | |
| Grok 4.5 (Reasoning, High) | 100% | $0.0018 | 5.6s | 100% | |
| Grok 4.5 (Reasoning, Low) | 100% | $0.0014 | 6.2s | 91% | |
| Gemini 3 Flash (Preview, Reasoning) | 99% | $0.0026 | 7.0s | 82% | |
| Claude Sonnet 4 | 96% | $0.0004 | 1.6s | 72% | |
| Aion 3.0 Mini | 98% | $0.0005 | 8.5s | 75% | |
| Claude Sonnet 5 (Reasoning, Low) | 95% | $0.0005 | 3.4s | 71% | |
| GPT-4o Mini (temp=0) | 96% | $0.0000 | 8.1s | 72% | |
| Gemma 4 31B | 97% | $0.0000 | 5.1s | 64% | |
| Claude Sonnet 5 (Reasoning) | 94% | $0.0005 | 3.2s | 68% | |
| Aion 3.0 | 96% | $0.0014 | 4.8s | 68% | |
| Claude Opus 4.7 | 96% | $0.0008 | 1.0s | 60% | |
| Gemma 4 26B | 95% | $0.0000 | 1.8s | 56% | |
| Claude Sonnet 5 | 92% | $0.0004 | 2.8s | 64% | |
| Model | Total ▼ | Who's the tallest? | What's the color of the car? | What instrument does Lucy play? | Guess the pet | Who's the sister? | Contextual pronoun | Indirect birth year | Fruits excluding citrus | Future event time | Highest-rated movie | All valid emails |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-5.6 Sol (Reasoning) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| Qwen3.7 Max | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| Gemini 3.1 Pro (Preview) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| GPT-5.4 (Reasoning) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| Z.AI GLM 5.1 | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| Qwen3.6 Max Preview | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| GPT-5.5 (Reasoning) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| Z.AI GLM 5.2 (Reasoning, High) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| Gemini 3.5 Flash (Reasoning) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| MoonshotAI: Kimi K2.6 | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| Claude Opus 4.7 (Reasoning) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| GPT-5.5 (Reasoning, Low) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| GPT-5.6 Terra (Reasoning) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
| GPT-5 Mini | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% |
Who's the tallest?
Performance Score Distribution (Top 20)
Click a model name to view its detail page.
Price-Performance Score Distribution (Top 20)
Click a model name to view its detail page.
| Score | Cost | Time | ||
|---|---|---|---|---|
| Gemma 3 4B | 100% | $0.0000 | 217ms | |
| Ministral 3B | 100% | $0.0000 | 269ms | |
| Gemma 3 12B | 100% | $0.0000 | 322ms | |
| Ministral 8B | 90% | $0.0000 | 261ms | |
| Ministral 3 3B | 100% | $0.0000 | 272ms | |
| Ministral 3 8B | 100% | $0.0000 | 368ms | |
| Mistral Small 3.2 24B | 100% | $0.0000 | 363ms | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 389ms | |
| Ministral 3 14B | 100% | $0.0000 | 392ms | |
| Mistral NeMO | 80% | $0.0000 | 318ms | |
| Gemma 3 27B | 100% | $0.0000 | 467ms | |
| Mistral Small 4 | 100% | $0.0000 | 463ms | |
| DeepSeek V3 (2024-12-26) | 100% | $0.0000 | 594ms | |
| Gemma 4 26B | 100% | $0.0000 | 1.8s | |
| Gemini 2.5 Flash | 100% | $0.0000 | 399ms | |
| DeepSeek V4 Flash | 100% | $0.0000 | 2.4s | |
| GPT-5.4 Nano | 100% | $0.0000 | 579ms | |
| Gemini 3.1 Flash Lite | 100% | $0.0000 | 742ms | |
| Gemma 4 31B | 100% | $0.0000 | 6.4s | |
| Qwen3 235B A22B Instruct 2507 | 100% | $0.0000 | 658ms | |
Most Stable Models (Top 20)
Ranked by stability (median × consistency). Click a model name to view its detail page.
| Score | Consistency | Stability | ||
|---|---|---|---|---|
| GPT-5.6 Sol (Reasoning) | 100% | 100% | 100% | |
| Claude Opus 4.6 (Reasoning) | 100% | 100% | 100% | |
| Qwen3.7 Max | 100% | 100% | 100% | |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.1 Pro (Preview) | 100% | 100% | 100% | |
| GPT-5.4 (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5.1 | 100% | 100% | 100% | |
| Qwen3.6 Max Preview | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning) | 100% | 100% | 100% | |
| Claude Sonnet 4.6 (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5.2 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.5 Flash (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5 Turbo | 100% | 100% | 100% | |
| MoonshotAI: Kimi K2.6 | 100% | 100% | 100% | |
| Claude Opus 4.7 (Reasoning) | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning, Low) | 100% | 100% | 100% | |
| GPT-5.6 Terra (Reasoning) | 100% | 100% | 100% | |
| Claude Opus 4.8 (Reasoning) | 100% | 100% | 100% | |
| Claude Opus 4.6 | 100% | 100% | 100% | |
| Claude Opus 4.8 (Reasoning, Low) | 100% | 100% | 100% | |
Top Overall Models (Top 20)
Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.
| Score | Cost | Speed | Stability | ||
|---|---|---|---|---|---|
| Gemma 3 4B | 100% | $0.0000 | 217ms | 100% | |
| Ministral 3B | 100% | $0.0000 | 269ms | 100% | |
| Ministral 3 3B | 100% | $0.0000 | 272ms | 100% | |
| Gemma 3 12B | 100% | $0.0000 | 322ms | 100% | |
| Mistral Small 3.2 24B | 100% | $0.0000 | 363ms | 100% | |
| Ministral 3 8B | 100% | $0.0000 | 368ms | 100% | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 389ms | 100% | |
| Ministral 3 14B | 100% | $0.0000 | 392ms | 100% | |
| Gemini 2.5 Flash | 100% | $0.0000 | 399ms | 100% | |
| Gemma 3 27B | 100% | $0.0000 | 467ms | 100% | |
| Mistral Small 4 | 100% | $0.0000 | 463ms | 100% | |
| Llama 3.1 70B | 100% | $0.0000 | 419ms | 100% | |
| GPT-5.4 Nano | 100% | $0.0000 | 579ms | 100% | |
| DeepSeek V3 (2024-12-26) | 100% | $0.0000 | 594ms | 100% | |
| Qwen3 235B A22B Instruct 2507 | 100% | $0.0000 | 658ms | 100% | |
| Qwen 2.5 72B | 100% | $0.0000 | 596ms | 100% | |
| Inception Mercury 2 | 100% | $0.0001 | 367ms | 100% | |
| Mistral Large 2 | 100% | $0.0001 | 314ms | 100% | |
| Gemini 3.1 Flash Lite | 100% | $0.0000 | 742ms | 100% | |
| Gemini 3.1 Flash Lite (Preview) | 100% | $0.0000 | 746ms | 100% | |
| Median | Evaluator | Top 3 | Flop 3 |
|---|---|---|---|
| 100.0% | Matches Regex | ||
| 100.0% | Matches text |
What's the color of the car?
Performance Score Distribution (Top 20)
Click a model name to view its detail page.
Price-Performance Score Distribution (Top 20)
Click a model name to view its detail page.
| Score | Cost | Time | ||
|---|---|---|---|---|
| Gemma 3 4B | 100% | $0.0000 | 221ms | |
| Ministral 3B | 100% | $0.0000 | 273ms | |
| Mistral NeMO | 100% | $0.0000 | 536ms | |
| Ministral 8B | 100% | $0.0000 | 256ms | |
| Ministral 3 3B | 100% | $0.0000 | 293ms | |
| Qwen3 235B A22B Instruct 2507 | 100% | $0.0000 | 758ms | |
| Ministral 3 8B | 100% | $0.0000 | 361ms | |
| Mistral Small 3.2 24B | 100% | $0.0000 | 406ms | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 375ms | |
| Ministral 3 14B | 100% | $0.0000 | 362ms | |
| Gemma 4 26B | 100% | $0.0000 | 1.5s | |
| Gemma 3 12B | 100% | $0.0000 | 403ms | |
| Gemma 4 31B | 100% | $0.0000 | 2.0s | |
| Gemma 3 27B | 100% | $0.0000 | 518ms | |
| Mistral Small 4 | 100% | $0.0000 | 488ms | |
| Mistral Medium 3.1 | 100% | $0.0000 | 441ms | |
| Cydonia 24B V4.1 | 100% | $0.0000 | 300ms | |
| Gemini 2.5 Flash | 100% | $0.0000 | 549ms | |
| GPT-5.4 Nano | 100% | $0.0000 | 578ms | |
| Gemini 3.1 Flash Lite (Reasoning) | 100% | $0.0000 | 610ms | |
Most Stable Models (Top 20)
Ranked by stability (median × consistency). Click a model name to view its detail page.
| Score | Consistency | Stability | ||
|---|---|---|---|---|
| GPT-5.6 Sol (Reasoning) | 100% | 100% | 100% | |
| Claude Opus 4.6 (Reasoning) | 100% | 100% | 100% | |
| Qwen3.7 Max | 100% | 100% | 100% | |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.1 Pro (Preview) | 100% | 100% | 100% | |
| GPT-5.4 (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5.1 | 100% | 100% | 100% | |
| Qwen3.6 Max Preview | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning) | 100% | 100% | 100% | |
| Claude Sonnet 4.6 (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5.2 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.5 Flash (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5 Turbo | 100% | 100% | 100% | |
| MoonshotAI: Kimi K2.6 | 100% | 100% | 100% | |
| Claude Opus 4.7 (Reasoning) | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning, Low) | 100% | 100% | 100% | |
| GPT-5.6 Terra (Reasoning) | 100% | 100% | 100% | |
| Claude Opus 4.8 (Reasoning) | 100% | 100% | 100% | |
| Claude Opus 4.6 | 100% | 100% | 100% | |
| Claude Opus 4.8 (Reasoning, Low) | 100% | 100% | 100% | |
Top Overall Models (Top 20)
Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.
| Score | Cost | Speed | Stability | ||
|---|---|---|---|---|---|
| Gemma 3 4B | 100% | $0.0000 | 221ms | 100% | |
| Ministral 3B | 100% | $0.0000 | 273ms | 100% | |
| Ministral 8B | 100% | $0.0000 | 256ms | 100% | |
| Ministral 3 3B | 100% | $0.0000 | 293ms | 100% | |
| Gemini 2.5 Flash Lite | 100% | $0.0000 | 375ms | 100% | |
| Ministral 3 8B | 100% | $0.0000 | 361ms | 100% | |
| Ministral 3 14B | 100% | $0.0000 | 362ms | 100% | |
| Gemma 3 12B | 100% | $0.0000 | 403ms | 100% | |
| Mistral Small 3.2 24B | 100% | $0.0000 | 406ms | 100% | |
| Cydonia 24B V4.1 | 100% | $0.0000 | 300ms | 100% | |
| Mistral Small 4 | 100% | $0.0000 | 488ms | 100% | |
| Gemma 3 27B | 100% | $0.0000 | 518ms | 100% | |
| Mistral NeMO | 100% | $0.0000 | 536ms | 100% | |
| Mistral Medium 3.1 | 100% | $0.0000 | 441ms | 100% | |
| Gemini 2.5 Flash | 100% | $0.0000 | 549ms | 100% | |
| GPT-5.4 Nano | 100% | $0.0000 | 578ms | 100% | |
| Gemini 3.1 Flash Lite (Reasoning) | 100% | $0.0000 | 610ms | 100% | |
| Gemini 3.1 Flash Lite (Preview) | 100% | $0.0000 | 638ms | 100% | |
| Llama 3.1 70B | 100% | $0.0001 | 413ms | 100% | |
| Mistral Large 3 | 100% | $0.0000 | 617ms | 100% |