Text Replacement
Tests deterministic text transformations: renaming characters/locations, expanding contractions, tense rewriting, POV shifts, gender swaps, combined transformations, and word avoidance. Scored by checking each expected change independently.
Performance Score Distribution (Top 20)
Click a model name to view its detail page.
Price-Performance Score Distribution (Top 20)
Click a model name to view its detail page.
| Score | Cost | Time | ||
|---|---|---|---|---|
| Gemini 2.5 Flash Lite | 97% | $0.0003 | 1.7s | |
| Gemini 3.1 Flash Lite (Preview) | 99% | $0.0010 | 1.8s | |
| Gemini 3.1 Flash Lite (Reasoning) | 99% | $0.0010 | 3.1s | |
| Mistral Small 4 | 96% | $0.0004 | 3.3s | |
| Gemini 3.1 Flash Lite | 99% | $0.0010 | 2.8s | |
| Mistral Small 3.2 24B | 97% | $0.0002 | 5.0s | |
| Gemini 2.5 Flash | 99% | $0.0015 | 2.2s | |
| DeepSeek V4 Flash | 97% | $0.0002 | 8.1s | |
| Gemini 3 Flash (Preview) | 99% | $0.0019 | 3.4s | |
| GPT-4.1 Mini | 98% | $0.0011 | 7.0s | |
| Mistral Large 3 | 98% | $0.0011 | 7.7s | |
| Gemma 3 12B | 95% | $0.0001 | 9.0s | |
| Inception Mercury 2 | 95% | $0.0017 | 2.3s | |
| Grok 4.20 | 98% | $0.0020 | 4.4s | |
| Qwen 2.5 72B | 98% | $0.0003 | 10.9s | |
| Qwen 3.5 Plus (2026-02-15) | 99% | $0.0015 | 7.2s | |
| GPT-4o Mini (temp=1) | 95% | $0.0004 | 9.5s | |
| Grok 4.3 | 95% | $0.0021 | 4.7s | |
| Mistral Medium 3.1 | 97% | $0.0013 | 5.9s | |
| Claude Haiku 4.5 | 99% | $0.0036 | 3.2s | |
Cost vs Performance
Compares total cost for this test against the test score. Quadrant lines are drawn at the median values. Only models with available cost data are shown.
12 low-scoring outliers hidden: DeepSeek V3.1 (89.5%), GPT-4.1 Nano (89.3%), Gemma 3 4B (89.3%), Ministral 3 8B (87.0%), Ministral 8B (86.7%), Mistral NeMO (86.6%), Arcee AI: Trinity Mini (85.7%), Nemotron 3 Nano (83.3%), Ministral 3 3B (81.2%), Ministral 3B (80.9%), Cohere Command R+ (Aug. 2024) (73.7%), Hermes 3 70B (69.5%).
Most Stable Models (Top 20)
Ranked by stability (median × consistency). Click a model name to view its detail page.
| Score | Consistency | Stability | ||
|---|---|---|---|---|
| Claude Opus 4.6 | 100% | 99% | 99% | |
| Gemma 4 31B | 100% | 98% | 98% | |
| Claude Opus 4.5 | 100% | 98% | 98% | |
| Gemma 4 31B (Reasoning) | 100% | 98% | 98% | |
| Claude Sonnet 4 | 100% | 98% | 98% | |
| Grok 4.5 (Reasoning, High) | 100% | 98% | 98% | |
| Claude Opus 4.6 (Reasoning) | 100% | 98% | 98% | |
| Claude Sonnet 4.5 | 100% | 98% | 98% | |
| Claude Opus 4.8 (Reasoning, Low) | 100% | 98% | 98% | |
| Claude Opus 4.8 (Reasoning) | 100% | 98% | 98% | |
| Claude Sonnet 5 (Reasoning, Low) | 99% | 98% | 98% | |
| Claude Sonnet 4.6 (Reasoning) | 100% | 98% | 98% | |
| Qwen3.6 Max Preview | 100% | 98% | 98% | |
| Z.AI GLM 5.1 | 100% | 98% | 98% | |
| Claude Sonnet 5 (Reasoning) | 99% | 98% | 98% | |
| GPT-5.6 Terra | 100% | 98% | 98% | |
| Qwen 3.5 27B | 99% | 98% | 98% | |
| Z.AI GLM 5 | 99% | 98% | 98% | |
| Claude Opus 4.7 (Reasoning) | 99% | 98% | 98% | |
| Gemma 4 26B (Reasoning) | 99% | 97% | 97% | |
Top Overall Models (Top 20)
Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.
| Score | Cost | Speed | Stability | ||
|---|---|---|---|---|---|
| Gemini 3.1 Flash Lite | 99% | $0.0010 | 2.8s | 96% | |
| Gemini 3.1 Flash Lite (Preview) | 99% | $0.0010 | 1.8s | 96% | |
| Gemini 3.1 Flash Lite (Reasoning) | 99% | $0.0010 | 3.1s | 96% | |
| Gemini 3 Flash (Preview) | 99% | $0.0019 | 3.4s | 97% | |
| Qwen 3.5 Plus (2026-02-15) | 99% | $0.0015 | 7.2s | 97% | |
| Claude Haiku 4.5 | 99% | $0.0036 | 3.2s | 97% | |
| Gemma 4 26B | 99% | $0.0003 | 17.1s | 97% | |
| Gemini 2.5 Flash | 99% | $0.0015 | 2.2s | 92% | |
| GPT-4.1 Mini | 98% | $0.0011 | 7.0s | 94% | |
| Grok 4.20 | 98% | $0.0020 | 4.4s | 92% | |
| GPT-5.6 Luna | 98% | $0.0038 | 3.0s | 94% | |
| DeepSeek V4 Pro | 99% | $0.0013 | 21.1s | 97% | |
| GPT-5.6 Terra | 100% | $0.0094 | 3.3s | 98% | |
| Gemma 4 31B | 100% | $0.0003 | 30.2s | 98% | |
| GPT-5.6 Luna (Reasoning) | 99% | $0.0052 | 5.0s | 94% | |
| Claude Sonnet 4.5 | 100% | $0.011 | 4.9s | 98% | |
| Mistral Medium 3.1 | 97% | $0.0013 | 5.9s | 91% | |
| Claude Sonnet 4 | 100% | $0.011 | 6.1s | 98% | |
| Gemini 3.5 Flash (Reasoning, Minimal) | 99% | $0.0057 | 2.6s | 91% | |
| GPT-4.1 | 98% | $0.0054 | 4.4s | 92% | |
| Specific Prompt | Generic Prompt | Specific Prompt | Generic Prompt | Specific Prompt | Generic Prompt | Specific Prompt | Generic Prompt | Specific Prompt | Generic Prompt | Specific Prompt | Generic Prompt | Specific Prompt | Generic Prompt | Specific Prompt | Generic Prompt | Specific Prompt | Generic Prompt | ||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | Total ▼ | Character rename: Elena->Mirabel, Gregor->Aldric | Character rename: Elena->Mirabel, Gregor->Aldric | Location rename: market square, outer ring, bridge, northern mines | Location rename: market square, outer ring, bridge, northern mines | Expand all contractions | Expand all contractions | Tense rewriting: past to present | Tense rewriting: past to present | POV shift: 3rd person to 1st person (Elena's perspective) | POV shift: 3rd person to 1st person (Elena's perspective) | Multi-character gender swap: Priya(F)->Rohan(M), Mara unchanged | Multi-character gender swap: Priya(F)->Rohan(M), Mara unchanged | Combined: 3rd person past → 1st person present | Combined: 3rd person past → 1st person present | Passive voice → active voice | Passive voice → active voice | Avoid said/asked/replied/answered | Avoid said/asked/replied/answered |
| Claude Sonnet 4 | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 99% | 98% | 97% | 100% | 100% |
| Gemma 4 31B | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 98% | 100% | 100% | 100% | 100% | 100% | 99% | 99% | 98% | 100% | 100% |
| Claude Opus 4.6 | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 99% | 100% | 99% | 100% | 100% | 100% | 99% | 98% | 98% | 100% | 100% |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 99% | 100% | 100% | 100% | 100% | 99% | 99% | 99% | 97% | 100% | 100% |
| Claude Sonnet 4.5 | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 99% | 99% | 99% | 96% | 100% | 100% |
| Claude Opus 4.6 (Reasoning) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 99% | 100% | 100% | 100% | 99% | 99% | 96% | 100% | 100% |
| Gemma 4 31B (Reasoning) | 100% | 100% | 100% | 100% | 100% | 100% | 99% | 100% | 99% | 100% | 100% | 100% | 100% | 100% | 99% | 99% | 97% | 100% | 100% |
| Z.AI GLM 5.2 (Reasoning, High) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 99% | 100% | 100% | 100% | 96% | 100% | 99% | 99% | 98% | 100% | 100% |
| Claude Opus 4.8 (Reasoning, Low) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 99% | 100% | 100% | 100% | 100% | 96% | 99% | 99% | 98% | 100% | 100% |
| GPT-5.6 Terra | 100% | 100% | 100% | 100% | 100% | 100% | 99% | 100% | 99% | 100% | 100% | 100% | 100% | 100% | 99% | 98% | 96% | 100% | 100% |
| Claude Sonnet 4.6 (Reasoning) | 100% | 100% | 100% | 100% | 100% | 100% | 99% | 100% | 99% | 100% | 100% | 100% | 100% | 100% | 99% | 99% | 96% | 100% | 100% |
| Claude Opus 4.5 | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 99% | 100% | 99% | 100% | 100% | 100% | 99% | 97% | 97% | 100% | 100% |
| Qwen3.6 Max Preview | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 98% | 100% | 100% | 100% | 100% | 100% | 99% | 98% | 96% | 100% | 100% |
| Claude Opus 4.8 (Reasoning) | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 99% | 99% | 100% | 100% | 100% | 100% | 96% | 99% | 99% | 98% | 100% | 100% |
| Z.AI GLM 5.1 | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 97% | 100% | 100% | 100% | 100% | 100% | 99% | 99% | 97% | 100% | 100% |
Generic Prompt
Character rename: Elena->Mirabel, Gregor->Aldric
Performance Score Distribution (Top 20)
Click a model name to view its detail page.
Price-Performance Score Distribution (Top 20)
Click a model name to view its detail page.
| Score | Cost | Time | ||
|---|---|---|---|---|
| Gemini 2.5 Flash Lite | 100% | $0.0003 | 1.6s | |
| Inception Mercury 2 | 100% | $0.0007 | 971ms | |
| Ministral 8B | 100% | $0.0001 | 3.2s | |
| Ministral 3 8B | 100% | $0.0002 | 2.9s | |
| Mistral Small 4 | 100% | $0.0004 | 2.8s | |
| GPT-4.1 Nano | 100% | $0.0003 | 3.7s | |
| Ministral 3 14B | 100% | $0.0002 | 4.0s | |
| Gemini 3.1 Flash Lite (Preview) | 100% | $0.0009 | 1.6s | |
| Gemini 3.1 Flash Lite (Reasoning) | 100% | $0.0009 | 4.3s | |
| Gemini 3.1 Flash Lite | 100% | $0.0009 | 1.9s | |
| GPT-5.4 Nano (Reasoning, Low) | 100% | $0.0007 | 2.8s | |
| GPT-5.4 Nano | 100% | $0.0007 | 2.7s | |
| GPT-5.4 Nano (Reasoning) | 100% | $0.0007 | 3.2s | |
| Gemma 3 4B | 100% | $0.0001 | 6.1s | |
| Mistral Small 3.2 24B | 100% | $0.0002 | 5.7s | |
| Mistral NeMO | 93% | $0.0002 | 2.3s | |
| Gemini 2.5 Flash | 100% | $0.0014 | 2.1s | |
| DeepSeek V4 Flash | 100% | $0.0002 | 8.3s | |
| Gemini 2.5 Flash Lite (Reasoning) | 100% | $0.0007 | 4.7s | |
| ByteDance Seed 1.6 Flash | 100% | $0.0003 | 6.6s | |
Most Stable Models (Top 20)
Ranked by stability (median × consistency). Click a model name to view its detail page.
| Score | Consistency | Stability | ||
|---|---|---|---|---|
| GPT-5.6 Sol (Reasoning) | 100% | 100% | 100% | |
| Claude Opus 4.6 (Reasoning) | 100% | 100% | 100% | |
| Qwen3.7 Max | 100% | 100% | 100% | |
| Grok 4.5 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.1 Pro (Preview) | 100% | 100% | 100% | |
| GPT-5.4 (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5.1 | 100% | 100% | 100% | |
| Qwen3.6 Max Preview | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning) | 100% | 100% | 100% | |
| Claude Sonnet 4.6 (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5.2 (Reasoning, High) | 100% | 100% | 100% | |
| Gemini 3.5 Flash (Reasoning) | 100% | 100% | 100% | |
| Z.AI GLM 5 Turbo | 100% | 100% | 100% | |
| MoonshotAI: Kimi K2.6 | 100% | 100% | 100% | |
| Claude Opus 4.7 (Reasoning) | 100% | 100% | 100% | |
| GPT-5.5 (Reasoning, Low) | 100% | 100% | 100% | |
| GPT-5.6 Terra (Reasoning) | 100% | 100% | 100% | |
| Claude Opus 4.8 (Reasoning) | 100% | 100% | 100% | |
| Claude Opus 4.6 | 100% | 100% | 100% | |
| Claude Opus 4.8 (Reasoning, Low) | 100% | 100% | 100% | |
Top Overall Models (Top 20)
Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.
| Score | Cost | Speed | Stability | ||
|---|---|---|---|---|---|
| Gemini 2.5 Flash Lite | 100% | $0.0003 | 1.6s | 100% | |
| Inception Mercury 2 | 100% | $0.0007 | 971ms | 100% | |
| Ministral 3 8B | 100% | $0.0002 | 2.9s | 100% | |
| Gemini 3.1 Flash Lite (Preview) | 100% | $0.0009 | 1.6s | 100% | |
| Ministral 8B | 100% | $0.0001 | 3.2s | 100% | |
| Mistral Small 4 | 100% | $0.0004 | 2.8s | 100% | |
| Gemini 3.1 Flash Lite | 100% | $0.0009 | 1.9s | 100% | |
| GPT-5.4 Nano | 100% | $0.0007 | 2.7s | 100% | |
| GPT-5.4 Nano (Reasoning, Low) | 100% | $0.0007 | 2.8s | 100% | |
| GPT-4.1 Nano | 100% | $0.0003 | 3.7s | 100% | |
| Ministral 3 14B | 100% | $0.0002 | 4.0s | 100% | |
| GPT-5.4 Nano (Reasoning) | 100% | $0.0007 | 3.2s | 100% | |
| Gemini 2.5 Flash | 100% | $0.0014 | 2.1s | 100% | |
| Gemini 3.1 Flash Lite (Reasoning) | 100% | $0.0009 | 4.3s | 100% | |
| Gemini 2.5 Flash Lite (Reasoning) | 100% | $0.0007 | 4.7s | 100% | |
| Mistral Small 3.2 24B | 100% | $0.0002 | 5.7s | 100% | |
| Gemma 3 4B | 100% | $0.0001 | 6.1s | 100% | |
| Gemini 3 Flash (Preview) | 100% | $0.0018 | 3.2s | 100% | |
| GPT-5.4 Mini | 100% | $0.0027 | 1.9s | 100% | |
| Mistral Medium 3.1 | 100% | $0.0012 | 4.8s | 100% | |