Non-passive narration preserved

Test: Text Replacement

Avg. Score
86.0%
Scenarios
2

Overall Performance

Rank ▲ Model Score Avg. Cost Avg. Time Stability
1Claude Haiku 4.5100.0%$0.00515.8s100%
2Claude Sonnet 4.6100.0%$0.0157.5s100%
3Claude Sonnet 4100.0%$0.0159.4s100%
4Qwen 3.5 Plus (2026-02-15)99.1%$0.002210.2s94%
5Gemini 2.5 Flash Lite98.2%$0.00042.5s91%
6Claude Opus 4.5100.0%$0.0268.1s100%
7Claude Opus 4.6100.0%$0.0268.7s100%
8Gemini 2.5 Flash95.5%$0.00213.0s88%
9Claude Sonnet 4.596.4%$0.0157.2s89%
10Qwen 2.5 72B94.6%$0.000416.3s84%
11Mistral Small 3.2 24B93.8%$0.00036.9s82%
12Gemini 3 Flash (Preview)93.8%$0.00274.3s82%
13Mistral Large 393.8%$0.001610.5s82%
14Grok 4 Fast93.8%$0.001411.7s82%
15Writer: Palmyra X593.8%$0.005013.7s84%
16DeepSeek V3.295.5%$0.000753.2s88%
17Ministral 3 14B92.9%$0.00035.7s79%
18Mistral Large93.8%$0.006210.4s82%
19Mistral Large 293.8%$0.006210.5s82%
20GPT-4o Mini (temp=1)87.5%$0.000612.4s88%
21Grok 4.1 Fast93.8%$0.001722.5s82%
22GPT-4o Mini (temp=0)87.5%$0.000613.8s88%
23Minimax M2.597.3%$0.00211.3m90%
24ByteDance Seed 1.6 Flash92.9%$0.001221.9s82%
25Gemma 3 12B92.9%$0.000112.8s77%
26GPT-4.1 Mini88.4%$0.001510.3s82%
27Hermes 3 405B87.5%$0.001731.7s88%
28Mistral Small Creative91.1%$0.00034.1s73%
29DeepSeek V3 (2024-12-26)91.1%$0.001223.1s78%
30Gemma 3 27B86.6%$0.000324.0s82%
31Arcee AI: Trinity Large (Preview)89.3%$0.000031.5s80%
32GPT-5 Mini93.8%$0.008250.8s82%
33Gemini 2.5 Flash Lite (Reasoning)92.0%$0.003332.9s77%
34Z.AI GLM 4.693.8%$0.007857.8s82%
35Arcee AI: Trinity Mini86.6%$0.000311.2s75%
36Claude 3.7 Sonnet86.6%$0.0158.5s82%
37GPT-4o, May 13th (temp=0)90.2%$0.0154.6s75%
38Gemini 2.5 Flash (Reasoning)90.2%$0.01422.4s79%
39Gemma 3 4B88.4%$0.000110.1s70%
40Mistral Medium 3.182.1%$0.00186.3s77%
41Z.AI GLM 4.592.0%$0.006352.4s77%
42Grok 495.5%$0.03947.6s88%
43ByteDance Seed 1.687.5%$0.00771.4m88%
44Gemini 3 Flash (Preview, Reasoning)88.4%$0.02236.9s82%
45GPT-4o, May 13th (temp=1)85.7%$0.0154.7s74%
46GPT-5.288.4%$0.02522.2s77%
47GPT-5.192.0%$0.02934.1s77%
48Ministral 3 8B83.9%$0.00024.2s57%
49GPT-4.180.4%$0.00766.4s66%
50Ministral 8B83.0%$0.00023.9s57%
51DeepSeek V3.188.4%$0.000935.4s52%
52Llama 3.1 70B77.7%$0.000718.9s61%
53Claude 3.5 Sonnet81.3%$0.03013.8s71%
54Gemini 2.5 Pro92.9%$0.05641.9s77%
55GPT-592.0%$0.0501.3m77%
56Claude Opus 491.1%$0.07713.4s78%
57Gemini 3 Pro (Preview)89.3%$0.06443.7s80%
58GPT-4o, Aug. 6th (temp=0)81.3%$0.00913.9s46%
59Z.AI GLM 4.792.0%$0.0173.0m77%
60DeepSeek V3 (2025-03-24)82.1%$0.000842.6s46%
61Llama 3.1 8B75.9%$0.000115.3s42%
62Qwen 3.5 397B A17B90.2%$0.0113.7m79%
63Z.AI GLM 591.1%$0.0233.2m78%
64GPT-4.1 Nano75.0%$0.00045.0s39%
65Aion 2.085.7%$0.00691.5m45%
66WizardLM 2 8x22b82.1%$0.00141.5m45%
67Z.AI GLM 4.7 Flash84.8%$0.00342.4m55%
68o4 Mini83.0%$0.02639.2s44%
69DeepSeek-V2 Chat75.9%$0.001121.5s33%
70Ministral 3B71.4%$0.00012.9s32%
71Llama 3.1 Nemotron 70B64.3%$0.002120.8s47%
72GPT-5 Nano78.6%$0.00521.9m52%
73GPT-4o, Aug. 6th (temp=1)71.4%$0.00864.1s31%
74Claude Sonnet 4.6 (Reasoning)93.8%$0.1231.5m82%
75o4 Mini High83.9%$0.0561.4m54%
76Claude Opus 4.6 (Reasoning)91.1%$0.1291.1m78%
77Ministral 3 3B60.7%$0.00022.9s25%
78Mistral NeMO53.6%$0.00023.1s14%
79MoonshotAI: Kimi K2.591.1%$0.0296.5m70%
80Gemini 3.1 Pro (Preview)93.8%$0.1682.6m82%
81Claude 3 Haiku42.0%$0.00136.8s6%
82Cohere Command R+ (Aug. 2024)36.6%$0.009218.4s17%
83Hermes 3 70B57.1%$0.00222.8m10%
84Rocinante 12B28.6%$0.000510.0s7%
86.00%

Individual Scenarios

Generic Prompt

Model # 1 # 2 # 3 # 4 # 5 # 6 # 7 Avg ▼
Claude Opus 4.6100100100100100100100100.0%
Claude Sonnet 4.6100100100100100100100100.0%
Claude Opus 4.5100100100100100100100100.0%
Claude Sonnet 4100100100100100100100100.0%
Claude Haiku 4.5100100100100100100100100.0%
Minimax M2.51001001001001001008898.2%
Qwen 3.5 Plus (2026-02-15)1001001001001001008898.2%
Gemini 2.5 Flash Lite100100100100100888896.4%
Z.AI GLM 4.61001001008888888892.9%
Claude Sonnet 4.51001001008888888892.9%
DeepSeek V3.21001001008888888892.9%
Grok 4100100888888888891.1%
Gemini 2.5 Flash100100888888888891.1%
Qwen 2.5 72B1001001008888887591.1%
Arcee AI: Trinity Large (Preview)100100888888888891.1%
GPT-4o, May 13th (temp=1)10088888888888889.3%
DeepSeek V3 (2024-12-26)10088888888888889.3%
Claude Opus 4.6 (Reasoning)8888888888888887.5%
Gemini 3.1 Pro (Preview)8888888888888887.5%
Claude Sonnet 4.6 (Reasoning)8888888888888887.5%
GPT-5 Mini8888888888888887.5%
GPT-5.18888888888888887.5%
GPT-58888888888888887.5%
Qwen 3.5 397B A17B8888888888888887.5%
Z.AI GLM 58888888888888887.5%
MoonshotAI: Kimi K2.58888888888888887.5%
ByteDance Seed 1.68888888888888887.5%
Gemini 3 Flash (Preview, Reasoning)8888888888888887.5%
Grok 4.1 Fast8888888888888887.5%
Aion 2.08888888888888887.5%
Gemini 3 Pro (Preview)8888888888888887.5%
Z.AI GLM 4.78888888888888887.5%
Gemini 2.5 Pro8888888888888887.5%
Claude Opus 48888888888888887.5%
Gemini 2.5 Flash (Reasoning)8888888888888887.5%
Z.AI GLM 4.58888888888888887.5%
Grok 4 Fast8888888888888887.5%
Gemini 2.5 Flash Lite (Reasoning)8888888888888887.5%
Mistral Large 38888888888888887.5%
GPT-4o, May 13th (temp=0)8888888888888887.5%
Gemini 3 Flash (Preview)8888888888888887.5%
DeepSeek-V2 Chat8888888888888887.5%
Z.AI GLM 4.7 Flash10088888888887587.5%
GPT-4.1 Mini8888888888888887.5%
Hermes 3 405B8888888888888887.5%
Mistral Large 28888888888888887.5%
DeepSeek V3.1100100100100100882587.5%
DeepSeek V3 (2025-03-24)10088888888887587.5%
Mistral Large8888888888888887.5%
Writer: Palmyra X510088888888887587.5%
GPT-4o Mini (temp=1)8888888888888887.5%
Mistral Small 3.2 24B8888888888888887.5%
Gemma 3 12B8888888888888887.5%
GPT-4o Mini (temp=0)8888888888888887.5%
GPT-5.28888888888887585.7%
Claude 3.7 Sonnet8888888888887585.7%
Gemma 3 27B8888888888887585.7%
ByteDance Seed 1.6 Flash10088888888757585.7%
Hermes 3 70B100100888888756385.7%
Ministral 3 14B8888888888887585.7%
GPT-4.18888888888757583.9%
Mistral Small Creative8888888888756382.1%
Arcee AI: Trinity Mini8888888875757582.1%
Llama 3.1 70B8888887575755076.8%
Mistral Medium 3.18875757575757576.8%
Gemma 3 4B8888757575756376.8%
Claude 3.5 Sonnet7575757575757575.0%
GPT-4o, Aug. 6th (temp=0)888888888888075.0%
o4 Mini High8888888875632573.2%
WizardLM 2 8x22b100100888888381373.2%
Llama 3.1 8B100100887563503873.2%
GPT-4o, Aug. 6th (temp=1)888888888863071.4%
o4 Mini8888888875252567.9%
Ministral 3 8B7575757563635067.9%
GPT-5 Nano8888757550503866.1%
Ministral 8B7575636363636366.1%
GPT-4.1 Nano1008875635038058.9%
Llama 3.1 Nemotron 70B8875636338382555.4%
Ministral 3B6363505050383850.0%
Ministral 3 3B6363502525251337.5%
Mistral NeMO38252513130016.1%
Cohere Command R+ (Aug. 2024)501313000010.7%
Rocinante 12B38131300008.9%
Claude 3 Haiku00000000.0%

Specific Prompt

Model # 1 # 2 # 3 # 4 # 5 # 6 # 7 Avg ▼
Gemini 3.1 Pro (Preview)100100100100100100100100.0%
Claude Sonnet 4.6 (Reasoning)100100100100100100100100.0%
GPT-5 Mini100100100100100100100100.0%
Claude Opus 4.6100100100100100100100100.0%
Claude Sonnet 4.6100100100100100100100100.0%
Claude Opus 4.5100100100100100100100100.0%
Grok 4.1 Fast100100100100100100100100.0%
Claude Sonnet 4100100100100100100100100.0%
Grok 4100100100100100100100100.0%
Claude Sonnet 4.5100100100100100100100100.0%
Grok 4 Fast100100100100100100100100.0%
Qwen 3.5 Plus (2026-02-15)100100100100100100100100.0%
Mistral Large 3100100100100100100100100.0%
Gemini 3 Flash (Preview)100100100100100100100100.0%
Claude Haiku 4.5100100100100100100100100.0%
Mistral Large 2100100100100100100100100.0%
Gemini 2.5 Flash Lite100100100100100100100100.0%
Gemini 2.5 Flash100100100100100100100100.0%
Mistral Large100100100100100100100100.0%
Writer: Palmyra X5100100100100100100100100.0%
Mistral Small 3.2 24B100100100100100100100100.0%
ByteDance Seed 1.6 Flash100100100100100100100100.0%
Mistral Small Creative100100100100100100100100.0%
Ministral 3 14B100100100100100100100100.0%
Ministral 3 8B100100100100100100100100.0%
Gemma 3 4B100100100100100100100100.0%
Ministral 8B100100100100100100100100.0%
Gemini 2.5 Pro1001001001001001008898.2%
o4 Mini1001001001001001008898.2%
DeepSeek V3.21001001001001001008898.2%
Gemma 3 12B1001001001001001008898.2%
Qwen 2.5 72B1001001001001001008898.2%
GPT-5.1100100100100100888896.4%
GPT-5100100100100100888896.4%
Minimax M2.5100100100100100888896.4%
Z.AI GLM 4.7100100100100100888896.4%
Z.AI GLM 4.5100100100100100888896.4%
Gemini 2.5 Flash Lite (Reasoning)100100100100100888896.4%
Claude Opus 4.6 (Reasoning)10010010010088888894.6%
Z.AI GLM 510010010010088888894.6%
MoonshotAI: Kimi K2.51001001001001001006394.6%
o4 Mini High10010010010088888894.6%
Z.AI GLM 4.610010010010088888894.6%
Claude Opus 410010010010088888894.6%
Qwen 3.5 397B A17B1001001008888888892.9%
Gemini 2.5 Flash (Reasoning)1001001008888888892.9%
GPT-4o, May 13th (temp=0)10010010010088887592.9%
DeepSeek V3 (2024-12-26)1001001008888888892.9%
Ministral 3B100100100100100886392.9%
GPT-5.2100100888888888891.1%
Gemini 3 Pro (Preview)100100888888888891.1%
GPT-5 Nano10010010010088886391.1%
GPT-4.1 Nano100100888888888891.1%
WizardLM 2 8x22b100100888888888891.1%
Arcee AI: Trinity Mini100100888888888891.1%
Mistral NeMO100100888888888891.1%
Gemini 3 Flash (Preview, Reasoning)10088888888888889.3%
GPT-4.1 Mini10088888888888889.3%
DeepSeek V3.1100100100100100883889.3%
ByteDance Seed 1.68888888888888887.5%
Claude 3.5 Sonnet8888888888888887.5%
Claude 3.7 Sonnet8888888888888887.5%
Hermes 3 405B8888888888888887.5%
GPT-4o, Aug. 6th (temp=0)1001001008875757587.5%
GPT-4o Mini (temp=1)8888888888888887.5%
GPT-4o Mini (temp=0)8888888888888887.5%
Gemma 3 27B8888888888888887.5%
Mistral Medium 3.18888888888888887.5%
Arcee AI: Trinity Large (Preview)8888888888888887.5%
Aion 2.010010010010010088083.9%
Claude 3 Haiku10088888875757583.9%
Ministral 3 3B10010010010063636383.9%
Z.AI GLM 4.7 Flash10010010010088503882.1%
GPT-4o, May 13th (temp=1)10088887575757582.1%
Llama 3.1 70B8888757575757578.6%
Llama 3.1 8B1001001008875632578.6%
GPT-4.18875757575757576.8%
DeepSeek V3 (2025-03-24)10010088888875076.8%
Llama 3.1 Nemotron 70B8875757575755073.2%
GPT-4o, Aug. 6th (temp=1)10010075757575071.4%
DeepSeek-V2 Chat100888888880064.3%
Cohere Command R+ (Aug. 2024)8875636363503862.5%
Rocinante 12B887575502525048.2%
Hermes 3 70B1001000000028.6%