Mara pronouns preserved (coreference test)

Test: Text Replacement

Avg. Score
82.0%
Scenarios
2

Overall Performance

Rank ▲ Model Score Avg. Cost Avg. Time Stability
1Mistral Small Creative100.0%$0.00023.2s100%
2Gemini 2.5 Flash100.0%$0.00172.4s100%
3Grok 4 Fast100.0%$0.00086.0s100%
4Gemini 3 Flash (Preview)100.0%$0.00213.6s100%
5GPT-4.1 Mini100.0%$0.00127.1s100%
6Mistral Medium 3.1100.0%$0.00156.5s100%
7Stealth: Healer Alpha100.0%$0.000014.7s100%
8Gemini 3.1 Flash Lite (Preview)99.4%$0.00111.9s95%
9Claude Haiku 4.5100.0%$0.00423.3s100%
10Llama 3.1 Nemotron 70B100.0%$0.001616.3s100%
11GPT-4.1100.0%$0.00614.7s100%
12GPT-4o, Aug. 6th (temp=0)100.0%$0.00762.8s100%
13GPT-4o, Aug. 6th (temp=1)100.0%$0.00763.2s100%
14Hermes 3 405B100.0%$0.001425.1s100%
15Qwen 3.5 Plus (2026-02-15)98.7%$0.00177.7s91%
16Grok 4.1 Fast98.7%$0.000910.9s91%
17Qwen 3 32B100.0%$0.000633.7s100%
18GPT-5.4100.0%$0.0106.1s100%
19Stealth: Hunter Alpha98.7%$0.000017.4s91%
20Gemini 3 Flash (Preview, Reasoning)100.0%$0.008313.3s100%
21GPT-5.4 (Reasoning, Low)100.0%$0.0116.7s100%
22GPT-4o, May 13th (temp=1)100.0%$0.0124.0s100%
23GPT-4o, May 13th (temp=0)100.0%$0.0124.1s100%
24Claude Sonnet 4.6100.0%$0.0135.1s100%
25Claude Sonnet 4.5100.0%$0.0135.2s100%
26Claude 3.7 Sonnet100.0%$0.0136.4s100%
27Claude Sonnet 4100.0%$0.0136.5s100%
28Z.AI GLM 4.5100.0%$0.004132.4s100%
29Grok 4.20 (Beta, Reasoning)100.0%$0.0136.4s100%
30Gemini 2.5 Flash (Reasoning)98.7%$0.00598.8s91%
31ByteDance Seed 1.6100.0%$0.003537.7s100%
32Aion 2.0100.0%$0.003847.6s100%
33o4 Mini100.0%$0.01421.5s100%
34ByteDance Seed 2.0 Lite100.0%$0.004552.2s100%
35GPT-5.1100.0%$0.01716.5s100%
36Claude Opus 4.5100.0%$0.0215.8s100%
37Claude Opus 4.6100.0%$0.0216.5s100%
38Z.AI GLM 5100.0%$0.007454.4s100%
39Z.AI GLM 4.7100.0%$0.00591.0m100%
40Qwen 3.5 27B100.0%$0.01249.5s100%
41Claude 3.5 Sonnet100.0%$0.02510.7s100%
42MiniMax M2.7100.0%$0.00651.1m100%
43Inception Mercury 287.0%$0.00192.6s68%
44Gemma 3 27B89.0%$0.000218.1s72%
45Z.AI GLM 4.698.7%$0.00501.0m91%
46o4 Mini High100.0%$0.02233.0s100%
47Gemini 2.5 Pro100.0%$0.02820.4s100%
48Qwen 3.5 122B100.0%$0.02046.9s100%
49Grok 4100.0%$0.02530.9s100%
50Gemini 3 Pro (Preview)100.0%$0.03121.3s100%
51ByteDance Seed 1.6 Flash90.3%$0.000712.3s47%
52Claude Sonnet 4.6 (Reasoning)100.0%$0.03820.1s100%
53MoonshotAI: Kimi K2.5100.0%$0.00851.8m100%
54Ministral 3 3B80.5%$0.00012.2s50%
55GPT-5.4 Nano (Reasoning, Low)86.4%$0.00098.4s44%
56Qwen3 235B A22B Instruct 250781.8%$0.000516.2s52%
57Claude Opus 4.6 (Reasoning)100.0%$0.04316.8s100%
58GPT-5.292.9%$0.0107.1s48%
59Llama 3.1 70B92.9%$0.000536.2s48%
60Arcee AI: Trinity Mini63.6%$0.00027.1s64%
61Gemini 3.1 Pro (Preview)100.0%$0.03836.2s100%
62GPT-5100.0%$0.03549.7s100%
63Z.AI GLM 5 Turbo92.9%$0.009519.7s48%
64Claude 3 Haiku78.6%$0.00105.0s42%
65Ministral 3B77.3%$0.00012.3s41%
66GPT-5.4 Nano81.8%$0.00093.5s37%
67Gemma 3 4B77.3%$0.00016.5s42%
68MiniMax M2.595.5%$0.00171.6m72%
69GPT-5 Mini92.2%$0.005732.9s49%
70Inception Mercury69.5%$0.00055.2s41%
71Mistral Small 4 (Reasoning)84.4%$0.001916.2s30%
72GPT-5.4 (Reasoning)92.9%$0.02016.0s48%
73GPT-5.4 Mini (Reasoning)85.7%$0.00716.2s30%
74Writer: Palmyra X577.3%$0.004013.0s37%
75Claude Opus 4100.0%$0.0638.9s100%
76Qwen 2.5 72B78.6%$0.000311.5s18%
77Grok 4.20 (Beta)75.3%$0.00382.0s21%
78GPT-5.4 Nano (Reasoning)68.8%$0.00103.7s20%
79Nemotron 3 Super82.5%$0.000057.7s31%
80GPT-5.4 Mini (Reasoning, Low)72.1%$0.00362.9s19%
81GPT-5.4 Mini55.8%$0.00312.3s27%
82Mistral Small 461.7%$0.00043.4s15%
83Qwen 3.5 9B89.6%$0.00141.9m45%
84DeepSeek-V2 Chat71.4%$0.000917.7s10%
85Ministral 3 14B61.0%$0.00034.5s13%
86Mistral Large71.4%$0.00508.4s10%
87DeepSeek V3 (2025-03-24)68.8%$0.000726.6s11%
88GPT-5 Nano74.0%$0.00291.0m24%
89Mistral Small 3.2 24B56.5%$0.00025.5s6%
90Qwen 3.5 35B80.5%$0.01652.7s30%
91DeepSeek V3 (2024-12-26)64.3%$0.000918.3s4%
92Ministral 3 8B52.6%$0.00023.7s3%
93Qwen 3.5 Flash71.4%$0.00331.0m17%
94Llama 3.1 8B54.5%$0.000110.1s3%
95Ministral 8B50.6%$0.00013.9s1%
96Gemini 2.5 Flash Lite50.0%$0.00031.9s0%
97Mistral NeMO50.0%$0.00022.7s0%
98Rocinante 12B50.0%$0.00049.0s3%
99Arcee AI: Trinity Large (Preview)54.5%$0.000027.1s3%
100Mistral Large 350.0%$0.00128.4s0%
101Z.AI GLM 4.7 Flash63.0%$0.00161.1m15%
102Mistral Large 250.0%$0.00508.4s0%
103Qwen 3.5 397B A17B85.7%$0.00752.2m30%
104GPT-4.1 Nano40.3%$0.00034.2s0%
105WizardLM 2 8x22b57.1%$0.000840.0s1%
106GPT-4o Mini (temp=0)31.8%$0.000510.8s12%
107GPT-4o Mini (temp=1)31.8%$0.000511.0s12%
108Gemini 2.5 Flash Lite (Reasoning)44.2%$0.002315.5s2%
109Gemma 3 12B29.2%$0.00019.6s8%
110DeepSeek V3.250.0%$0.000638.9s0%
111DeepSeek V3.146.1%$0.000643.9s1%
112Hermes 3 70B35.7%$0.000421.2s0%
113Cohere Command R+ (Aug. 2024)50.0%$0.007932.6s0%
114LFM2 24B6.5%$0.000114.7s0%
115ByteDance Seed 2.0 Mini50.0%$0.00202.0m0%
116Nemotron 3 Nano57.8%$0.00223.2m17%
81.96%

Individual Scenarios

Generic Prompt

Model # 1 # 2 # 3 # 4 # 5 # 6 # 7 Avg ▼
Claude Opus 4.6 (Reasoning)100100100100100100100100.0%
Gemini 3.1 Pro (Preview)100100100100100100100100.0%
Claude Sonnet 4.6 (Reasoning)100100100100100100100100.0%
GPT-5.1100100100100100100100100.0%
Claude Opus 4.6100100100100100100100100.0%
GPT-5100100100100100100100100.0%
Qwen 3.5 122B100100100100100100100100.0%
Grok 4.20 (Beta, Reasoning)100100100100100100100100.0%
GPT-5.4 (Reasoning, Low)100100100100100100100100.0%
Z.AI GLM 5100100100100100100100100.0%
Claude Sonnet 4.6100100100100100100100100.0%
MoonshotAI: Kimi K2.5100100100100100100100100.0%
Qwen 3.5 27B100100100100100100100100.0%
ByteDance Seed 1.6100100100100100100100100.0%
Gemini 3 Flash (Preview, Reasoning)100100100100100100100100.0%
o4 Mini High100100100100100100100100.0%
Claude Opus 4.5100100100100100100100100.0%
Grok 4.1 Fast100100100100100100100100.0%
Aion 2.0100100100100100100100100.0%
Z.AI GLM 4.6100100100100100100100100.0%
MiniMax M2.7100100100100100100100100.0%
Gemini 3 Pro (Preview)100100100100100100100100.0%
Claude Sonnet 4100100100100100100100100.0%
Z.AI GLM 4.7100100100100100100100100.0%
GPT-4.1100100100100100100100100.0%
Gemini 2.5 Pro100100100100100100100100.0%
o4 Mini100100100100100100100100.0%
Grok 4100100100100100100100100.0%
Claude Sonnet 4.5100100100100100100100100.0%
Claude Opus 4100100100100100100100100.0%
Stealth: Hunter Alpha100100100100100100100100.0%
Z.AI GLM 4.5100100100100100100100100.0%
Grok 4 Fast100100100100100100100100.0%
Qwen 3.5 Plus (2026-02-15)100100100100100100100100.0%
Stealth: Healer Alpha100100100100100100100100.0%
Gemini 3.1 Flash Lite (Preview)100100100100100100100100.0%
GPT-4o, May 13th (temp=0)100100100100100100100100.0%
Gemini 3 Flash (Preview)100100100100100100100100.0%
Claude Haiku 4.5100100100100100100100100.0%
ByteDance Seed 2.0 Lite100100100100100100100100.0%
GPT-5.4100100100100100100100100.0%
Claude 3.5 Sonnet100100100100100100100100.0%
GPT-4o, May 13th (temp=1)100100100100100100100100.0%
Claude 3.7 Sonnet100100100100100100100100.0%
GPT-4.1 Mini100100100100100100100100.0%
Hermes 3 405B100100100100100100100100.0%
GPT-4o, Aug. 6th (temp=1)100100100100100100100100.0%
GPT-4o, Aug. 6th (temp=0)100100100100100100100100.0%
Qwen 3 32B100100100100100100100100.0%
Gemini 2.5 Flash100100100100100100100100.0%
Mistral Medium 3.1100100100100100100100100.0%
Llama 3.1 Nemotron 70B100100100100100100100100.0%
Mistral Small Creative100100100100100100100100.0%
Gemini 2.5 Flash (Reasoning)1001001001001001008297.4%
Inception Mercury 210010010010082828292.2%
MiniMax M2.5100100100100100914590.9%
Z.AI GLM 5 Turbo100100100100100100085.7%
GPT-5.4 (Reasoning)100100100100100100085.7%
GPT-5.2100100100100100100085.7%
Qwen 3.5 9B100100100100100100085.7%
Llama 3.1 70B100100100100100100085.7%
ByteDance Seed 1.6 Flash100100100100100100085.7%
GPT-5 Mini10010010010010091084.4%
GPT-5.4 Nano (Reasoning, Low)10010010010010045077.9%
Gemma 3 27B9191737373737377.9%
Inception Mercury10082828282732775.3%
Qwen 3.5 397B A17B1001001001001000071.4%
GPT-5.4 Mini (Reasoning)1001001001001000071.4%
Mistral Small 4 (Reasoning)100100100100820068.8%
Nemotron 3 Super1001009182820064.9%
Qwen3 235B A22B Instruct 25076464646464646463.6%
GPT-5.4 Nano100100100733627963.6%
Arcee AI: Trinity Mini6464646464646463.6%
Qwen 3.5 35B10010010064640061.0%
Ministral 3 3B6464646464555561.0%
GPT-5 Nano1001001009199058.4%
Ministral 3B6464645555555558.4%
GPT-5.4 Mini (Reasoning, Low)10010010010000057.1%
Qwen 2.5 72B10010010010000057.1%
Claude 3 Haiku6464646464641857.1%
Nemotron 3 Nano1001008255450054.5%
Writer: Palmyra X5646464646464054.5%
Gemma 3 4B5555555555555554.5%
Grok 4.20 (Beta)10010010027270050.6%
GPT-5.4 Nano (Reasoning)100100735500046.8%
Qwen 3.5 Flash100100643600042.9%
DeepSeek-V2 Chat100100100000042.9%
Mistral Large100100100000042.9%
Z.AI GLM 4.7 Flash10010073000039.0%
DeepSeek V3 (2025-03-24)10010064000037.7%
GPT-5.4 Mini363636363627029.9%
DeepSeek V3 (2024-12-26)1001000000028.6%
Hermes 3 70B1001000000028.6%
Mistral Small 3.2 24B2727272727272727.3%
Mistral Small 4454518181818023.4%
Ministral 3 14B5555181890022.1%
Rocinante 12B100270000018.2%
Gemini 2.5 Flash Lite (Reasoning)10000000014.3%
WizardLM 2 8x22b10000000014.3%
Arcee AI: Trinity Large (Preview)640000009.1%
Llama 3.1 8B640000009.1%
Ministral 3 8B1818000005.2%
Ministral 8B90000001.3%
ByteDance Seed 2.0 Mini00000000.0%
Mistral Large 300000000.0%
Mistral Large 200000000.0%
DeepSeek V3.100000000.0%
DeepSeek V3.200000000.0%
Gemini 2.5 Flash Lite00000000.0%
GPT-4o Mini (temp=1)00000000.0%
Gemma 3 12B00000000.0%
GPT-4o Mini (temp=0)00000000.0%
GPT-4.1 Nano00000000.0%
Cohere Command R+ (Aug. 2024)00000000.0%
Mistral NeMO00000000.0%
LFM2 24B00000000.0%

Specific Prompt

Model # 1 # 2 # 3 # 4 # 5 # 6 # 7 Avg ▼
Claude Opus 4.6 (Reasoning)100100100100100100100100.0%
Gemini 3.1 Pro (Preview)100100100100100100100100.0%
Z.AI GLM 5 Turbo100100100100100100100100.0%
Claude Sonnet 4.6 (Reasoning)100100100100100100100100.0%
GPT-5.4 (Reasoning)100100100100100100100100.0%
GPT-5 Mini100100100100100100100100.0%
GPT-5.1100100100100100100100100.0%
Claude Opus 4.6100100100100100100100100.0%
GPT-5100100100100100100100100.0%
Qwen 3.5 397B A17B100100100100100100100100.0%
Qwen 3.5 122B100100100100100100100100.0%
Grok 4.20 (Beta, Reasoning)100100100100100100100100.0%
GPT-5.4 (Reasoning, Low)100100100100100100100100.0%
Z.AI GLM 5100100100100100100100100.0%
Claude Sonnet 4.6100100100100100100100100.0%
MoonshotAI: Kimi K2.5100100100100100100100100.0%
Qwen 3.5 27B100100100100100100100100.0%
ByteDance Seed 1.6100100100100100100100100.0%
GPT-5.4 Mini (Reasoning)100100100100100100100100.0%
Gemini 3 Flash (Preview, Reasoning)100100100100100100100100.0%
o4 Mini High100100100100100100100100.0%
GPT-5.2100100100100100100100100.0%
Claude Opus 4.5100100100100100100100100.0%
Aion 2.0100100100100100100100100.0%
MiniMax M2.7100100100100100100100100.0%
Gemini 3 Pro (Preview)100100100100100100100100.0%
Claude Sonnet 4100100100100100100100100.0%
MiniMax M2.5100100100100100100100100.0%
Z.AI GLM 4.7100100100100100100100100.0%
GPT-4.1100100100100100100100100.0%
Gemini 2.5 Pro100100100100100100100100.0%
o4 Mini100100100100100100100100.0%
Grok 4100100100100100100100100.0%
Claude Sonnet 4.5100100100100100100100100.0%
Qwen 3.5 35B100100100100100100100100.0%
Claude Opus 4100100100100100100100100.0%
ByteDance Seed 2.0 Mini100100100100100100100100.0%
Gemini 2.5 Flash (Reasoning)100100100100100100100100.0%
Qwen 3.5 Flash100100100100100100100100.0%
Z.AI GLM 4.5100100100100100100100100.0%
Grok 4 Fast100100100100100100100100.0%
Stealth: Healer Alpha100100100100100100100100.0%
Mistral Large 3100100100100100100100100.0%
GPT-4o, May 13th (temp=0)100100100100100100100100.0%
Gemini 3 Flash (Preview)100100100100100100100100.0%
Claude Haiku 4.5100100100100100100100100.0%
DeepSeek-V2 Chat100100100100100100100100.0%
ByteDance Seed 2.0 Lite100100100100100100100100.0%
Nemotron 3 Super100100100100100100100100.0%
GPT-5.4100100100100100100100100.0%
Claude 3.5 Sonnet100100100100100100100100.0%
Grok 4.20 (Beta)100100100100100100100100.0%
GPT-4o, May 13th (temp=1)100100100100100100100100.0%
DeepSeek V3 (2024-12-26)100100100100100100100100.0%
Claude 3.7 Sonnet100100100100100100100100.0%
GPT-4.1 Mini100100100100100100100100.0%
Hermes 3 405B100100100100100100100100.0%
GPT-4o, Aug. 6th (temp=1)100100100100100100100100.0%
GPT-4o, Aug. 6th (temp=0)100100100100100100100100.0%
Mistral Large 2100100100100100100100100.0%
Mistral Small 4 (Reasoning)100100100100100100100100.0%
DeepSeek V3.2100100100100100100100100.0%
Qwen 3 32B100100100100100100100100.0%
DeepSeek V3 (2025-03-24)100100100100100100100100.0%
Gemini 2.5 Flash Lite100100100100100100100100.0%
Gemini 2.5 Flash100100100100100100100100.0%
Mistral Large100100100100100100100100.0%
Qwen3 235B A22B Instruct 2507100100100100100100100100.0%
Writer: Palmyra X5100100100100100100100100.0%
Llama 3.1 70B100100100100100100100100.0%
Gemma 3 27B100100100100100100100100.0%
Mistral Medium 3.1100100100100100100100100.0%
Mistral Small 4100100100100100100100100.0%
Qwen 2.5 72B100100100100100100100100.0%
Llama 3.1 Nemotron 70B100100100100100100100100.0%
GPT-5.4 Nano100100100100100100100100.0%
Arcee AI: Trinity Large (Preview)100100100100100100100100.0%
Mistral Small Creative100100100100100100100100.0%
Ministral 3 14B100100100100100100100100.0%
Ministral 3 8B100100100100100100100100.0%
Claude 3 Haiku100100100100100100100100.0%
WizardLM 2 8x22b100100100100100100100100.0%
Cohere Command R+ (Aug. 2024)100100100100100100100100.0%
Gemma 3 4B100100100100100100100100.0%
Ministral 3 3B100100100100100100100100.0%
Mistral NeMO100100100100100100100100.0%
Ministral 8B100100100100100100100100.0%
Llama 3.1 8B100100100100100100100100.0%
Gemini 3.1 Flash Lite (Preview)1001001001001001009198.7%
Grok 4.1 Fast1001001001001001008297.4%
Z.AI GLM 4.61001001001001001008297.4%
Stealth: Hunter Alpha1001001001001001008297.4%
Qwen 3.5 Plus (2026-02-15)1001001001001001008297.4%
Ministral 3B1001001001001001007396.1%
GPT-5.4 Nano (Reasoning, Low)100100100100100828294.8%
ByteDance Seed 1.6 Flash1001001001001001006494.8%
Qwen 3.5 9B1001001001001001005593.5%
DeepSeek V3.1100100100100100915592.2%
GPT-5.4 Nano (Reasoning)1001001009182828290.9%
GPT-5 Nano1001001008282828289.6%
GPT-5.4 Mini (Reasoning, Low)100100828282828287.0%
Z.AI GLM 4.7 Flash10091918282828287.0%
Mistral Small 3.2 24B100100100100100100085.7%
Inception Mercury 28282828282828281.8%
GPT-5.4 Mini8282828282828281.8%
Rocinante 12B1001001001009182081.8%
GPT-4.1 Nano1001001001008282080.5%
Gemini 2.5 Flash Lite (Reasoning)1001001001007345074.0%
Inception Mercury10082827355361863.6%
GPT-4o Mini (temp=1)6464646464646463.6%
GPT-4o Mini (temp=0)6464646464646463.6%
Arcee AI: Trinity Mini6464646464646463.6%
Nemotron 3 Nano1009182644545061.0%
Gemma 3 12B10055555555454558.4%
Hermes 3 70B100100100000042.9%
LFM2 24B1818181899013.0%