Mara pronouns preserved (coreference test)

Test: Text Replacement

Avg. Score
81.8%
Scenarios
2

Overall Performance

Rank ▲ Model Score Avg. Cost Avg. Time Stability
1Mistral Small Creative100.0%$0.00023.2s100%
2Gemini 2.5 Flash100.0%$0.00172.4s100%
3Grok 4 Fast100.0%$0.00086.0s100%
4Gemini 3 Flash (Preview)100.0%$0.00213.6s100%
5Mistral Medium 3.1100.0%$0.00156.5s100%
6GPT-4.1 Mini100.0%$0.00127.1s100%
7Claude Haiku 4.5100.0%$0.00423.3s100%
8GPT-4.1100.0%$0.00614.7s100%
9GPT-4o, Aug. 6th (temp=0)100.0%$0.00762.8s100%
10GPT-4o, Aug. 6th (temp=1)100.0%$0.00763.2s100%
11Llama 3.1 Nemotron 70B100.0%$0.001616.3s100%
12Qwen 3.5 Plus (2026-02-15)98.7%$0.00177.7s91%
13Grok 4.1 Fast98.7%$0.000910.9s91%
14Hermes 3 405B100.0%$0.001425.1s100%
15GPT-4o, May 13th (temp=1)100.0%$0.0124.0s100%
16GPT-4o, May 13th (temp=0)100.0%$0.0124.1s100%
17Gemini 3 Flash (Preview, Reasoning)100.0%$0.008313.3s100%
18Claude Sonnet 4.6100.0%$0.0135.1s100%
19Claude Sonnet 4.5100.0%$0.0135.2s100%
20Claude 3.7 Sonnet100.0%$0.0136.4s100%
21Claude Sonnet 4100.0%$0.0136.5s100%
22Gemini 2.5 Flash (Reasoning)98.7%$0.00598.8s91%
23Z.AI GLM 4.5100.0%$0.004132.4s100%
24ByteDance Seed 1.6100.0%$0.003537.7s100%
25Claude Opus 4.5100.0%$0.0215.8s100%
26Claude Opus 4.6100.0%$0.0216.5s100%
27GPT-5.1100.0%$0.01716.5s100%
28o4 Mini100.0%$0.01421.5s100%
29Aion 2.0100.0%$0.003847.6s100%
30Claude 3.5 Sonnet100.0%$0.02510.7s100%
31Z.AI GLM 5100.0%$0.007454.4s100%
32Gemma 3 27B89.0%$0.000218.1s72%
33Z.AI GLM 4.7100.0%$0.00591.0m100%
34Gemini 2.5 Pro100.0%$0.02820.4s100%
35o4 Mini High100.0%$0.02233.0s100%
36Grok 4100.0%$0.02530.9s100%
37Z.AI GLM 4.698.7%$0.00501.0m91%
38Gemini 3 Pro (Preview)100.0%$0.03121.3s100%
39Claude Sonnet 4.6 (Reasoning)100.0%$0.03820.1s100%
40ByteDance Seed 1.6 Flash90.3%$0.000712.3s47%
41Ministral 3 3B80.5%$0.00012.2s50%
42Claude Opus 4.6 (Reasoning)100.0%$0.04316.8s100%
43GPT-5.292.9%$0.0107.1s48%
44Gemini 3.1 Pro (Preview)100.0%$0.03836.2s100%
45Llama 3.1 70B92.9%$0.000536.2s48%
46Ministral 3B77.3%$0.00012.3s41%
47Arcee AI: Trinity Mini63.6%$0.00027.1s64%
48Claude 3 Haiku78.6%$0.00105.0s42%
49GPT-5100.0%$0.03549.7s100%
50Gemma 3 4B77.3%$0.00016.5s42%
51GPT-5 Mini92.2%$0.005732.9s49%
52MoonshotAI: Kimi K2.5100.0%$0.00851.8m100%
53Claude Opus 4100.0%$0.0638.9s100%
54Minimax M2.595.5%$0.00171.6m72%
55Writer: Palmyra X577.3%$0.004013.0s37%
56Qwen 2.5 72B78.6%$0.000311.5s18%
57Mistral Large71.4%$0.00508.4s10%
58DeepSeek-V2 Chat71.4%$0.000917.7s10%
59Ministral 3 14B61.0%$0.00034.5s13%
60DeepSeek V3 (2025-03-24)68.8%$0.000726.6s11%
61Mistral Small 3.2 24B56.5%$0.00025.5s6%
62DeepSeek V3 (2024-12-26)64.3%$0.000918.3s4%
63GPT-5 Nano74.0%$0.00291.0m24%
64Ministral 3 8B52.6%$0.00023.7s3%
65Llama 3.1 8B54.5%$0.000110.1s3%
66Ministral 8B50.6%$0.00013.9s1%
67Gemini 2.5 Flash Lite50.0%$0.00031.9s0%
68Mistral NeMO50.0%$0.00022.7s0%
69Rocinante 12B50.0%$0.00049.0s3%
70Mistral Large 350.0%$0.00128.4s0%
71Arcee AI: Trinity Large (Preview)54.5%$0.000027.1s3%
72Mistral Large 250.0%$0.00508.4s0%
73GPT-4.1 Nano40.3%$0.00034.2s0%
74Z.AI GLM 4.7 Flash63.0%$0.00161.1m15%
75WizardLM 2 8x22b57.1%$0.000840.0s1%
76Gemini 2.5 Flash Lite (Reasoning)44.2%$0.002315.5s2%
77GPT-4o Mini (temp=0)31.8%$0.000510.8s12%
78GPT-4o Mini (temp=1)31.8%$0.000511.0s12%
79Gemma 3 12B29.2%$0.00019.6s8%
80DeepSeek V3.250.0%$0.000638.9s0%
81Qwen 3.5 397B A17B85.7%$0.00752.2m30%
82Hermes 3 70B35.7%$0.000421.2s0%
83Cohere Command R+ (Aug. 2024)50.0%$0.007932.6s0%
84DeepSeek V3.146.1%$0.000643.9s1%
81.85%

Individual Scenarios

Generic Prompt

Model # 1 # 2 # 3 # 4 # 5 # 6 # 7 Avg ▼
Claude Opus 4.6 (Reasoning)100100100100100100100100.0%
Gemini 3.1 Pro (Preview)100100100100100100100100.0%
Claude Sonnet 4.6 (Reasoning)100100100100100100100100.0%
GPT-5.1100100100100100100100100.0%
Claude Opus 4.6100100100100100100100100.0%
GPT-5100100100100100100100100.0%
Z.AI GLM 5100100100100100100100100.0%
Claude Sonnet 4.6100100100100100100100100.0%
MoonshotAI: Kimi K2.5100100100100100100100100.0%
ByteDance Seed 1.6100100100100100100100100.0%
Gemini 3 Flash (Preview, Reasoning)100100100100100100100100.0%
o4 Mini High100100100100100100100100.0%
Claude Opus 4.5100100100100100100100100.0%
Grok 4.1 Fast100100100100100100100100.0%
Aion 2.0100100100100100100100100.0%
Z.AI GLM 4.6100100100100100100100100.0%
Gemini 3 Pro (Preview)100100100100100100100100.0%
Claude Sonnet 4100100100100100100100100.0%
Z.AI GLM 4.7100100100100100100100100.0%
GPT-4.1100100100100100100100100.0%
Gemini 2.5 Pro100100100100100100100100.0%
o4 Mini100100100100100100100100.0%
Grok 4100100100100100100100100.0%
Claude Sonnet 4.5100100100100100100100100.0%
Claude Opus 4100100100100100100100100.0%
Z.AI GLM 4.5100100100100100100100100.0%
Grok 4 Fast100100100100100100100100.0%
Qwen 3.5 Plus (2026-02-15)100100100100100100100100.0%
GPT-4o, May 13th (temp=0)100100100100100100100100.0%
Gemini 3 Flash (Preview)100100100100100100100100.0%
Claude Haiku 4.5100100100100100100100100.0%
Claude 3.5 Sonnet100100100100100100100100.0%
GPT-4o, May 13th (temp=1)100100100100100100100100.0%
Claude 3.7 Sonnet100100100100100100100100.0%
GPT-4.1 Mini100100100100100100100100.0%
Hermes 3 405B100100100100100100100100.0%
GPT-4o, Aug. 6th (temp=1)100100100100100100100100.0%
GPT-4o, Aug. 6th (temp=0)100100100100100100100100.0%
Gemini 2.5 Flash100100100100100100100100.0%
Mistral Medium 3.1100100100100100100100100.0%
Llama 3.1 Nemotron 70B100100100100100100100100.0%
Mistral Small Creative100100100100100100100100.0%
Gemini 2.5 Flash (Reasoning)1001001001001001008297.4%
Minimax M2.5100100100100100914590.9%
GPT-5.2100100100100100100085.7%
Llama 3.1 70B100100100100100100085.7%
ByteDance Seed 1.6 Flash100100100100100100085.7%
GPT-5 Mini10010010010010091084.4%
Gemma 3 27B9191737373737377.9%
Qwen 3.5 397B A17B1001001001001000071.4%
Arcee AI: Trinity Mini6464646464646463.6%
Ministral 3 3B6464646464555561.0%
GPT-5 Nano1001001009199058.4%
Ministral 3B6464645555555558.4%
Qwen 2.5 72B10010010010000057.1%
Claude 3 Haiku6464646464641857.1%
Writer: Palmyra X5646464646464054.5%
Gemma 3 4B5555555555555554.5%
DeepSeek-V2 Chat100100100000042.9%
Mistral Large100100100000042.9%
Z.AI GLM 4.7 Flash10010073000039.0%
DeepSeek V3 (2025-03-24)10010064000037.7%
DeepSeek V3 (2024-12-26)1001000000028.6%
Hermes 3 70B1001000000028.6%
Mistral Small 3.2 24B2727272727272727.3%
Ministral 3 14B5555181890022.1%
Rocinante 12B100270000018.2%
Gemini 2.5 Flash Lite (Reasoning)10000000014.3%
WizardLM 2 8x22b10000000014.3%
Arcee AI: Trinity Large (Preview)640000009.1%
Llama 3.1 8B640000009.1%
Ministral 3 8B1818000005.2%
Ministral 8B90000001.3%
Mistral Large 300000000.0%
Mistral Large 200000000.0%
DeepSeek V3.100000000.0%
DeepSeek V3.200000000.0%
Gemini 2.5 Flash Lite00000000.0%
GPT-4o Mini (temp=1)00000000.0%
Gemma 3 12B00000000.0%
GPT-4o Mini (temp=0)00000000.0%
GPT-4.1 Nano00000000.0%
Cohere Command R+ (Aug. 2024)00000000.0%
Mistral NeMO00000000.0%

Specific Prompt

Model # 1 # 2 # 3 # 4 # 5 # 6 # 7 Avg ▼
Claude Opus 4.6 (Reasoning)100100100100100100100100.0%
Gemini 3.1 Pro (Preview)100100100100100100100100.0%
Claude Sonnet 4.6 (Reasoning)100100100100100100100100.0%
GPT-5 Mini100100100100100100100100.0%
GPT-5.1100100100100100100100100.0%
Claude Opus 4.6100100100100100100100100.0%
GPT-5100100100100100100100100.0%
Qwen 3.5 397B A17B100100100100100100100100.0%
Z.AI GLM 5100100100100100100100100.0%
Claude Sonnet 4.6100100100100100100100100.0%
MoonshotAI: Kimi K2.5100100100100100100100100.0%
ByteDance Seed 1.6100100100100100100100100.0%
Gemini 3 Flash (Preview, Reasoning)100100100100100100100100.0%
o4 Mini High100100100100100100100100.0%
GPT-5.2100100100100100100100100.0%
Claude Opus 4.5100100100100100100100100.0%
Aion 2.0100100100100100100100100.0%
Gemini 3 Pro (Preview)100100100100100100100100.0%
Claude Sonnet 4100100100100100100100100.0%
Minimax M2.5100100100100100100100100.0%
Z.AI GLM 4.7100100100100100100100100.0%
GPT-4.1100100100100100100100100.0%
Gemini 2.5 Pro100100100100100100100100.0%
o4 Mini100100100100100100100100.0%
Grok 4100100100100100100100100.0%
Claude Sonnet 4.5100100100100100100100100.0%
Claude Opus 4100100100100100100100100.0%
Gemini 2.5 Flash (Reasoning)100100100100100100100100.0%
Z.AI GLM 4.5100100100100100100100100.0%
Grok 4 Fast100100100100100100100100.0%
Mistral Large 3100100100100100100100100.0%
GPT-4o, May 13th (temp=0)100100100100100100100100.0%
Gemini 3 Flash (Preview)100100100100100100100100.0%
Claude Haiku 4.5100100100100100100100100.0%
DeepSeek-V2 Chat100100100100100100100100.0%
Claude 3.5 Sonnet100100100100100100100100.0%
GPT-4o, May 13th (temp=1)100100100100100100100100.0%
DeepSeek V3 (2024-12-26)100100100100100100100100.0%
Claude 3.7 Sonnet100100100100100100100100.0%
GPT-4.1 Mini100100100100100100100100.0%
Hermes 3 405B100100100100100100100100.0%
GPT-4o, Aug. 6th (temp=1)100100100100100100100100.0%
GPT-4o, Aug. 6th (temp=0)100100100100100100100100.0%
Mistral Large 2100100100100100100100100.0%
DeepSeek V3.2100100100100100100100100.0%
DeepSeek V3 (2025-03-24)100100100100100100100100.0%
Gemini 2.5 Flash Lite100100100100100100100100.0%
Gemini 2.5 Flash100100100100100100100100.0%
Mistral Large100100100100100100100100.0%
Writer: Palmyra X5100100100100100100100100.0%
Llama 3.1 70B100100100100100100100100.0%
Gemma 3 27B100100100100100100100100.0%
Mistral Medium 3.1100100100100100100100100.0%
Qwen 2.5 72B100100100100100100100100.0%
Llama 3.1 Nemotron 70B100100100100100100100100.0%
Arcee AI: Trinity Large (Preview)100100100100100100100100.0%
Mistral Small Creative100100100100100100100100.0%
Ministral 3 14B100100100100100100100100.0%
Ministral 3 8B100100100100100100100100.0%
Claude 3 Haiku100100100100100100100100.0%
WizardLM 2 8x22b100100100100100100100100.0%
Cohere Command R+ (Aug. 2024)100100100100100100100100.0%
Gemma 3 4B100100100100100100100100.0%
Ministral 3 3B100100100100100100100100.0%
Mistral NeMO100100100100100100100100.0%
Ministral 8B100100100100100100100100.0%
Llama 3.1 8B100100100100100100100100.0%
Grok 4.1 Fast1001001001001001008297.4%
Z.AI GLM 4.61001001001001001008297.4%
Qwen 3.5 Plus (2026-02-15)1001001001001001008297.4%
Ministral 3B1001001001001001007396.1%
ByteDance Seed 1.6 Flash1001001001001001006494.8%
DeepSeek V3.1100100100100100915592.2%
GPT-5 Nano1001001008282828289.6%
Z.AI GLM 4.7 Flash10091918282828287.0%
Mistral Small 3.2 24B100100100100100100085.7%
Rocinante 12B1001001001009182081.8%
GPT-4.1 Nano1001001001008282080.5%
Gemini 2.5 Flash Lite (Reasoning)1001001001007345074.0%
GPT-4o Mini (temp=1)6464646464646463.6%
GPT-4o Mini (temp=0)6464646464646463.6%
Arcee AI: Trinity Mini6464646464646463.6%
Gemma 3 12B10055555555454558.4%
Hermes 3 70B100100100000042.9%