Language Writing

Can the model generate text in different languages?

Price-Performance Score Distribution (Top 20)

Click a model name to view its detail page.

ScoreCostTime
GPT-4.1 Nano93%$0.00014.0s
Inception Mercury 2100%$0.00061.4s
GPT-4.1 Mini99%$0.00043.4s
GPT-4o Mini (temp=0)100%$0.00034.8s
Mistral NeMO67%$0.00014.3s
GPT-4o Mini (temp=1)100%$0.00035.6s
Laguna S 2.196%$0.000113.7s
Arcee AI: Trinity Mini81%$0.000215.9s
Grok 4.394%$0.00093.8s
Gemini 3.1 Flash Lite (Preview)95%$0.00113.7s
Gemini 3.1 Flash Lite97%$0.00116.2s
Gemini 3.1 Flash Lite (Reasoning)98%$0.00115.3s
Laguna XS 2.194%$0.000213.4s
Nemotron 3 Nano95%$0.000210.8s
DeepSeek V4 Flash (Reasoning)90%$0.000220.8s
DeepSeek V4 Flash87%$0.000212.4s
Nemotron 3 Super98%$0.000021.7s
Mistral Small 3.2 24B71%$0.000311.0s
GPT-5.4 Mini97%$0.00202.5s
DeepSeek-V2 Chat100%$0.000116.1s
0.600.700.800.901.00

Cost vs Performance

Compares total cost for this test against the test score. Quadrant lines are drawn at the median values. Only models with available cost data are shown.

22 low-scoring outliers hidden: DeepSeek V3 (2024-12-26) (75.8%), DeepSeek V4 Pro (75.6%), Gemma 3 4B (74.6%), Cohere Command R+ (Aug. 2024) (73.2%), DeepSeek V3 (2025-03-24) (72.8%), Mistral Small 4 (Reasoning) (71.1%), Mistral Small 3.2 24B (70.5%), Mistral Large 2 (70.4%), MiniMax M2.7 (69.6%), Qwen 2.5 72B (67.9%), Mistral NeMO (66.6%), WizardLM 2 8x22b (61.1%), Ministral 3B (59.5%), Ministral 8B (52.8%), Cydonia 24B V4.1 (50.0%), Mistral Medium 3.1 (49.0%), Mistral Small 4 (48.9%), Ministral 3 8B (47.9%), Qwen3 235B A22B Instruct 2507 (46.7%), Writer: Palmyra X5 (43.2%), Ministral 3 3B (36.2%), Ministral 3 14B (10.0%).

Most Stable Models (Top 20)

Ranked by stability (median × consistency). Click a model name to view its detail page.

ScoreConsistencyStability
Muse Spark 1.1 (Reasoning, Medium)100%100%100%
Qwen3.6 Max Preview100%100%100%
MoonshotAI: Kimi K3 (Reasoning, High)100%100%100%
Muse Spark 1.1 (Reasoning, Minimal)100%100%100%
Grok 4.3 (Reasoning)100%100%100%
Claude Sonnet 4.6100%100%100%
Gemini 3.6 Flash (Reasoning, Minimal)100%100%100%
o4 Mini100%100%100%
Gemini 3.5 Flash (Reasoning, Minimal)100%100%100%
Gemini 3 Flash (Preview)100%100%100%
Gemma 4 31B100%100%100%
DeepSeek-V2 Chat100%100%100%
GPT-4o, Aug. 6th (temp=0)100%100%100%
GPT-4o Mini (temp=1)100%100%100%
GPT-4o Mini (temp=0)100%100%100%
GPT-5.6 Sol (Reasoning)100%99%99%
GPT-5.4 Mini (Reasoning, Low)100%99%99%
Z.AI GLM 5 Turbo100%98%98%
MoonshotAI: Kimi K3 (Reasoning, Low)100%98%98%
GPT-5.5 (Reasoning)99%98%98%
100%

Top Overall Models (Top 20)

Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.

ScoreCostSpeedStability
GPT-4o Mini (temp=0)100%$0.00034.8s100%
GPT-4o Mini (temp=1)100%$0.00035.6s100%
Inception Mercury 2100%$0.00061.4s96%
Gemini 3 Flash (Preview)100%$0.00205.6s100%
GPT-5.4 Mini (Reasoning, Low)100%$0.00223.5s99%
DeepSeek-V2 Chat100%$0.000116.1s100%
GPT-4.1 Mini99%$0.00043.4s93%
Gemini 3.6 Flash (Reasoning, Minimal)100%$0.00505.1s100%
GPT-4o, Aug. 6th (temp=0)100%$0.00526.1s100%
Muse Spark 1.1 (Reasoning, Minimal)100%$0.00478.0s100%
Z.AI GLM 4.5100%$0.001314.5s97%
Gemini 3.5 Flash (Reasoning, Minimal)100%$0.00735.3s100%
Z.AI GLM 5 Turbo100%$0.003714.7s98%
GPT-5.6 Luna99%$0.00525.1s95%
Hermes 3 405B99%$0.000021.0s94%
Gemma 4 31B100%$0.000332.1s100%
GPT-4o, Aug. 6th (temp=1)99%$0.00566.5s94%
GPT-5.4 Mini97%$0.00202.5s87%
Muse Spark 1.1 (Reasoning, Medium)100%$0.007312.8s100%
GPT-OSS 120B99%$0.000328.0s96%
80%90%100%
Model Total â–¼Character dialogue (Spanish) in a storyCharacter dialogue (French) in a storyCharacter dialogue (German) in a storyCharacter dialogue (Italian) in a storyCharacter dialogue (Hindi) in a story
Muse Spark 1.1 (Reasoning, Medium)100%100%100%100%100%100%
Qwen3.6 Max Preview100%100%100%100%100%100%
MoonshotAI: Kimi K3 (Reasoning, High)100%100%100%100%100%100%
Muse Spark 1.1 (Reasoning, Minimal)100%100%100%100%100%100%
Grok 4.3 (Reasoning)100%100%100%100%100%100%
Claude Sonnet 4.6100%100%100%100%100%100%
Gemini 3.6 Flash (Reasoning, Minimal)100%100%100%100%100%100%
o4 Mini100%100%100%100%100%100%
Gemini 3.5 Flash (Reasoning, Minimal)100%100%100%100%100%100%
Gemini 3 Flash (Preview)100%100%100%100%100%100%
Gemma 4 31B100%100%100%100%100%100%
DeepSeek-V2 Chat100%100%100%100%100%100%
GPT-4o, Aug. 6th (temp=0)100%100%100%100%100%100%
GPT-4o Mini (temp=1)100%100%100%100%100%100%
GPT-4o Mini (temp=0)100%100%100%100%100%100%
1–15 of 161
Page 1 / 11

Character dialogue (Spanish) in a story

Character dialogue (French) in a story

Character dialogue (German) in a story

Character dialogue (Italian) in a story

Character dialogue (Hindi) in a story