Codex Extraction

Evaluates a model's ability to extract structured codex entries (characters, locations, objects, lore) from prose passages and return them as well-formed XML.

Price-Performance Score Distribution (Top 20)

Click a model name to view its detail page.

ScoreCostTime
Gemini 3 Flash (Preview)97%$0.00273.9s
DeepSeek V4 Flash95%$0.00037.7s
Qwen 3.5 Plus (2026-02-15)98%$0.003010.6s
Mistral Medium 3.196%$0.00265.8s
Gemini 3.6 Flash (Reasoning, Minimal)98%$0.0102.9s
Xiaomi MIMO v2.597%$0.003413.4s
Gemini 3.1 Flash Lite (Reasoning)95%$0.00182.0s
Mistral Large 394%$0.00278.2s
Gemini 2.5 Flash94%$0.00232.5s
Gemini 3.5 Flash (Reasoning, Minimal)98%$0.0103.1s
Gemini 3.5 Flash Lite (Reasoning, Minimal)93%$0.00251.5s
Gemini 3.1 Flash Lite (Preview)94%$0.00172.0s
Z.AI GLM 5 Turbo97%$0.006816.0s
Z.AI GLM 4.596%$0.002816.8s
DeepSeek V4 Pro95%$0.002116.0s
Z.AI GLM 5.2 (Reasoning, High)98%$0.007121.2s
Ministral 3 8B94%$0.00063.3s
Gemini 3.1 Flash Lite94%$0.00175.1s
Xiaomi MIMO v2.5 Pro96%$0.004818.7s
Laguna S 2.195%$0.000524.7s
0.901.00

Cost vs Performance

Compares total cost for this test against the test score. Quadrant lines are drawn at the median values. Only models with available cost data are shown.

11 low-scoring outliers hidden: GPT-4o Mini (temp=1) (84.3%), Cohere Command R+ (Aug. 2024) (84.2%), Llama 3.1 70B (83.8%), GPT-5.4 Nano (Reasoning, Low) (83.8%), Gemma 3 4B (83.3%), GPT-5.4 Nano (83.2%), Gemma 4 26B (80.5%), Gemma 3 12B (77.7%), Aion 3.0 Mini (76.1%), GPT-4.1 Nano (75.3%), Mistral NeMO (26.1%).

Top Overall Models (Top 20)

Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.

ScoreCostSpeedStability
Gemini 3 Flash (Preview)97%$0.00273.9s95%
Qwen 3.5 Plus (2026-02-15)98%$0.003010.6s95%
Gemini 3.1 Flash Lite (Reasoning)95%$0.00182.0s92%
Gemini 3.6 Flash (Reasoning, Minimal)98%$0.0102.9s96%
Gemini 3.1 Flash Lite (Preview)94%$0.00172.0s92%
Mistral Medium 3.196%$0.00265.8s92%
Gemini 3.5 Flash (Reasoning, Minimal)98%$0.0103.1s95%
Xiaomi MIMO v2.597%$0.003413.4s94%
Gemini 3.1 Flash Lite94%$0.00175.1s92%
GPT-5.6 Luna95%$0.00433.6s92%
Muse Spark 1.1 (Reasoning, Minimal)97%$0.00897.1s94%
Z.AI GLM 5.2 (Reasoning, High)98%$0.007121.2s97%
Ministral 3 8B94%$0.00063.3s88%
Z.AI GLM 5 Turbo97%$0.006816.0s95%
Gemini 2.5 Flash94%$0.00232.5s89%
Grok 4.2095%$0.00484.9s91%
Mistral Large 394%$0.00278.2s91%
DeepSeek V4 Pro95%$0.002116.0s92%
Z.AI GLM 4.596%$0.002816.8s92%
Mistral Small 3.2 24B93%$0.00054.4s89%
80%90%100%
Model Total â–¼Short: The Rusty Lantern (Explicit)Medium: Through the Thornveil (Scattered)Medium: The Hollow (Inferred)Long: The Spire of Echoes (Dense)
Claude Opus 4.599%99%98%99%98%
Claude Opus 4.6 (Reasoning)98%99%99%99%98%
Claude Opus 4.8 (Reasoning)98%97%98%99%99%
Claude Opus 4.8 (Reasoning, Low)98%97%99%99%99%
Z.AI GLM 5.2 (Reasoning, High)98%99%99%97%98%
Gemini 3.5 Flash (Reasoning)98%98%99%98%99%
Claude Opus 4.698%99%97%99%98%
Grok 4.5 (Reasoning, High)98%98%99%98%98%
Z.AI GLM 5.198%99%99%97%97%
GPT-598%99%98%99%96%
MoonshotAI: Kimi K3 (Reasoning, Low)98%98%98%98%97%
MoonshotAI: Kimi K3 (Reasoning, High)98%98%97%98%99%
Gemini 3 Flash (Preview, Reasoning)98%98%97%98%99%
Gemini 3.6 Flash (Reasoning)98%97%98%98%98%
Gemini 3.5 Flash (Reasoning, Minimal)98%98%99%97%97%
1–15 of 161
Page 1 / 11

Short: The Rusty Lantern (Explicit)

Medium: Through the Thornveil (Scattered)

Medium: The Hollow (Inferred)

Long: The Spire of Echoes (Dense)