Codex Red Herring (False Positive Detection)

Tests whether models correctly report "no violations" when a codex is fully consistent with the prose passage. Models that hallucinate false violations (false positives) fail. Uses a 2×2 matrix of text length × codex size, with bare and detailed-entry variants.

Price-Performance Score Distribution (Top 20)

Click a model name to view its detail page.

ScoreCostTime
GPT-5.6 Luna70%$0.00201.2s
GPT-4.1 Nano63%$0.00054.6s
GPT-5.4 Nano68%$0.00051.5s
Ministral 8B68%$0.00074.4s
Ministral 3 8B78%$0.001317.9s
GPT-5.4 Nano (Reasoning, Low)97%$0.00083.9s
Inception Mercury 296%$0.00273.8s
GPT-5.4 Mini (Reasoning, Low)95%$0.00344.0s
GPT-5.4 Nano (Reasoning)94%$0.001711.4s
Gemma 4 31B71%$0.000911.7s
ByteDance Seed 1.6 Flash94%$0.00089.1s
Arcee AI: Trinity Mini73%$0.000926.3s
GPT-5.6 Luna (Reasoning)92%$0.00484.7s
GPT-4.186%$0.00481.2s
Gemini 3.6 Flash (Reasoning, Minimal)79%$0.00781.6s
Laguna XS 2.191%$0.000516.4s
GPT-5.6 Sol75%$0.00911.7s
Laguna S 2.192%$0.000841.9s
GPT-5.4 Mini (Reasoning)89%$0.009010.8s
Gemini 2.5 Flash Lite (Reasoning)92%$0.002316.6s
0.600.700.800.901.00

Cost vs Performance

Compares total cost for this test against the test score. Quadrant lines are drawn at the median values. Only models with available cost data are shown.

Most Stable Models (Top 20)

Ranked by stability (median × consistency). Click a model name to view its detail page.

ScoreConsistencyStability
Nemotron 3 Super99%83%83%
o4 Mini High97%72%72%
Muse Spark 1.1 (Reasoning, Minimal)97%72%72%
GPT-5.4 Nano (Reasoning, Low)97%70%70%
o4 Mini96%67%67%
Inception Mercury 296%67%67%
Z.AI GLM 5 Turbo96%65%65%
GPT-5.195%64%64%
GPT-5.4 Mini (Reasoning, Low)95%64%64%
Muse Spark 1.1 (Reasoning, Medium)95%61%61%
Claude Opus 4.6 (Reasoning)94%60%60%
Grok 4.5 (Reasoning, High)94%60%60%
ByteDance Seed 1.6 Flash94%59%59%
Grok 4.5 (Reasoning, Low)94%58%58%
Z.AI GLM 5.2 (Reasoning, High)94%58%58%
GPT-5.4 Nano (Reasoning)94%57%57%
Laguna S 2.192%56%56%
Z.AI GLM 593%56%56%
GPT-5 Mini93%56%56%
GPT-5 Nano93%55%55%
10%20%30%40%50%60%70%80%90%100%

Top Overall Models (Top 20)

Ranked by composite score (performance, cost, speed & stability). Click a model name to view its detail page.

ScoreCostSpeedStability
Nemotron 3 Super99%$0.00001.3m83%
GPT-5.4 Nano (Reasoning, Low)97%$0.00083.9s70%
Inception Mercury 296%$0.00273.8s67%
Muse Spark 1.1 (Reasoning, Minimal)97%$0.01211.2s72%
GPT-5.4 Mini (Reasoning, Low)95%$0.00344.0s64%
Z.AI GLM 5 Turbo96%$0.007116.0s65%
o4 Mini96%$0.01425.0s67%
ByteDance Seed 1.6 Flash94%$0.00089.1s59%
GPT-5.4 Nano (Reasoning)94%$0.001711.4s57%
o4 Mini High97%$0.02752.5s72%
Grok 4.5 (Reasoning, Low)94%$0.01217.0s58%
Muse Spark 1.1 (Reasoning, Medium)95%$0.01919.4s61%
Z.AI GLM 5.2 (Reasoning, High)94%$0.008125.8s58%
GPT-5.195%$0.02526.1s64%
GPT-5.6 Luna (Reasoning)92%$0.00484.7s51%
Gemini 2.5 Flash Lite (Reasoning)92%$0.002316.6s53%
Laguna XS 2.191%$0.000516.4s52%
Laguna S 2.192%$0.000841.9s56%
GPT-5 Mini93%$0.005937.8s56%
GPT-5 Nano93%$0.00351.1m55%
10%20%30%40%50%60%70%80%90%100%
basic entriesdetailed entries
Model Total ▼Short text (~524 words), small codex (11 entries)Short text (~524 words), big codex (51 entries)Long text (~1594 words), small codex (11 entries)Long text (~1594 words), big codex (51 entries)Short text (~524 words), small codex (11 detailed entries)Short text (~524 words), big codex (51 detailed entries)Long text (~1594 words), small codex (11 detailed entries)Long text (~1594 words), big codex (51 detailed entries)
Nemotron 3 Super99%100%100%100%93%100%100%100%100%
Muse Spark 1.1 (Reasoning, Minimal)97%93%85%100%100%100%100%100%100%
o4 Mini High97%93%100%100%85%100%100%100%100%
GPT-5.4 Nano (Reasoning, Low)97%100%90%100%100%93%93%100%100%
o4 Mini96%100%100%100%85%93%100%93%100%
Inception Mercury 296%100%100%100%100%85%100%100%85%
Z.AI GLM 5 Turbo96%100%100%100%68%100%100%100%100%
GPT-5.195%93%93%100%100%93%85%100%100%
GPT-5.4 Mini (Reasoning, Low)95%93%85%100%85%100%100%100%100%
Muse Spark 1.1 (Reasoning, Medium)95%85%100%100%83%100%100%100%92%
Claude Opus 4.6 (Reasoning)94%100%100%100%100%85%70%100%100%
Grok 4.5 (Reasoning, High)94%100%100%100%100%70%85%100%100%
ByteDance Seed 1.6 Flash94%100%100%100%93%84%93%100%84%
Grok 4.5 (Reasoning, Low)94%85%100%100%75%93%100%100%100%
Z.AI GLM 5.2 (Reasoning, High)94%93%100%100%59%100%100%100%100%
1–15 of 161
Page 1 / 11

basic entries

Short text (~524 words), small codex (11 entries)

Short text (~524 words), big codex (51 entries)

Long text (~1594 words), small codex (11 entries)

Long text (~1594 words), big codex (51 entries)

detailed entries

Short text (~524 words), small codex (11 detailed entries)

Short text (~524 words), big codex (51 detailed entries)

Long text (~1594 words), small codex (11 detailed entries)

Long text (~1594 words), big codex (51 detailed entries)