Tests

This page shows the performance of each test and its scenarios.

Bad Writing Habits

Detects common prose quality anti-patterns in AI-generated creative writing, including passive voice, past progressive overuse, weak dialogue tags, filter words, purple prose, cliches, AI-ism words/adverbs/names, and more.

Scenario Best Model Score
Detailed Writing Rules
Fantasy: entering an ancient ruin Z.AI GLM 5.3 Flash (Reasoning, Max) 93.26%
Horror: alone in an eerie place at night Z.AI GLM 5.3 Flash (Reasoning, Max) 94.73%
Literary fiction: old friends reunite Claude Opus 4 91.78%
Mystery: examining a crime scene GPT-5.4 (Reasoning, Low) 93.24%
Romance: separated couple reunites GPT-5.4 (Reasoning, Low) 92.41%
Thriller: chase through city streets Z.AI GLM 5.3 (Reasoning, Max) 94.85%
genre
Fantasy: entering an ancient ruin GPT-5.6 Sol 91.58%
Horror: alone in an eerie place at night GPT-6 Sol (Reasoning, Medium) 92.62%
Literary fiction: old friends reunite GPT-6 Luna (Reasoning, High) 91.24%
Mystery: examining a crime scene GPT-5.4 (Reasoning) 92.47%
Romance: separated couple reunites GPT-5.4 92.16%
Thriller: chase through city streets GPT-6 Sol 92.61%
Novelcrafter Default Prompt
Fantasy: entering an ancient ruin GPT-5.4 90.45%
Horror: alone in an eerie place at night Claude Opus 5.5 (Reasoning) 95.27%
Literary fiction: old friends reunite Grok 4.20 (Reasoning) 94.33%
Mystery: examining a crime scene Claude Opus 5.5 (Reasoning) 93.00%
Romance: separated couple reunites Qwen 3.7 Flash (Reasoning) 90.46%
Thriller: chase through city streets GPT-5.4 (Reasoning) 92.77%

Codex Extraction

Evaluates a model's ability to extract structured codex entries (characters, locations, objects, lore) from prose passages and return them as well-formed XML.

Scenario Best Model Score
Long: The Spire of Echoes (Dense) Claude Opus 4.8 (Reasoning) 99.21%
Medium: The Hollow (Inferred) Claude Opus 4.8 (Reasoning) 99.41%
Medium: Through the Thornveil (Scattered) Aion 3.5 (Reasoning, High) 99.12%
Short: The Rusty Lantern (Explicit) Z.AI GLM 5.1 99.49%

Tool usage within Novelcrafter

Output messages that are related to tool usage within Novelcrafter

Scenario Best Model Score
Create alternate prose sections Grok 4.5 (Reasoning, Low) 100.00%

Relationship tree

Extracts a deterministic XML family and relationship tree from cumulative literary prose.

Scenario Best Model Score
Core relationship tree GPT-5.6 Sol (Reasoning, Medium) 99.37%
Family relationship tree Gemini 3.8 Flash (Reasoning, Medium) 94.08%

Text Replacement

Tests deterministic text transformations: renaming characters/locations, expanding contractions, tense rewriting, POV shifts, gender swaps, combined transformations, and word avoidance. Scored by checking each expected change independently.

Scenario Best Model Score
Generic Prompt
Avoid said/asked/replied/answered Claude Opus 4.7 100.00%
Character rename: Elena->Mirabel, Gregor->Aldric o4 Mini High 100.00%
Combined: 3rd person past → 1st person present Qwen 3.8 Max (Reasoning, XHigh) 100.00%
Expand all contractions ByteDance Seed 2.0 Lite 100.00%
Location rename: market square, outer ring, bridge, northern mines Z.AI GLM 5.2 (Reasoning, High) 100.00%
Multi-character gender swap: Priya(F)->Rohan(M), Mara unchanged DeepSeek V4 Flash (Reasoning) 100.00%
Passive voice → active voice Claude Opus 4.8 (Reasoning) 98.46%
POV shift: 3rd person to 1st person (Elena's perspective) Hy4 Preview (Reasoning, High) 100.00%
Tense rewriting: past to present Hy4 Preview (Reasoning, High) 99.91%
Specific Prompt
Avoid said/asked/replied/answered Muse Spark 1.1 (Reasoning, Medium) 100.00%
Character rename: Elena->Mirabel, Gregor->Aldric Z.AI GLM 4.7 100.00%
Combined: 3rd person past → 1st person present Claude Opus 4.6 (Reasoning) 100.00%
Expand all contractions Mistral Large 3 100.00%
Location rename: market square, outer ring, bridge, northern mines GPT-OSS 120B 100.00%
Multi-character gender swap: Priya(F)->Rohan(M), Mara unchanged Ministral 3 8B 100.00%
Passive voice → active voice Qwen 3.8 Max (Reasoning, XHigh) 99.56%
POV shift: 3rd person to 1st person (Elena's perspective) Gemini 3.8 Flash (Reasoning, Medium) 100.00%
Tense rewriting: past to present DeepSeek V4 Pro 100.00%