NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M37PJTX98DVRZ2SJ7S9EGM7W
Run Details
Model:
Claude Opus 5.5 (Reasoning)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
23.6s
Cost:
$0.041160
Provider:
Claude Platform on AWS
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
99.9989%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly unrolled the blueprints across the hood of the van, pinning one corner with a thermos and the other with Bill's elbow. The target was 14 Pemberton Close, a modest semi-detached currently wrapped in more balloons than a hot-air festival. Inside, according to their intel, lay forty-two gift-wrapped boxes, a bouncy castle, and one Tobias Finch, who was turning seven at three o'clock sharp. "Entry point's the cat flap, obviously," Molly said. Bill looked down at his own considerable stomach, then back at her, deeply wounded. She ignored him and tapped the diagram. The plan was elegant. Bill would pose as the hired magician, a role he had rehearsed for nine days, mostly by pulling the same exhausted rabbit from the same hat. While thirty children watched him fail spectacularly, Molly would slip through the kitchen, load the presents into a laundry trolley, and wheel them out past the parents, who would be too busy holding paper plates of cocktail sausages to notice anything. "What if the kids recognise me?" Bill asked. Molly paused. Bill had, after all, been arrested at a bowling alley birthday party last spring, still holding cake. "Wear the fake beard, the big one," she said.
created_at
"2026-09-23T17:56:27.707Z"
9/23/2026, 5:56:27 PM
model
"anthropic/claude-opus-5.5-20260921"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
6024
moderation_latency
(null)
generation_time
23539
tokens_prompt
75
tokens_completion
600
native_tokens_prompt
95
native_tokens_completion
2039
native_tokens_completion_images
(null)
native_tokens_reasoning
1610
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"end_turn"
service_tier
"default"
usage
0.04116
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790186187-z5yCOHc3o4QykLjhE8IL"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790186187-9cLFyjLYNIGhc78Qy9ib"
upstream_id
"msg_011CfLoVrm1UoA9DJ6EVGNEg"
provider_responses
0
endpoint_id
"3a2388bc-3740-4e64-a1f5-4e301726b6b9"
id
"msg_011CfLoVrm1UoA9DJ6EVGNEg"
is_byok
false
latency
1917
model_permaslug
"anthropic/claude-opus-5.5-20260921"
provider_name
"Claude Platform on AWS"
status
200
total_cost
0.04116
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Claude Platform on AWS"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
100.00%
Dialogue to Total Word Ratio
Ratio: 9.31%, Deviation: 0.69%
neededClean
false
wordsTotal
204
wordsDialogue
19
99.9989%