NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-29-01M3P2E0Q1MWM8CH4WA9BHN54T
Run Details
Model:
Claude Sonnet 5.5 (Reasoning)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
18.9s
Cost:
$0.024900
Provider:
Claude Platform on AWS
Evaluation summary
Score
Evaluator
Details Link
99.92%
Matches word count
(details)
12.57%
Dialogue to Total Word Ratio
(details)
56.2461%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly crouched behind a hedge outside the Hendersons' suburban colonial, binoculars pressed to her eyes, while Bill wheezed beside her in a foam dinosaur costume that was two sizes too small. Through the bay window, forty screaming seven-year-olds circled a bouncy castle, a cake shaped like a unicorn, and a mountain of gift-wrapped boxes stacked taller than the birthday girl's father. "That's the biggest pile of loot I've ever seen," Bill whispered, tail knocking over a garden gnome. Molly lowered the binoculars and studied the clown by the pizza table. The clown studied her back, then slowly drew a finger across his throat. Everyone in the trade knew Pickles. He had retired from safecracking after a bad divorce and now worked birthdays for cash and revenge. "Pickles is inside," she said. "Abort the front door. Switch to Plan B." Bill adjusted his foam snout, which had slid over one eye. He tried to answer, but a passing toddler pointed at him and shrieked with delight, and Molly hissed at him to stay still. Bill roared weakly. The toddler applauded. Molly grinned. Thanks to her, the piƱata was already full of smoke bombs, and nobody suspected a thing.
created_at
"2026-09-29T07:52:54.758Z"
9/29/2026, 7:52:54 AM
model
"anthropic/claude-sonnet-5.5-20260928"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
6141
moderation_latency
(null)
generation_time
18884
tokens_prompt
75
tokens_completion
586
native_tokens_prompt
95
native_tokens_completion
2471
native_tokens_completion_images
(null)
native_tokens_reasoning
2037
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"end_turn"
service_tier
"default"
usage
0.0249
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.21.0; linux; x64))"
http_referer
(null)
request_id
"req-1790668374-OfingGc6JFA5STibzyy4"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790668374-w8gEOKwQROgsOJBRJ0Ef"
upstream_id
"msg_011CfXNKwCGFuNX1GRoHPzFY"
provider_responses
0
endpoint_id
"99aaad94-923b-4fc1-b763-271ed5486f7a"
id
"msg_011CfXNKwCGFuNX1GRoHPzFY"
is_byok
false
latency
675
model_permaslug
"anthropic/claude-sonnet-5.5-20260928"
provider_name
"Claude Platform on AWS"
status
200
total_cost
0.0249
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Claude Platform on AWS"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
99.92%
Matches word count
n/a
neededClean
false
words
197
12.57%
Dialogue to Total Word Ratio
Ratio: 22.00%, Deviation: 12.00%
neededClean
false
wordsTotal
200
wordsDialogue
44
56.2461%