NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-29-01M3P2GZC3MB88SZNZ9FF0XKW7
Run Details
Model:
Claude Sonnet 5.5 (Reasoning)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
9.0s
Cost:
$0.011280
Provider:
Claude Platform on AWS
Evaluation summary
Score
Evaluator
Details Link
95.99%
Matches word count
(details)
52.52%
Dialogue to Total Word Ratio
(details)
74.2519%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
"Okay, Bill, repeat the plan." "We sneak into Tyler's fifth birthday party, steal every present, and leave before cake." "Before cake? Why before cake?" "Because you get emotional at cake, Molly." "That was one time, and it was a dinosaur cake. Nobody should have to watch that get cut." "Focus. How do we get in?" "Costumes. You're the clown." "I hate clowns." "Everyone hates clowns. That's why nobody will look at you directly." "And you?" "I'm the bouncy castle inspector. I brought a clipboard." "That's brilliant. Terrifying, but brilliant." "Thank you. Now, the presents are in the living room, guarded by Grandma Peg." "How do we handle Grandma Peg?" "You fold balloon animals until she's mesmerized." "I only know one animal." "Which?" "A sad giraffe." "Perfect. Sad giraffes are universally distracting." "What if the kids catch us?" "Toddlers are unreliable witnesses. Last time one described me as 'a tall dog.'" "You were wearing a fur hood." "Exactly. Details matter. Ready?" "Wait. If we steal everything, what do we do with forty Lego sets?" "Bill, we're professionals. We build a getaway vehicle." Bill sighed. "Can we at least keep the cake?" "No cake."
created_at
"2026-09-29T07:54:31.688Z"
9/29/2026, 7:54:31 AM
model
"anthropic/claude-sonnet-5.5-20260928"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
3967
moderation_latency
(null)
generation_time
8971
tokens_prompt
75
tokens_completion
375
native_tokens_prompt
95
native_tokens_completion
1109
native_tokens_completion_images
(null)
native_tokens_reasoning
618
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"end_turn"
service_tier
"default"
usage
0.01128
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.21.0; linux; x64))"
http_referer
(null)
request_id
"req-1790668471-KS6NDBLWlxrTShlLg6LJ"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790668471-lw0wHYniOoYBj3XhU8r6"
upstream_id
"msg_011CfXNT5YWAcnGBoGdc9g5Z"
provider_responses
0
endpoint_id
"99aaad94-923b-4fc1-b763-271ed5486f7a"
id
"msg_011CfXNT5YWAcnGBoGdc9g5Z"
is_byok
false
latency
705
model_permaslug
"anthropic/claude-sonnet-5.5-20260928"
provider_name
"Claude Platform on AWS"
status
200
total_cost
0.01128
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Claude Platform on AWS"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
95.99%
Matches word count
n/a
neededClean
false
words
192
52.52%
Dialogue to Total Word Ratio
Ratio: 98.96%, Deviation: 8.96%
neededClean
false
wordsTotal
192
wordsDialogue
190
74.2519%