NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-29-01M3P2CMCX9WQWF8X1DZ7ETNBR
Run Details
Model:
Claude Sonnet 5.5 (Reasoning)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
8.3s
Cost:
$0.010020
Provider:
Claude Platform on AWS
Evaluation summary
Score
Evaluator
Details Link
0.00%
Matches word count
(details)
99.92%
Dialogue to Total Word Ratio
(details)
49.9614%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly slapped the blueprint onto the picnic table. "Bill, tell me the plan again." "We enter as clowns, grab every present, and leave before cake." "Why before cake?" "Because last time you cried at the cake, Molly." "It had a tiny sugar dinosaur, Bill. He looked terrified." "Focus. Seven-year-old Timmy has forty gifts. Forty." "Forty? What kind of monster invites forty friends?" "The kind with a bouncy castle." "Fine. Who distracts the parents?" "You do. With balloon animals." "I only know one animal." "Which?" "A sad, lumpy snake." "Perfect. Parents love sad snakes." "And what will you do?" "I'll distract the children with magic." "You can't do magic." "Exactly. They'll be so confused they'll forget presents exist." "Bill, that's almost smart." "Thank you. Now, what's our code word if it goes wrong?" "Piñata." "Why piñata?" "Because if we panic, I want something to hit." "Beautiful." Bill grabbed his oversized shoes. "Let's go steal childhood."
created_at
"2026-09-29T07:52:09.394Z"
9/29/2026, 7:52:09 AM
model
"anthropic/claude-sonnet-5.5-20260928"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
4164
moderation_latency
(null)
generation_time
8280
tokens_prompt
75
tokens_completion
302
native_tokens_prompt
95
native_tokens_completion
983
native_tokens_completion_images
(null)
native_tokens_reasoning
579
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"end_turn"
service_tier
"default"
usage
0.01002
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.21.0; linux; x64))"
http_referer
(null)
request_id
"req-1790668329-PY64zgWZHbb2pjbda4vC"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790668329-RjWACU7ehemlIA4gWagB"
upstream_id
"msg_011CfXNGbfRsWHHXPzmib5gF"
provider_responses
0
endpoint_id
"99aaad94-923b-4fc1-b763-271ed5486f7a"
id
"msg_011CfXNGbfRsWHHXPzmib5gF"
is_byok
false
latency
794
model_permaslug
"anthropic/claude-sonnet-5.5-20260928"
provider_name
"Claude Platform on AWS"
status
200
total_cost
0.01002
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Claude Platform on AWS"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
0.00%
Matches word count
n/a
neededClean
false
words
154
99.92%
Dialogue to Total Word Ratio
Ratio: 91.67%, Deviation: 1.67%
neededClean
false
wordsTotal
156
wordsDialogue
143
49.9614%