NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-29-01M3P2CHPTFHEF92ZN23KTF8MW
Run Details
Model:
Claude Sonnet 5.5 (Reasoning)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
9.3s
Cost:
$0.009960
Provider:
Claude Platform on AWS
Evaluation summary
Score
Evaluator
Details Link
2.01%
Matches word count
(details)
36.79%
Dialogue to Total Word Ratio
(details)
19.3998%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
"Okay, Bill, repeat the plan." "We sneak into Tommy Peterson's sixth birthday party, steal every present, and leave before cake." "Before cake? Why before cake?" "Because you get sentimental about cake, Molly." "That was one time, and it was a really good cake." "You cried into the frosting." "It said 'Happy Birthday Gary.' Nobody remembers Gary." "Focus. I'm the clown. You're the bouncy castle inspector." "Why am I the inspector?" "Nobody questions a woman with a clipboard." "And what does the clown do?" "Distracts the children with balloon animals." "You can't make balloon animals." "I can make one. It's a pointy dog." "Bill, that's a sword." "Children love swords." "Their parents don't. What's the escape route?" "The piƱata van. Loaded, gone, done in ninety seconds." "And if a kid spots us?" "We give him a balloon sword and a wink." "Bill, they're six. He'll scream." "Then we scream back. Louder. Professionals always win." "Fine. But if there's cake, I'm taking a slice." "Molly." "A small slice. For Gary." "...Fine. But we take it to go."
created_at
"2026-09-29T07:52:06.629Z"
9/29/2026, 7:52:06 AM
model
"anthropic/claude-sonnet-5.5-20260928"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
3786
moderation_latency
(null)
generation_time
9281
tokens_prompt
75
tokens_completion
459
native_tokens_prompt
95
native_tokens_completion
977
native_tokens_completion_images
(null)
native_tokens_reasoning
545
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"end_turn"
service_tier
"default"
usage
0.00996
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.21.0; linux; x64))"
http_referer
(null)
request_id
"req-1790668326-TQvVUJDUsdBNEGiaf3fJ"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790668326-ohj10Rv72mK0O4ayQBLf"
upstream_id
"msg_011CfXNGPWvx9CadLH8ZSD5E"
provider_responses
0
endpoint_id
"99aaad94-923b-4fc1-b763-271ed5486f7a"
id
"msg_011CfXNGPWvx9CadLH8ZSD5E"
is_byok
false
latency
691
model_permaslug
"anthropic/claude-sonnet-5.5-20260928"
provider_name
"Claude Platform on AWS"
status
200
total_cost
0.00996
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Claude Platform on AWS"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
2.01%
Matches word count
n/a
neededClean
false
words
175
36.79%
Dialogue to Total Word Ratio
Ratio: 100.00%, Deviation: 10.00%
neededClean
false
wordsTotal
175
wordsDialogue
175
19.3998%