NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M37NVTMJP4EAH8WH9PB60DD9
Run Details
Model:
Claude Opus 5.5 (Reasoning)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
20.5s
Cost:
$0.038700
Provider:
Claude Platform on AWS
Evaluation summary
Score
Evaluator
Details Link
99.74%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
99.8721%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
"Right," said Molly, unrolling the blueprint. "Timmy Pemberton's seventh birthday. Two o'clock. Forty guests, sixty presents minimum." "Sixty?" Bill whistled. "What's the kid done to deserve sixty?" "Nothing. His parents feel guilty about the divorce. That's our window." "And how do we get in?" "You're the entertainment." She handed him a folder. "Bongo the Clown." "I hate clowns." "Children hate clowns too, Bill. That's the beauty. Nobody will look at you directly." "What about the bouncy castle?" "Decoy. At two-fifteen, I puncture it. Mass hysteria. Parents scramble. You grab the gift table and walk out the back gate." "Sixty presents? I've got two arms, Molly." "You've got balloon animals. Tie them together, make a sort of... raft." "A raft. Molly, this is a garden, not the Atlantic." "A floating raft of stolen joy." Bill stared. "What if the kid cries?" "Timmy won't cry. Timmy bit a magician last year. That's why they need a clown." "He bit a magician?" "Drew blood. The rabbit's still in therapy. Went in the hat, came out with a limp. Nobody talks about it." Bill donned the red nose. "I want hazard pay." "You'll get cake." "What kind of cake?" "Stolen."
created_at
"2026-09-23T17:43:53.763Z"
9/23/2026, 5:43:53 PM
model
"anthropic/claude-opus-5.5-20260921"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
5385
moderation_latency
(null)
generation_time
20478
tokens_prompt
75
tokens_completion
573
native_tokens_prompt
95
native_tokens_completion
1916
native_tokens_completion_images
(null)
native_tokens_reasoning
1421
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"end_turn"
service_tier
"default"
usage
0.0387
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790185433-p8JIz1ML8WrbLfzBWSyo"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790185433-LdM6VVxfa249NVDCQct6"
upstream_id
"msg_011CfLnYHBt4DeFkPZVDss58"
provider_responses
0
endpoint_id
"3a2388bc-3740-4e64-a1f5-4e301726b6b9"
id
"msg_011CfLnYHBt4DeFkPZVDss58"
is_byok
false
latency
1225
model_permaslug
"anthropic/claude-opus-5.5-20260921"
provider_name
"Claude Platform on AWS"
status
200
total_cost
0.0387
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Claude Platform on AWS"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
99.74%
Matches word count
n/a
neededClean
false
words
196
100.00%
Dialogue to Total Word Ratio
Ratio: 90.36%, Deviation: 0.36%
neededClean
false
wordsTotal
197
wordsDialogue
178
99.8721%