NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M37Q2WVWEN65MGSRCD68SNW4
Run Details
Model:
Claude Opus 5.5 (Reasoning)
Test:
Dialogue tags
Scenario:
Write 200 words with 50% dialogue
Duration:
18.6s
Cost:
$0.035040
Provider:
Claude Platform on AWS
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
99.91%
Dialogue to Total Word Ratio
(details)
99.9558%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 50% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly unrolled the blueprint across the hood of the van. It was drawn in purple crayon. "Where did you get this?" Bill asked. "The birthday boy drew it for his invitations. Very detailed. He even labeled the snack table." Bill squinted. "Is that a moat?" "Paddling pool. Treat it like a moat." A dozen balloons bobbed over the fence next door, each printed with a grinning dinosaur. Bill watched them with the wary respect of a man who had once been bitten by a clown. "Walk me through it," he said. "Two o'clock, the magician starts. Every child and parent faces the patio. You go over the fence, I grab the gift pile, we're gone before the rabbit comes out of the hat." "And if someone sees us?" Molly handed him a folded bundle of felt. "Then you're Rexy, the party dinosaur." Bill held up the costume. The tail dragged on the asphalt. "Molly, this is a crime against dignity." "We're stealing birthday presents from a six-year-old, Bill. Dignity left hours ago." He sighed, stepped into the legs, and zipped himself up to the snout. Somewhere beyond the fence, a small voice shrieked, "DINOSAUR!" Bill froze solid. "It's only one-thirty."
created_at
"2026-09-23T18:05:14.017Z"
9/23/2026, 6:05:14 PM
model
"anthropic/claude-opus-5.5-20260921"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
6875
moderation_latency
(null)
generation_time
18550
tokens_prompt
75
tokens_completion
495
native_tokens_prompt
95
native_tokens_completion
1733
native_tokens_completion_images
(null)
native_tokens_reasoning
1287
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"end_turn"
service_tier
"default"
usage
0.03504
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790186714-AIkHeB48UVJx4H17CtIX"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790186714-j0WglOJTq8IGC2SVUAWb"
upstream_id
"msg_011CfLpAf7rvvhTjaiKGtXRj"
provider_responses
0
endpoint_id
"3a2388bc-3740-4e64-a1f5-4e301726b6b9"
id
"msg_011CfLpAf7rvvhTjaiKGtXRj"
is_byok
false
latency
1341
model_permaslug
"anthropic/claude-opus-5.5-20260921"
provider_name
"Claude Platform on AWS"
status
200
total_cost
0.03504
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Claude Platform on AWS"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
99.91%
Dialogue to Total Word Ratio
Ratio: 51.72%, Deviation: 1.72%
neededClean
false
wordsTotal
203
wordsDialogue
105
99.9558%