NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-07-25-01KYCRKCXA96ED0EM9F3CHA6MT
Run Details
Model:
Claude Opus 5 (Reasoning)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
10.2s
Cost:
$0.014040
Provider:
Amazon Bedrock
Evaluation summary
Score
Evaluator
Details Link
86.38%
Matches word count
(details)
86.19%
Dialogue to Total Word Ratio
(details)
86.2868%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
"Bouncy castle's the weak point," Bill said. "It's a children's party, Bill. There is no weak point. It's all weak point." "So we go in through the bouncy castle." "Why the bouncy castle?" "Nobody suspects a man emerging from a bouncy castle." "Because nobody's ever done it. Because it's a sealed inflatable object." "Then I'll be the first, Molly. Pioneers get remembered." "Pioneers get arrested with their trousers full of party bags." "The presents are stacked on the trestle table by the window. Forty, maybe fifty. Some of them are big." "How big?" "Trampoline big." "We're not stealing a trampoline." "I'm just saying, the intelligence suggests trampoline." "Who gave you the intelligence?" "My nephew. He's eight. He's got no reason to lie." "He's got every reason to lie, he's eight." "He wants a cut." "Of what?" "The presents." "So he's stealing his own birthday presents." "It's his cousin's party. There's bad blood." Molly rubbed her eyes. "Fine. Entry?" "Front door. I'll be a clown." "You can't do balloons." "I can do a sausage." "A clown who can only do sausages is a suspicious clown." "Then you distract them." "With what?" "Cake." "Whose cake?" "Theirs. We steal the cake first, then give it back. They'll love us." Molly stood. "Get the wig."
created_at
"2026-07-25T13:49:26.324Z"
7/25/2026, 1:49:26 PM
model
"anthropic/claude-opus-5-20260723"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
2349
moderation_latency
(null)
generation_time
10100
tokens_prompt
75
tokens_completion
359
native_tokens_prompt
93
native_tokens_completion
543
native_tokens_completion_images
(null)
native_tokens_reasoning
19
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"end_turn"
service_tier
"default"
usage
0.01404
router
(null)
provider_responses
0
endpoint_id
"76cb4608-f48c-483d-8da8-9957fb44244e"
id
"msg_011CdNt8NWTKavWyHecsJuZB"
is_byok
false
latency
1199
model_permaslug
"anthropic/claude-opus-5-20260723"
provider_name
"Amazon Bedrock"
status
200
user_agent
"langchainjs-openai/1.0.0 ((node/v24.18.0; linux; x64))"
http_referer
(null)
request_id
"req-1784987366-bwEFTUOZkxjtHTskWxhD"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1784987366-ldmbwnVpeq2aHMJcr1bX"
upstream_id
"msg_011CdNt8NWTKavWyHecsJuZB"
total_cost
0.01404
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Amazon Bedrock"
response_cache_source_id
(null)
data_region
"global"
Evaluation details
Result
Evaluator
Details
Meta Data
86.38%
Matches word count
n/a
neededClean
false
words
211
86.19%
Dialogue to Total Word Ratio
Ratio: 96.21%, Deviation: 6.21%
neededClean
false
wordsTotal
211
wordsDialogue
203
86.2868%