NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-07-25-01KYCRE7JK0SX70CS00NDY5AX5
Run Details
Model:
Claude Opus 5 (Reasoning)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
11.1s
Cost:
$0.014340
Provider:
Amazon Bedrock
Evaluation summary
Score
Evaluator
Details Link
93.65%
Matches word count
(details)
98.77%
Dialogue to Total Word Ratio
(details)
96.2113%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
The blueprints were drawn on the back of a takeout menu, which Molly had flattened against the dashboard with the heel of her palm. Bill studied them the way a man studies a legal document he intends to ignore. Circles indicated the bouncy castle. An X marked the folding table where the presents would be stacked, glossy and beribboned, unguarded but for a gauntlet of grandmothers. Molly tapped the menu twice. "Piñata at four," she said. "That's our window." Bill nodded, though he was, in fact, thinking about cake. He had been thinking about cake since Tuesday, when the reconnaissance mission—conducted through a hedge with binoculars and a thermos of soup—had revealed a three-tiered dinosaur confection of unusual ambition. A professional does not get distracted. A professional also does not skip lunch. They had costumes. Molly's was a clown suit acquired under circumstances neither of them discussed anymore. Bill's was a dinosaur, which he insisted was thematically appropriate and Molly insisted was two sizes too small, a debate that had consumed most of their planning meetings and all of their goodwill. Outside, balloons bobbed against a mailbox like sentries who had given up. Bill zipped himself in. "Remember," Molly whispered, "no eye contact with the birthday boy."
created_at
"2026-07-25T13:46:37.018Z"
7/25/2026, 1:46:37 PM
model
"anthropic/claude-opus-5-20260723"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
3742
moderation_latency
(null)
generation_time
10988
tokens_prompt
75
tokens_completion
381
native_tokens_prompt
93
native_tokens_completion
555
native_tokens_completion_images
(null)
native_tokens_reasoning
51
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"end_turn"
service_tier
"default"
usage
0.01434
router
(null)
provider_responses
0
endpoint_id
"76cb4608-f48c-483d-8da8-9957fb44244e"
id
"msg_011CdNsuu4cqKCmyf15bAg7E"
is_byok
false
latency
1371
model_permaslug
"anthropic/claude-opus-5-20260723"
provider_name
"Amazon Bedrock"
status
200
user_agent
"langchainjs-openai/1.0.0 ((node/v24.18.0; linux; x64))"
http_referer
(null)
request_id
"req-1784987197-4UBcGo35rVIXskb1PGPJ"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1784987197-MCn0OPBY5OPTk26dbnrb"
upstream_id
"msg_011CdNsuu4cqKCmyf15bAg7E"
total_cost
0.01434
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Amazon Bedrock"
response_cache_source_id
(null)
data_region
"global"
Evaluation details
Result
Evaluator
Details
Meta Data
93.65%
Matches word count
n/a
neededClean
false
words
209
98.77%
Dialogue to Total Word Ratio
Ratio: 6.67%, Deviation: 3.33%
neededClean
false
wordsTotal
210
wordsDialogue
14
96.2113%