NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-07-25-01KYCQQ3YB2HBMMVPCMQE8037H
Run Details
Model:
Claude Opus 5 (Reasoning)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
26.2s
Cost:
$0.013590
Provider:
Amazon Bedrock
Evaluation summary
Score
Evaluator
Details Link
93.65%
Matches word count
(details)
99.70%
Dialogue to Total Word Ratio
(details)
96.6740%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly had drawn the layout of the Pemberton back garden on the back of a napkin, using ketchup for the bouncy castle and a mustard smear to indicate the trestle table where the presents would be stacked like a glittering paper ziggurat. Bill studied it with the gravity of a man reading his own autopsy report, then dipped a chip in the bouncy castle. "That's structural," Molly said. The plan, as she explained it over the following forty minutes, involved a rented clown costume, a magician's collapsible sack, and a diversion involving the neighbour's Labrador and eleven sausages. Bill's contribution was a pair of binoculars he had already broken and an insistence that they synchronise watches, despite neither of them owning a watch. He synchronised his phone with the microwave instead, which put them both nine minutes into the future. There were, Molly conceded, complications. The Pemberton child was turning six, which meant thirty guests, four hovering mothers, and a professional entertainer named Mister Sparkle who reportedly boxed on weekends. Bill wiped mustard from his thumb and considered the ruined napkin. "We're stealing from children," he said. "Small children," Molly agreed. "They can't identify us. They can't even read." Bill nodded slowly, reassured, and ordered another plate of chips.
created_at
"2026-07-25T13:33:59.642Z"
7/25/2026, 1:33:59 PM
model
"anthropic/claude-opus-5-20260723"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
3971
moderation_latency
(null)
generation_time
25992
tokens_prompt
75
tokens_completion
385
native_tokens_prompt
93
native_tokens_completion
525
native_tokens_completion_images
(null)
native_tokens_reasoning
50
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"end_turn"
service_tier
"default"
usage
0.01359
router
(null)
provider_responses
0
endpoint_id
"76cb4608-f48c-483d-8da8-9957fb44244e"
id
"msg_011CdNrx6SffqEcarznXcCb6"
is_byok
false
latency
1808
model_permaslug
"anthropic/claude-opus-5-20260723"
provider_name
"Amazon Bedrock"
status
200
user_agent
"langchainjs-openai/1.0.0 ((node/v24.18.0; linux; x64))"
http_referer
(null)
request_id
"req-1784986439-8yzYaw3QbXlgafSxbxfq"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1784986439-4I3IevHxPwQ43VLy1t0m"
upstream_id
"msg_011CdNrx6SffqEcarznXcCb6"
total_cost
0.01359
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Amazon Bedrock"
response_cache_source_id
(null)
data_region
"global"
Evaluation details
Result
Evaluator
Details
Meta Data
93.65%
Matches word count
n/a
neededClean
false
words
209
99.70%
Dialogue to Total Word Ratio
Ratio: 7.66%, Deviation: 2.34%
neededClean
false
wordsTotal
209
wordsDialogue
16
96.6740%