NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M37NY98V8MFDZDXKSY59M02G
Run Details
Model:
Claude Opus 5.5 (Reasoning)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
18.4s
Cost:
$0.037940
Provider:
Claude Platform on AWS
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
99.96%
Dialogue to Total Word Ratio
(details)
99.9783%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly spread the blueprints across the hood of the getaway van, which was really her mother's minivan with a "Baby on Board" sticker peeled halfway off. The blueprints were actually a birthday invitation, decorated with glitter unicorns and the words Tyler Turns Seven. Bill leaned over it, squinting through the ski mask he had insisted on wearing since breakfast. "The bouncy castle is our weak point," he said. Molly stared at him for a long moment. "It's inflatable, Bill." "Exactly. Nobody expects a thief in a bouncy castle." She had to admit, grudgingly, that this was true. Nobody expected a thief anywhere near a bouncy castle, mostly because no thief in recorded history had ever been stupid enough to plan around one. Still, the gift table sat twelve feet from its netted entrance, piled high with boxes wrapped in dinosaur paper and bows the size of dinner plates. Somewhere in that pile, according to reliable intelligence from a gossiping soccer mom, was a limited-edition robot worth four hundred dollars. Molly checked her watch. Cake was at three. Every child would be distracted by frosting, and every parent by their phones. "Fine." Bill pumped his fist and hit the side mirror.
created_at
"2026-09-23T17:45:14.283Z"
9/23/2026, 5:45:14 PM
model
"anthropic/claude-opus-5.5-20260921"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
5315
moderation_latency
(null)
generation_time
18318
tokens_prompt
75
tokens_completion
496
native_tokens_prompt
95
native_tokens_completion
1878
native_tokens_completion_images
(null)
native_tokens_reasoning
1446
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"end_turn"
service_tier
"default"
usage
0.03794
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790185514-rZiXk2gXKTeDJf9D3NQN"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790185514-TyR8P8sAvmW89SwL761n"
upstream_id
"msg_011CfLneDGA17LYXVh6q2VS5"
provider_responses
0
endpoint_id
"3a2388bc-3740-4e64-a1f5-4e301726b6b9"
id
"msg_011CfLneDGA17LYXVh6q2VS5"
is_byok
false
latency
1191
model_permaslug
"anthropic/claude-opus-5.5-20260921"
provider_name
"Claude Platform on AWS"
status
200
total_cost
0.03794
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Claude Platform on AWS"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
99.96%
Dialogue to Total Word Ratio
Ratio: 11.44%, Deviation: 1.44%
neededClean
false
wordsTotal
201
wordsDialogue
23
99.9783%