NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-29-01M3P2DYEKQHC9ME671PFWA231
Run Details
Model:
Claude Sonnet 5.5 (Reasoning)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
24.9s
Cost:
$0.034140
Provider:
Claude Platform on AWS
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
100.0000%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly adjusted her fake mustache and peered through binoculars at the suburban house, where forty balloons bobbed above the porch. Bill sat beside her in the van, sweating through a rented clown costume that smelled of nacho cheese and regret. Blueprints covered the dashboard, along with a cake diagram and a photo of the birthday boy, Tyler, who reportedly had a vicious right hook. Bill honked his rubber nose nervously. Molly slapped his hand away. "Relax. Nobody suspects a clown," she whispered. "Everybody suspects a clown," Bill said. "That's the whole problem with clowns." Molly ignored him and checked her watch. The piƱata would be swung at two o'clock, and in the chaos of candy and screaming, two professionals could empty the gift table before a toddler could say mine. Bill squeezed out of the van, his enormous shoes flapping like startled fish. Molly followed, carrying a balloon animal named Kevin, which secretly contained a crowbar. "Just follow Kevin," she said. "Ready?" she asked. "Never," Bill said. Together they marched up the walkway toward the front door, ready for the greatest heist ever attempted against people who still needed naps. Somewhere inside, a tiny voice began singing off-key. Bill winced.
created_at
"2026-09-29T07:52:52.44Z"
9/29/2026, 7:52:52 AM
model
"anthropic/claude-sonnet-5.5-20260928"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
3732
moderation_latency
(null)
generation_time
24918
tokens_prompt
75
tokens_completion
781
native_tokens_prompt
95
native_tokens_completion
3395
native_tokens_completion_images
(null)
native_tokens_reasoning
2950
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"end_turn"
service_tier
"default"
usage
0.03414
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.21.0; linux; x64))"
http_referer
(null)
request_id
"req-1790668372-a1HlnCfKoE8QrTj4O0Wk"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790668372-9kcgCUdDJNjkljmn8hHf"
upstream_id
"msg_011CfXNKmHhcHoNiX5cXoqJD"
provider_responses
0
endpoint_id
"99aaad94-923b-4fc1-b763-271ed5486f7a"
id
"msg_011CfXNKmHhcHoNiX5cXoqJD"
is_byok
false
latency
645
model_permaslug
"anthropic/claude-sonnet-5.5-20260928"
provider_name
"Claude Platform on AWS"
status
200
total_cost
0.03414
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Claude Platform on AWS"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
100.00%
Dialogue to Total Word Ratio
Ratio: 9.95%, Deviation: 0.05%
neededClean
false
wordsTotal
201
wordsDialogue
20
100.0000%