NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-29-01M3P2CT0WERBS8BJFPHPF27MH
Run Details
Model:
Claude Sonnet 5.5 (Reasoning)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
20.9s
Cost:
$0.029050
Provider:
Claude Platform on AWS
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
99.9989%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly lay flat on the hedge-top, binoculars trained on the bouncy castle across the lawn. Bill wheezed beside her, wearing a balloon-animal hat he insisted was tactical camouflage. Below, forty children shrieked around a table buried in gift-wrapped treasure. Somewhere in that pile sat a remote-controlled monster truck, three scooters, and a mysterious box the size of a refrigerator. Molly had been planning this job for six weeks, which was five weeks longer than the birthday girl had been planning her own. A clown drifted past, juggling with sinister confidence. Bill shuddered. He had never recovered from the incident at the county fair, and he refused to discuss it. "Clown at two o'clock," he whispered. "I'm not doing this." Molly lowered the binoculars and gave him the look she reserved for amateurs and people who chewed loudly. Then she pulled two foam pool noodles from her duffel bag, handed him one, and pointed at the castle, where the children were beginning a game of pin the tail on the donkey. "Blindfolds, Bill," she said. "Nobody sees anything." Bill considered this, adjusted his hat, and slowly nodded. "Fine, but I'm keeping the hat," he said. Somewhere, a piƱata began to weep.
created_at
"2026-09-29T07:52:15.137Z"
9/29/2026, 7:52:15 AM
model
"anthropic/claude-sonnet-5.5-20260928"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
3944
moderation_latency
(null)
generation_time
20916
tokens_prompt
75
tokens_completion
596
native_tokens_prompt
95
native_tokens_completion
2886
native_tokens_completion_images
(null)
native_tokens_reasoning
2445
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"end_turn"
service_tier
"default"
usage
0.02905
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.21.0; linux; x64))"
http_referer
(null)
request_id
"req-1790668335-0nKlpcFMmuVZ9eZEDBwh"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790668335-auQkymEImWwZWTeOrHFJ"
upstream_id
"msg_011CfXNH1qUm2mnvq5JP95Tj"
provider_responses
0
endpoint_id
"99aaad94-923b-4fc1-b763-271ed5486f7a"
id
"msg_011CfXNH1qUm2mnvq5JP95Tj"
is_byok
false
latency
695
model_permaslug
"anthropic/claude-sonnet-5.5-20260928"
provider_name
"Claude Platform on AWS"
status
200
total_cost
0.02905
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Claude Platform on AWS"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
100.00%
Dialogue to Total Word Ratio
Ratio: 9.31%, Deviation: 0.69%
neededClean
false
wordsTotal
204
wordsDialogue
19
99.9989%