NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M37PPVEQT13C8FFYKQC559VH
Run Details
Model:
Claude Opus 5.5 (Reasoning)
Test:
Dialogue tags
Scenario:
Write 200 words with 50% dialogue
Duration:
23.4s
Cost:
$0.040040
Provider:
Claude Platform on AWS
Evaluation summary
Score
Evaluator
Details Link
99.98%
Matches word count
(details)
99.22%
Dialogue to Total Word Ratio
(details)
99.6043%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 50% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly unrolled the blueprint across the hood of the getaway van, which was really her mother's minivan with a sticker reading "Proud Grandma of an Honor Student." A juice box rolled off the dashboard in solidarity. "Timmy Pembrook turns seven at two o'clock," she said. "Fourteen presents, minimum. Sixteen if Aunt Carol feels guilty again." Bill squinted at the drawing. "Is this the house or a giraffe?" "It's the house. The neck is the driveway." He nodded, as though that settled it. Bill had once robbed a bank and left with only the complimentary pens, so Molly did not expect much. "So what's my job this time?" he asked. "You're the clown." "I hate clowns. They have too many pockets." "Children hate clowns too. That's why nobody will look at you directly. While they're screaming, I slip out the back with the gift table." "The whole table?" "It has wheels, Bill. I checked." Bill adjusted the rubber nose she'd handed him. It honked every time he breathed, which seemed like a design flaw for a stealth operation. "And if the birthday boy catches us?" Molly smiled and patted her bag. "Then we use the emergency bribe." "Cake?" "Better. His mother's iPad password."
created_at
"2026-09-23T17:58:39.337Z"
9/23/2026, 5:58:39 PM
model
"anthropic/claude-opus-5.5-20260921"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
5859
moderation_latency
(null)
generation_time
23368
tokens_prompt
75
tokens_completion
505
native_tokens_prompt
95
native_tokens_completion
1983
native_tokens_completion_images
(null)
native_tokens_reasoning
1540
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"end_turn"
service_tier
"default"
usage
0.04004
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790186319-Di57PFQSz27h8ZlWxaxO"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790186319-GNmrMPSkCJlh7UVnAbPN"
upstream_id
"msg_011CfLofZdeZFnNG4CawfPfK"
provider_responses
0
endpoint_id
"3a2388bc-3740-4e64-a1f5-4e301726b6b9"
id
"msg_011CfLofZdeZFnNG4CawfPfK"
is_byok
false
latency
1235
model_permaslug
"anthropic/claude-opus-5.5-20260921"
provider_name
"Claude Platform on AWS"
status
200
total_cost
0.04004
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Claude Platform on AWS"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
99.98%
Matches word count
n/a
neededClean
false
words
202
99.22%
Dialogue to Total Word Ratio
Ratio: 52.97%, Deviation: 2.97%
neededClean
false
wordsTotal
202
wordsDialogue
107
99.6043%