NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M37PDMZ7KTQ1BACC81YVBJZY
Run Details
Model:
Claude Opus 5.5 (Reasoning)
Test:
Dialogue tags
Scenario:
Write unattributed dialogue
Duration:
12.8s
Cost:
$0.019596
Provider:
Claude Platform on AWS
Evaluation summary
Score
Evaluator
Details Link
100.00%
Count dialogue tags
(details)
Model input
User/Human
You are an expert novel writer in the cozy mystery genre. Write a scene of about 200 words, with dialogue between two characters. The dialogue should be unattributed, meaning the speaker is not explicitly mentioned. The scene should be set in a small town bakery. Only return the story, nothing else.
Model response
Messages
Meta
AI
The bell over the door of the Crumb & Kettle jingled just as the first batch of cinnamon rolls came out of the oven. "You're open early." "I'm always open early. You're the one who's never up before nine, Walter." "Couldn't sleep. Not after what happened at the church bake sale." "If this is about Margaret's lemon bars, I already told the sheriff everything I know." "Did you tell him the recipe card was missing?" "How do you know about that?" "Because I found it. Tucked inside a hymnal, third pew from the back." "Well, don't just stand there dripping rain on my clean floor. Let me see it." "Look at the bottom. That's not Margaret's handwriting." "No, it isn't. See that lowercase g with the fancy loop? Only one person in Maple Hollow writes like that." "The new choir director." "Who, I might add, bought two dozen of my blueberry muffins yesterday and asked an awful lot of questions about which nuts I keep in the back." "So what do we do?" "We do what any sensible person would do. We have a cinnamon roll, we pour ourselves some coffee, and then we pay him a friendly little visit." "Friendly?" "Walter, dear, I'm always friendly. Right up until I'm not."
created_at
"2026-09-23T17:53:37.785Z"
9/23/2026, 5:53:37 PM
model
"anthropic/claude-opus-5.5-20260921"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
5988
moderation_latency
(null)
generation_time
12763
tokens_prompt
78
tokens_completion
424
native_tokens_prompt
109
native_tokens_completion
958
native_tokens_completion_images
(null)
native_tokens_reasoning
527
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"end_turn"
service_tier
"default"
usage
0.019596
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790186017-Sv8mBuRRdwm0dP1LmcYa"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790186017-fs9NVjnjCV3UNPlwwOoW"
upstream_id
"msg_011CfLoHLEPMdKz1F5mZeUNH"
provider_responses
0
endpoint_id
"3a2388bc-3740-4e64-a1f5-4e301726b6b9"
id
"msg_011CfLoHLEPMdKz1F5mZeUNH"
is_byok
false
latency
1700
model_permaslug
"anthropic/claude-opus-5.5-20260921"
provider_name
"Claude Platform on AWS"
status
200
total_cost
0.019596
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Claude Platform on AWS"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Count dialogue tags
n/a
neededClean
false
dialogueTags
(empty)