NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-10-08-01M4E09E8FWKWXMATMJF3A6R0W
Run Details
Model:
Claude Haiku 5.5 (Reasoning, Low)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
3.1s
Cost:
$0.000187
Provider:
Claude Platform on AWS
Evaluation summary
Score
Evaluator
Details Link
0.00%
Matches word count
(details)
82.08%
Dialogue to Total Word Ratio
(details)
41.0377%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
"Right, so the plan is simple," Molly whispered, tapping the blueprint. "We go in through the chimney." "Molly, it's a bouncy castle place. There's no chimney." "There's always a chimney. Everyone has a chimney." "Fine. Then what?" "Then you distract the clown." "Why me? I'm the one who sneezes around balloons." "Because you have the trustworthy face, Bill. Nobody suspects a face like yours." "That's a terrible compliment." "Take it anyway. Once the clown's busy, we grab the presents from the corner table." "And the birthday kid? What if he sees us?" "Five-year-olds don't count as witnesses, Bill." "He bit me last time I met one." "Focus. Wrapped boxes, ribbons, shiny bows. We're in and out in four minutes." "Four minutes? The cake alone takes four minutes." "We're not taking the cake. Bill, I swear, if you eat a single frosting rose..." "I'm only checking the quality control."
created_at
"2026-10-08T14:57:14.019Z"
10/8/2026, 2:57:14 PM
model
"anthropic/claude-haiku-5.5-20261007"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
792
moderation_latency
(null)
generation_time
3091
tokens_prompt
75
tokens_completion
234
native_tokens_prompt
95
native_tokens_completion
354
native_tokens_completion_images
(null)
native_tokens_reasoning
0
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"end_turn"
service_tier
"default"
usage
0.0001865
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.21.0; linux; x64))"
http_referer
(null)
request_id
"req-1791471434-5kbbvoPUljJsAs52r9qK"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1791471434-La0ilTuWle62KskyexL9"
upstream_id
"msg_011CfpxxtzMqKnquvLQmBuVW"
provider_responses
0
endpoint_id
"b6cdf493-3a61-441f-a134-e7944fddb980"
id
"msg_011CfpxxtzMqKnquvLQmBuVW"
is_byok
false
latency
791
model_permaslug
"anthropic/claude-haiku-5.5-20261007"
provider_name
"Claude Platform on AWS"
status
200
total_cost
0.0001865
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Claude Platform on AWS"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
0.00%
Matches word count
n/a
neededClean
false
words
148
82.08%
Dialogue to Total Word Ratio
Ratio: 96.67%, Deviation: 6.67%
neededClean
false
wordsTotal
150
wordsDialogue
145
41.0377%