NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M37HG2BQFF7R9S9G4GZ8D5CD
Run Details
Model:
GPT-6 Sol (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 500 words with 70% dialogue
Duration:
18.1s
Cost:
$0.008702
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
14.30%
Matches word count
(details)
1.31%
Dialogue to Total Word Ratio
(details)
7.8064%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 500 word scene that has 70% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly spread a crayon map across the café table. Someone had drawn a smiling sun in the corner and labeled it “DO NOT TRUST.” Bill squinted. “Why does the sun have a mustache?” “Because I drew this during my nephew’s art lesson. Focus. Saturday, two o’clock. Children’s party. Twenty guests. One enormous pile of presents.” “Isn’t stealing birthday presents a bit low, even for us?” “They aren’t birthday presents. It’s a gift-exchange party. Everyone brings one, and they leave them under a banner that says PRESENT MOUNTAIN.” “That does sound like it’s asking for trouble.” “Exactly. We’ll be the trouble.” Bill stirred his tea. “Why are we doing this? Last week you said you’d found a score that could change our lives.” “I had. Then I checked the address and learned the diamond exhibition was in Belgium.” “So this is the backup?” “The presents are wrapped. There could be anything in there.” “There could be socks.” “Designer socks.” “There could be a recorder.” “A vintage musical instrument.” “There could be a handmade card saying ‘I love you, Grandma.’” Molly paused. “We return those.” “Good. I don’t want Grandma after us.” A waitress set down two slices of cake. Molly waited until she’d gone, then pointed at the map. “We go in disguised as entertainers.” “What kind?” “A magician and his assistant.” Bill brightened. “I’m the magician.” “You’re the assistant.” “I know a trick.” “Making rent disappear isn’t a trick.” “I can pull a coin from somebody’s ear.” “You tried that on the bus and got slapped.” “The man had very small ears. I panicked.” Molly took a bite of cake. “Fine. You can be the magician. I’ll announce the big finale, everyone looks at you, and I wheel Present Mountain out of sight.” “Where do you wheel it?” “That’s a detail I’m still perfecting.” “Is this map finished?” “It has the location of the cake.” “That’s not a route.” “It’s a landmark.” Bill leaned closer. “What’s this square?” “Bouncy castle.” “And the skull beside it?” “That’s the birthday host.” “You said it wasn’t a birthday party.” “She’s a parent named Karen. I thought the skull captured her energy.” From the next table came a shriek of laughter. Molly folded the map so the skull vanished. Bill lowered his voice. “How do we get invited?” “We’re already booked.” “You booked us as entertainers?” “I sent an email.” “Can you do magic?” “No. I told them we perform educational comedy about sharing.” Bill looked at the map, then at Molly. “Sharing what?” “That’s where your coin trick comes in.” “My coin trick ends with me owning somebody else’s coin.” “Educational. We’ll demonstrate what not to do.” A phone buzzed. Molly read the message and went still. “What?” “Karen says the children are bringing presents for the hospital toy drive.” Bill pushed away his cake. “We’re not stealing those.” “Obviously not.” “So what’s the plan now?” Molly stared at the smiling sun. “We do the show.” “You can’t be serious.” “Twenty children, Bill. One magician. A lesson about sharing.” “And your part?” “I’m the assistant. I’ll take the coins away before you get slapped.”
created_at
"2026-09-23T16:27:34.142Z"
9/23/2026, 4:27:34 PM
model
"openai/gpt-6-sol-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
5042
moderation_latency
(null)
generation_time
18113
tokens_prompt
75
tokens_completion
934
native_tokens_prompt
66
native_tokens_completion
857
native_tokens_completion_images
0
native_tokens_reasoning
144
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.008702
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790180854-HFxTRlmTuSRTK35tEV45"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790180854-rA5H2NU08PV9uroMZIYJ"
upstream_id
"resp_05e4e19189f17c40016ab3fdf6443887d1b39abc5daf251fcb"
provider_responses
0
endpoint_id
"1cfc7d9d-4404-4b8e-9ee4-58ae45c9dcd4"
id
"resp_05e4e19189f17c40016ab3fdf6443887d1b39abc5daf251fcb"
is_byok
false
latency
542
model_permaslug
"openai/gpt-6-sol-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.008702
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
14.30%
Matches word count
n/a
neededClean
false
words
521
1.31%
Dialogue to Total Word Ratio
Ratio: 79.31%, Deviation: 9.31%
neededClean
false
wordsTotal
522
wordsDialogue
414
7.8064%