NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M37HN6FV0WF2QA8YEF8K4M3F
Run Details
Model:
GPT-6 Sol (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 500 words with 70% dialogue
Duration:
19.3s
Cost:
$0.008272
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
81.27%
Matches word count
(details)
6.32%
Dialogue to Total Word Ratio
(details)
43.7957%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 500 word scene that has 70% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly spread a crayon drawing across the café table. It showed a castle, twelve stick figures, and a dog the size of the moon. “That’s the floor plan,” she said. Bill studied it. “Why is the kitchen wearing a crown?” “That’s the castle.” “And the thing labeled KITCHEN?” “The kitchen. Keep up.” Bill turned the paper sideways. “Who drew this?” “My nephew. He’s attending the party.” “We’re using intelligence from a six-year-old?” “He’s seven. His birthday was in March.” “Sorry. A seasoned operative.” Molly tapped the crown. “Party starts at two. We arrive at two fifteen, when everyone’s distracted by the magician.” “What magician?” “The invitation says ‘magical entertainment.’” “That could mean anything. At my cousin’s wedding, it meant a chocolate fountain.” “Was it magical?” “It broke down and sprayed the vicar. So yes.” A waitress passed their table. Molly folded the drawing beneath her menu. Bill lowered his voice. “Let’s discuss the presents. How many?” “Twenty, maybe thirty.” “That’s a lot to carry.” “Which is why we need a disguise.” Bill brightened. “I still have the gorilla suit.” “No.” “It’s convincing.” “It has a zip across the forehead and smells like onions.” “Gorillas eat onions.” “Not at children’s parties. We’re going as entertainers. You’ll make balloon animals.” “I can’t.” “You made one last week.” “That was a balloon. It was shaped like a balloon.” “Call it a snake.” “What will you do?” “Face painting.” “You can’t paint.” “I can draw a cat.” “You can draw whiskers. Last time, you put them on my forehead.” “You kept moving.” “You told me to sneeze.” The waitress returned with two coffees. Bill waited until she left, then leaned closer. “Why are we stealing presents from children, exactly?” Molly looked offended. “We’re stealing presents from one child. The others brought them.” “That clears it up completely.” “And we’re not keeping them. Her father owes my sister four thousand pounds.” “So this is debt collection?” “Technically. We’ll return everything when he pays.” “What if there’s a bicycle?” “Then we’ll return it with the bell intact. I’m not a monster.” Bill stirred his coffee. “Your sister knows about this?” “No. She suggested a polite letter.” “A radical approach.” “Her last polite letter included six exclamation marks and a photograph of her looking disappointed. He framed it.” Bill glanced at the drawing again. “What’s the enormous dog?” “That’s Princess.” “Princess has eight legs.” “My nephew ran out of room for her tail.” A phone buzzed. Molly read the message and went still. “What?” Bill asked. “The party’s been moved.” “Where?” “My sister’s house. The birthday girl has chickenpox.” “So there won’t be a party?” “There will. Small one. Family only.” Bill nudged the drawing back to her. “We could visit your sister instead. Take cake. Write another letter.” Molly considered this. “With seven exclamation marks?” “Eight. We’re professionals.” She folded the floor plan into her pocket. “Fine. But I’m bringing the balloons.” “Good. Your nephew deserves a snake.” “And you’re not wearing the gorilla suit.” Bill looked wounded. “Even if I wash it?” “Especially then. It might shrink.”
created_at
"2026-09-23T16:30:22.428Z"
9/23/2026, 4:30:22 PM
model
"openai/gpt-6-sol-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
5110
moderation_latency
(null)
generation_time
19012
tokens_prompt
75
tokens_completion
1088
native_tokens_prompt
66
native_tokens_completion
814
native_tokens_completion_images
0
native_tokens_reasoning
88
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.008272
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790181022-NveynutzgZkHLaagKhFD"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790181022-6ueIOCw9l7wyMQ4Rz5ao"
upstream_id
"resp_05432c260a4063aa016ab3fe9e8db887d1b3ff330526b84d0c"
provider_responses
0
endpoint_id
"1cfc7d9d-4404-4b8e-9ee4-58ae45c9dcd4"
id
"resp_05432c260a4063aa016ab3fe9e8db887d1b3ff330526b84d0c"
is_byok
false
latency
1792
model_permaslug
"openai/gpt-6-sol-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.008272
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
81.27%
Matches word count
n/a
neededClean
false
words
512
6.32%
Dialogue to Total Word Ratio
Ratio: 77.43%, Deviation: 7.43%
neededClean
false
wordsTotal
514
wordsDialogue
398
43.7957%