NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M37HH311W2YQDRC78EB3NJGG
Run Details
Model:
GPT-6 Sol
Test:
Dialogue tags
Scenario:
Write 500 words with 70% dialogue
Duration:
16.5s
Cost:
$0.007362
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
0.00%
Matches word count
(details)
2.78%
Dialogue to Total Word Ratio
(details)
1.3906%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 500 word scene that has 70% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly spread a crayon map across the café table. A purple dinosaur grinned from the paper’s corner. “Tell me you didn’t draw this during the party.” “I drew it before the party,” Bill said. “During the party, I was conducting reconnaissance.” “You were wearing a balloon hat.” “Effective camouflage.” “You cried when the magician pulled a coin from your ear.” “It was my coin.” Molly tapped the map. “Sixteen children, two adults, one mountain of presents. We enter through the kitchen at three fifteen.” “That’s when they serve cake.” “Exactly. Everyone’s distracted.” “By cake,” Bill said. “Including me.” “You are not allowed cake until we’re gone.” “That wasn’t in the original plan.” “The original plan had you fitting inside a gift bag.” “I tested it. My shoulders are wider than anticipated.” Rain rattled the café windows. Bill adjusted his balloon hat, which squeaked against the chair behind him. “First problem,” he said. “The kitchen door has a lock.” “I have a key.” “How?” “The parents put one under a ceramic frog.” “Second problem. The frog is suspicious.” “The frog is ceramic.” “Exactly what it wants us to think.” Molly ignored him. “We cross the kitchen, collect the presents from the dining room, and leave through the garden gate.” “All the presents?” “All of them.” “Even the homemade ones?” “Especially the homemade ones. We can’t leave evidence of selective taste.” Bill studied the drawing. “What’s the big red circle?” “The dog.” “That dog is roughly the size of the house.” “It’s drawn to scale.” “And the little blue circle?” “You.” “That feels personal.” “The dog likes sausages. Bring sausages.” “I brought sausages last time.” “You ate them.” “I was establishing whether they were good sausages.” A waitress arrived with two coffees and paused at the map. Molly flipped it over, revealing a menu Bill had doodled on the back. “Planning a party?” the waitress asked. “Surprise party,” Molly said. “For a frog,” Bill added. The waitress nodded slowly and retreated. Molly leaned forward. “At three fourteen, you distract the dog. At three fifteen, I open the door. At three sixteen, you carry the presents.” “Why do I carry them?” “Because I’m carrying the cake.” “We’re stealing the cake now?” “No. We’re moving it away from the presents so nobody sees us.” “Where are you moving it?” “Outside.” “Near the garden gate?” “Possibly.” Bill narrowed his eyes. “You’re planning to eat it.” “Focus.” “What happens if a child sees us?” “We say we’re party helpers.” “What if they ask us to help?” “We help.” “What if they ask us to perform magic?” “You already know the coin trick.” “I know how to lose a coin. The magician did the rest.” Molly folded the map. “You can still back out.” Bill looked at the balloon hat in his lap. “What did you say the presents were for?” “The children’s shelter. This party is collecting donations.” “So we’re stealing presents from a donation party?” “We’re taking them to the shelter. The host asked us to make it a surprise delivery.” Bill stared. “We have permission?” “From everyone except the dog.” “Oh.” He picked up his coffee. “Can I have cake?” “If there’s any left.” “I’m buying more sausages.”
created_at
"2026-09-23T16:28:07.594Z"
9/23/2026, 4:28:07 PM
model
"openai/gpt-6-sol-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
2785
moderation_latency
(null)
generation_time
16500
tokens_prompt
75
tokens_completion
967
native_tokens_prompt
66
native_tokens_completion
723
native_tokens_completion_images
0
native_tokens_reasoning
0
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.007362
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790180887-u6Je2n6IFxgdbw1IdvJg"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790180887-3z03yoP5e63ASEMQ0VzT"
upstream_id
"resp_00a13d6efbfb07d3016ab3fe17c08487d1bfd81ee62802dfb3"
provider_responses
0
endpoint_id
"1cfc7d9d-4404-4b8e-9ee4-58ae45c9dcd4"
id
"resp_00a13d6efbfb07d3016ab3fe17c08487d1bfd81ee62802dfb3"
is_byok
false
latency
766
model_permaslug
"openai/gpt-6-sol-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.007362
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
0.00%
Matches word count
n/a
neededClean
false
words
534
2.78%
Dialogue to Total Word Ratio
Ratio: 78.46%, Deviation: 8.46%
neededClean
false
wordsTotal
534
wordsDialogue
419
1.3906%