NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M37J07EJ8EGNC7XV30HFZ5AE
Run Details
Model:
GPT-6 Sol
Test:
Dialogue tags
Scenario:
Write 500 words with 50% dialogue
Duration:
13.5s
Cost:
$0.007212
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
98.71%
Matches word count
(details)
0.70%
Dialogue to Total Word Ratio
(details)
49.7066%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 500 word scene that has 50% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly spread a crayon-drawn floor plan across the café table. A purple dinosaur smiled from one corner. Across the top, in enormous letters, someone had written HAPPY BIRTHDAY, MAISIE. “Where did you get this?” Bill asked. “Community noticeboard.” “This is a child’s drawing.” “It’s a surprisingly accurate child’s drawing. Look. Kitchen, living room, back garden.” Bill squinted. “That’s a dragon.” “That’s the garden shed. Focus.” Rain tapped the window. Around them, respectable people drank coffee, unaware that two criminals were plotting the least respectable raid in suburban history. “The party starts at two,” Molly said. “At two fifteen, the magician arrives. At two thirty, he pulls a rabbit from a hat.” “We’re stealing from children while they watch a magician?” “We’re borrowing from children permanently. And we wait until they’re outside for the rabbit.” Bill studied the plan. “How many presents?” “Forty-two, according to the invitation.” “Invitations don’t say that.” “This one says ‘Forty-two guests, gifts welcome.’ I made an inference.” “That’s not planning. That’s arithmetic wearing a trench coat.” Molly tapped a square marked TABLE. “Presents go here. We enter through the side gate disguised as entertainers.” “I thought we were disguised as caterers.” “The caterers have been booked. The clown cancelled.” Bill leaned back. “Why did the clown cancel?” “Personal reasons.” “Did you ask?” “He said the birthday girl had bitten him last year.” At the next table, a woman glanced over. Molly smiled brightly and turned the drawing upside down. “Fine,” Bill whispered. “I’m the clown. What are you?” “The clown’s assistant.” “What does a clown’s assistant do?” “Helps carry bags.” “Suspiciously large bags.” “Balloon bags.” “You can’t call a sack a balloon bag just because you put one balloon in it.” “I’ll put in two.” A waiter arrived with their order. Molly slid the plan under her plate while Bill accepted a slice of cake. The waiter paused at Bill’s red nose, which he had been testing since breakfast. “Dress rehearsal,” Bill explained. The waiter nodded without enthusiasm and left. Molly waited until he was out of earshot. “At three, the children line up for cake. We collect the presents and leave.” “And when someone asks where we’re going?” “We say we’re taking them to the car.” “Whose car?” “The clown car.” “We drive a hatchback.” “We’ll park far away.” Bill took a bite of cake. “What if the rabbit sees us?” “The rabbit won’t object.” “You don’t know that. Animals have instincts.” Molly folded the drawing. Bill noticed something on the back: a second picture, showing a girl holding a trophy beside a police officer. He read the caption aloud. “Maisie’s mum catches the bad guys.” Molly grabbed the paper. “Her mum’s a police officer?” Bill said. “Apparently.” “Does Maisie bite her too?” Molly looked at the rain, then at Bill’s cake. “New plan. We send a card.” “With what inside?” “Money for a present.” Bill removed the red nose. “Excellent. We can rob an adult for it.” “No, Bill.” “A wealthy adult?” “No.” He sighed and reached for his wallet.
created_at
"2026-09-23T16:36:23.641Z"
9/23/2026, 4:36:23 PM
model
"openai/gpt-6-sol-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
1830
moderation_latency
(null)
generation_time
13507
tokens_prompt
75
tokens_completion
918
native_tokens_prompt
66
native_tokens_completion
708
native_tokens_completion_images
0
native_tokens_reasoning
0
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.007212
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790181383-V4Q9XptGQlHOcVinuTnS"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790181383-S9ZDPWzKrS0QmdZIvD7y"
upstream_id
"resp_0337cfbf7689c7c3016ab40007bf6487d1b9e1cc5a5c293ab5"
provider_responses
0
endpoint_id
"1cfc7d9d-4404-4b8e-9ee4-58ae45c9dcd4"
id
"resp_0337cfbf7689c7c3016ab40007bf6487d1b9e1cc5a5c293ab5"
is_byok
false
latency
522
model_permaslug
"openai/gpt-6-sol-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.007212
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
98.71%
Matches word count
n/a
neededClean
false
words
506
0.70%
Dialogue to Total Word Ratio
Ratio: 59.96%, Deviation: 9.96%
neededClean
false
wordsTotal
507
wordsDialogue
304
49.7066%