NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M36NPEGB0T7A7E4ZGG5FCSN2
Run Details
Model:
GPT-6 Luna (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 500 words with 70% dialogue
Duration:
14.0s
Cost:
$0.000543
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
97.63%
Matches word count
(details)
2.42%
Dialogue to Total Word Ratio
(details)
50.0219%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 500 word scene that has 70% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
“Tonight,” said Molly, unfolding a map of the community hall across the café table, “we enter the children’s party, seize every present, and leave before anyone notices.” Bill studied the map upside down. “What if they notice before we leave?” “Then we’ll have to be charming.” “I’ve never been charming under pressure.” “You once talked your way out of a parking ticket.” “I was parked in a bakery.” “Exactly. You made the officer hungry.” Bill turned the map around. It showed a cheerful building surrounded by balloons. “This looks less like a vault and more like somewhere a clown might live.” “The presents are the vault,” Molly said. “Wrapped boxes. Untouched. Uncounted. Practically begging.” “Presents don’t beg.” “You’ve never heard one.” “I’ve heard you talk to a vending machine.” “It gave me crisps.” “That was because you shook it.” “That was a negotiation.” Molly tapped the map. “The party starts at six. We walk in wearing our best festive expressions, blend with the grown-ups, and carry out the loot.” “What are our best festive expressions?” “I have one.” She smiled. Bill recoiled. “That looks like you’ve just found a finger in a trifle.” “Then you do it.” He tried. His face settled into a grimace of cautious regret. “Perfect,” Molly said. “You look like a disappointed uncle.” “I am a disappointed uncle. I’m disappointed in this plan.” “You said you wanted a big score.” “I imagined a museum. Maybe a yacht. Not a room full of birthday cards addressed to Toby.” “Toby could have excellent taste.” “Last year’s presents were mostly socks.” “Socks are portable.” “They’re also not worth stealing.” Molly leaned closer. “We take the presents, sell the valuable ones, and disappear.” “What valuable ones?” “Anything with batteries.” “That includes the smoke alarm.” “Bill, focus.” “I am focused. I’m picturing a six-year-old discovering an empty present table.” For a moment, Molly stopped smiling. From the next room came a child’s squeal as someone dropped a spoon. Both thieves glanced toward the sound. “That’s the thing about children,” Bill said. “They notice when their cake goes missing.” “We’re not stealing the cake.” “Not today.” Molly folded the map. “Fine. We revise the plan.” “Into something less criminal?” “Into something more profitable.” “Which is?” “We arrive early, help carry things in, and get paid.” Bill blinked. “Paid to move presents?” “Paid to move tables. Maybe balloons. Possibly cake.” “And we don’t steal anything?” “We steal the leftover slices after the guests leave.” “That’s still theft.” “It’s dessert with paperwork missing.” Bill stood, straightening his jacket. “I’ll bring the trolley.” “For the presents?” “For the tables.” Molly picked up the map. “And if someone asks why two suspicious adults are helping at a children’s party?” “We’ll tell them we’re volunteers.” “You’re a terrible liar.” “I’m an excellent uncle.” “You don’t have nieces or nephews.” “Then I’ll have to be charming.” Molly considered this, then nodded. “Just don’t shake the vending machine.” “I make no promises.”
created_at
"2026-09-23T08:21:43.056Z"
9/23/2026, 8:21:43 AM
model
"openai/gpt-6-luna-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
3710
moderation_latency
(null)
generation_time
14001
tokens_prompt
75
tokens_completion
1169
native_tokens_prompt
66
native_tokens_completion
1073
native_tokens_completion_images
0
native_tokens_reasoning
387
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.0005431
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790151703-9snn5Je22FAPnCKfJzEz"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790151703-7lQg6PFCJK00j5MMS5lD"
upstream_id
"resp_0b0d9d69a40fdfed016ab38c1724b887d19583a15bdf105c1d"
provider_responses
0
endpoint_id
"05e94e02-b9c2-4bbb-ba55-4082ee9ad687"
id
"resp_0b0d9d69a40fdfed016ab38c1724b887d19583a15bdf105c1d"
is_byok
false
latency
332
model_permaslug
"openai/gpt-6-luna-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.0005431
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
97.63%
Matches word count
n/a
neededClean
false
words
493
2.42%
Dialogue to Total Word Ratio
Ratio: 78.63%, Deviation: 8.63%
neededClean
false
wordsTotal
496
wordsDialogue
390
50.0219%