NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M36R8GKDK4ENE9YZYGJ4E25B
Run Details
Model:
GPT-6 Luna (Reasoning, High)
Test:
Dialogue tags
Scenario:
Write 500 words with 50% dialogue
Duration:
1m 7s
Cost:
$0.003725
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
95.99%
Matches word count
(details)
84.59%
Dialogue to Total Word Ratio
(details)
90.2902%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 500 word scene that has 50% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Rain stitched silver lines across the bakery window, where Molly spread a map of the neighborhood beside two cooling cinnamon rolls and one suspicious raisin. “Bill, the invitation says the birthday party has forty children, a magician, and enough presents to make our pension plan look respectable for once, perhaps.” Bill examined the map upside down, then rotated it until the bakery occupied the ocean and the school floated near a mysterious legend of dragons. “Excellent. If we arrive by submarine, nobody will suspect a thing, except the children, who are famously observant and armed with juice boxes at parties.” Molly tapped the party address, which was the community hall, a cheerful building with painted balloons and a sign reading WELCOME, GROWNUPS in several colors. “Here’s my strategy: we enter as entertainers, distract everyone with a tasteful routine, then leave carrying every gift beneath our dignity and our matching hats.” She produced two paper crowns from her coat. One had glittery stars; the other said Birthday Princess, though Bill wore it without complaint or fuss. “Why am I Princess?” “Because you have the posture.” “I have the posture of a damp umbrella.” “An elegant damp umbrella.” Outside, a passing child dragged a balloon shaped like a dinosaur. Bill flinched, and the balloon bobbed after him like an unpaid debt, smiling maliciously. “Children detect fear.” “Then I shall look fearless,” Bill whispered, tightening his crown. “You look like royalty facing a tax audit in a thunderstorm, perhaps.” Molly circled the word ALL on her napkin, then crossed it out and wrote MOST, then MAYBE NONE, as her expression softened like warm butter. “Wait. Those gifts are for children. We could take the cake instead.” “That seems less lucrative and more crumbly than expected.” From the hall came a burst of children’s laughter, followed by applause. Molly listened, looking like someone remembering her own birthday fondly from long ago. “We’re thieves, not monsters.” “The distinction is important, though your crown is committing a separate crime against public taste and several mirrors in every town.” He folded the map into a paper boat, perhaps surrendering to geography. It sailed one inch across the table and capsized in cinnamon without warning. “New plan. We go inside, announce ourselves, and offer to perform a magic trick.” “Can we vanish before the cake arrives without any witnesses ever?” At the hall, a volunteer opened the door before they could knock. She mistook their crowns for costumes and handed them tambourines with a smile. “Are we expected to steal anything?” “Only the show.” “Then let’s steal the show and leave everything else where it belongs, and act like professionals.” By the finale, Bill juggled oranges while Molly performed a trick with a handkerchief and one patient rabbit borrowed from the magician for the occasion. “Your payment is applause and leftover cake.” “We accept both, though the cake’s security seems unusually lax tonight, with gratitude and without any suspicious questions.”
created_at
"2026-09-23T09:06:32.184Z"
9/23/2026, 9:06:32 AM
model
"openai/gpt-6-luna-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
4707
moderation_latency
(null)
generation_time
67122
tokens_prompt
75
tokens_completion
1448
native_tokens_prompt
66
native_tokens_completion
7436
native_tokens_completion_images
0
native_tokens_reasoning
6808
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.0037246
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.21.0; linux; x64))"
http_referer
(null)
request_id
"req-1790154392-HzugoAFV1jGgDyeZoeq3"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790154392-OYuJ4JkG19z5lLGL7Njf"
upstream_id
"resp_039ba0e3e9cef31b016ab396984ad487d18a299d7d5c208a55"
provider_responses
0
endpoint_id
"05e94e02-b9c2-4bbb-ba55-4082ee9ad687"
id
"resp_039ba0e3e9cef31b016ab396984ad487d18a299d7d5c208a55"
is_byok
false
latency
845
model_permaslug
"openai/gpt-6-luna-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.0037246
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
95.99%
Matches word count
n/a
neededClean
false
words
492
84.59%
Dialogue to Total Word Ratio
Ratio: 48.17%, Deviation: 1.83%
neededClean
false
wordsTotal
492
wordsDialogue
237
90.2902%