NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M36RA0D04VMPHNVF3APBYJDB
Run Details
Model:
GPT-6 Luna (Reasoning, High)
Test:
Dialogue tags
Scenario:
Write 500 words with 50% dialogue
Duration:
1m 17s
Cost:
$0.004317
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
99.80%
Dialogue to Total Word Ratio
(details)
99.9004%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 500 word scene that has 50% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
“Tonight, Bill, we infiltrate the birthday party, locate the presents, and leave with every box before anyone notices our unusual enthusiasm for wrapping paper bows.” Molly spread her map across the laundromat table, pinning its corners with quarters and a suspiciously damp cupcake. Bill had found it nearby that morning, apparently. “Every present? Molly, even the enormous one shaped like a pony? My conscience has limits, and apparently they are equine, though not very sturdy either.” Bill examined the map upside down, which was impressive because Molly had drawn it on the back of a pizza menu from last Tuesday night. “The pony is not a present; it is a child’s ride, and stealing it would make us villains in a very specific nursery rhyme forever.” She tapped the garden gate, the side door, and a penciled cloud labeled “possible escape,” then frowned at the cloud’s tiny umbrella drawn by Bill. “Fine. We take only the wrapped gifts. No bikes, no pets, no educational toys, and absolutely no homemade jam from the terrifying aunt in pearls.” Molly nodded, already mentally adopting the pony and naming it Tax Evasion, a name that sounded like bad branding and a tax audit for ponies. “Guests arrive at three. We enter disguised as entertainers, deliver one magic trick, then collect the gifts while everyone applauds our professionalism and excellent timing.” She produced two paper hats from her coat. One said MAGICIAN; the other said LIABILITY, in lettering that looked legally binding under the circumstances, somehow. “I can juggle. I can also apologize to a fruit bowl if the trick becomes emotionally complicated for everyone involved, even when the oranges disapprove.” Bill lifted three oranges; they rolled beneath a dryer. He bowed to the machine, and Molly revised their act to involve fewer spherical objects immediately. “Do not improvise. Last time you improvised, you escaped a museum wearing a bust of Napoleon as a hat and received complimentary parking for it.” Bill adjusted the LIABILITY hat, looking less like a magician than a man whose accountant had predicted this exact afternoon and charged him beforehand anyway. “Then you do the trick. I’ll guard the gifts, and if anyone asks, I’m the responsible adult accompanying the suspicious magician for community service credit.” Molly stared at him, recalling he was thirty-six and his last responsible act involved returning a library book before it became a weapon in court. “Exactly. We’ll do this cleanly, quietly, and with the dignity of professionals who have never been banned from a bounce house for unrelated insurance reasons.” A distant kazoo sounded from the party hall. Both thieves froze, then practiced innocent smiles at the vending machine’s reflection, which looked profoundly unconvinced, too. “On second thought, Molly, perhaps we should donate our services as clowns and leave with exactly zero presents and one balloon each, for morale purposes.” Molly considered the balloon, the map, and Bill’s hat. She approved the plan, provided they stole every present and the cake before anyone noticed them.
created_at
"2026-09-23T09:07:21.134Z"
9/23/2026, 9:07:21 AM
model
"openai/gpt-6-luna-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
27114
moderation_latency
(null)
generation_time
76941
tokens_prompt
75
tokens_completion
2135
native_tokens_prompt
66
native_tokens_completion
8621
native_tokens_completion_images
0
native_tokens_reasoning
7995
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.0043171
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.21.0; linux; x64))"
http_referer
(null)
request_id
"req-1790154441-iMINinJzaGJpynJ0NWKn"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790154441-obzOf9gFI1orwrbsUh3E"
upstream_id
"resp_0724cf8977d12a0c016ab396c9422c87d192373f18e88da0a1"
provider_responses
0
endpoint_id
"05e94e02-b9c2-4bbb-ba55-4082ee9ad687"
id
"resp_0724cf8977d12a0c016ab396c9422c87d192373f18e88da0a1"
is_byok
false
latency
445
model_permaslug
"openai/gpt-6-luna-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.0043171
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
501
99.80%
Dialogue to Total Word Ratio
Ratio: 50.20%, Deviation: 0.20%
neededClean
false
wordsTotal
502
wordsDialogue
252
99.9004%