NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M36NTQV9PSRD5C5AS0GHF9YX
Run Details
Model:
GPT-6 Luna (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 500 words with 30% dialogue
Duration:
40.6s
Cost:
$0.001516
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
97.63%
Matches word count
(details)
48.73%
Dialogue to Total Word Ratio
(details)
73.1790%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 500 word scene that has 30% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
At eleven o’clock on Saturday night, Molly and Bill crouched behind the hedge outside the Rainbow Room, where a children’s party had been underway since noon. Through the windows, paper stars trembled above a forest of balloons. A mountain of wrapped presents occupied the stage. Bill pressed his nose to the glass and left a perfect foggy circle, like a guilty moon. “Why are we doing this?” Molly whispered. “Unattended presents,” Bill said. “And I lost a bet.” “Those aren’t the same reason.” “They are when you bet a fortune teller.” “She predicted you’d become a better person.” “She has a terrible record.” “These are gifts for children.” “Small, portable, wrapped.” “Your moral compass points toward jail.” “It’s antique. It sticks.” Molly unfolded a napkin map of the building, drawn in ketchup during dinner. It showed the front door, the kitchen, and an enormous question mark labeled “security.” Bill had added a tiny sketch of himself wearing a crown. The plan, such as it was, involved walking in confidently and claiming to be entertainers. Their costumes consisted of one magician’s cape and a foam nose that smelled faintly of cheese. “Who’s the entertainer?” Molly asked. “Me. Cape.” “And foam nose?” “An entertainer needs range.” “You look like a magician who lost a fight with a tomato.” “Then be my assistant.” “I’d rather be the emergency contact.” “We’re collecting presents for charity.” “What charity?” “The Bill Needs Rent Foundation.” “Fraudulent.” “It has a website.” “On a napkin?” “Napkins take ink.” At that moment, a door swung open and a small boy stepped into the corridor, carrying a cupcake with the grave concentration of a bomb-disposal expert. He looked at the hedge. Bill froze in his cape; Molly gave a cheerful wave. The boy waved back, then offered them the cupcake. “For the performers,” he said. Bill accepted it, visibly defeated by frosting. “Are you here to take the presents?” the boy asked. Molly blinked. “We were considering it.” “Don’t. I made one for my grandma. It’s a rock with googly eyes.” Bill glanced at the pile inside. “That’s probably the best one.” “It is.” Molly folded the napkin map. “We’re leaving.” “Can you perform first?” Bill sighed. “I can make the thieves disappear.” Inside, the children gathered as Bill produced a coin from behind his own ear, then pretended to be astonished. Molly volunteered to juggle three balloons and managed two before one escaped toward the ceiling. Their audience applauded the performance, the frosting, and, most loudly, the moment Bill’s foam nose sprang off and landed in the punch. By midnight, Molly and Bill were carrying the gifts—not away, but from the stage to the coatroom, where parents could find them. The boy’s googly-eyed rock rode in Bill’s pocket as payment for the cupcake. Molly handed Bill the last present, a box that rattled ominously. “Next time, we rob a bank.” “Only if it has a children’s party,” Bill said. Then they walked home, two thieves with empty hands and a suspiciously generous collection of party hats.
created_at
"2026-09-23T08:24:03.695Z"
9/23/2026, 8:24:03 AM
model
"openai/gpt-6-luna-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
7383
moderation_latency
(null)
generation_time
40598
tokens_prompt
75
tokens_completion
1085
native_tokens_prompt
66
native_tokens_completion
3019
native_tokens_completion_images
0
native_tokens_reasoning
2338
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.0015161
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790151843-IVn6hseEUhIvS3DF9qaR"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790151843-L8l5fxZHOvYohn3jHnML"
upstream_id
"resp_0c85f15663169289016ab38ca3cb2887d1af42b7b19a0953e4"
provider_responses
0
endpoint_id
"05e94e02-b9c2-4bbb-ba55-4082ee9ad687"
id
"resp_0c85f15663169289016ab38ca3cb2887d1af42b7b19a0953e4"
is_byok
false
latency
346
model_permaslug
"openai/gpt-6-luna-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.0015161
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
97.63%
Matches word count
n/a
neededClean
false
words
507
48.73%
Dialogue to Total Word Ratio
Ratio: 33.79%, Deviation: 3.79%
neededClean
false
wordsTotal
509
wordsDialogue
172
73.1790%