NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M36R85BK0EE1XJSRA73M7Q4G
Run Details
Model:
GPT-6 Luna (Reasoning, High)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
53.1s
Cost:
$0.002684
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
99.92%
Matches word count
(details)
73.34%
Dialogue to Total Word Ratio
(details)
86.6307%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
“Molly, the presents are stacked beneath the banner.” “Excellent. Our plan?” “We enter the party, take every gift, and vanish.” “Bill, that’s not a plan. That’s a sentence wearing a tiny burglar mask.” “You can’t juggle.” “I can drop things in a way that suggests juggling.” “Who distracts the parents?” “The cake.” “The cake distracts nobody.” “It’s a spectacular cake. It has a fondant dragon.” “Does the dragon breathe fire?” “Only if someone lights the candles.” “Absolutely not. We are thieves, not arsonists.” “We take them all.” “Even the knitted sweater?” “Especially the knitted sweater. It looks expensive.” “It has a label saying ‘Grandma made this.’” “Then we return it to Grandma.” “Bill, perhaps we should leave the presents alone.” “Because of the sweater?” “Because those kids are watching us through the window.” “Oh. I thought that was a very judgmental balloon.” “It has eyebrows.” “Right. New plan: we bring presents.” “Where did you get those?” “Borrowed from our own apartment.” “Bill, that’s my lamp.” “Wrapped beautifully, though.” “And your socks?” “Gift-bagged, with a bow.” “Let’s go.” “To the party?” “To apologize.” “Should we keep the hats?” “Only if the balloon approves.” Outside, a balloon bobbed ominously.
created_at
"2026-09-23T09:06:20.67Z"
9/23/2026, 9:06:20 AM
model
"openai/gpt-6-luna-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
13232
moderation_latency
(null)
generation_time
53097
tokens_prompt
75
tokens_completion
1351
native_tokens_prompt
66
native_tokens_completion
5354
native_tokens_completion_images
0
native_tokens_reasoning
5053
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.0026836
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.21.0; linux; x64))"
http_referer
(null)
request_id
"req-1790154380-YlopWP73U6lDqxMsYqb3"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790154380-JuGYAFwIBItnjxLLU3qa"
upstream_id
"resp_0223966a193d44f9016ab3968cc62c87d1a2c8e7f0c52cb631"
provider_responses
0
endpoint_id
"05e94e02-b9c2-4bbb-ba55-4082ee9ad687"
id
"resp_0223966a193d44f9016ab3968cc62c87d1a2c8e7f0c52cb631"
is_byok
false
latency
1556
model_permaslug
"openai/gpt-6-luna-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.0026836
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
99.92%
Matches word count
n/a
neededClean
false
words
197
73.34%
Dialogue to Total Word Ratio
Ratio: 97.46%, Deviation: 7.46%
neededClean
false
wordsTotal
197
wordsDialogue
192
86.6307%