NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M36P2PG5S8XV848B92H4Y8NC
Run Details
Model:
GPT-6 Luna (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
21.6s
Cost:
$0.000910
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
36.79%
Dialogue to Total Word Ratio
(details)
68.3935%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
“Molly, the invitation says ‘pirate treasure hunt.’ That’s practically a job description.” “Bill, it says ‘bring a toy to donate.’” “Exactly. We bring toys, then leave with the presents.” “That is not what donate means.” “It is what ‘creative accounting’ means.” “You’re an accountant now?” “Only when the law is nearby.” “Besides, there are dozens of gifts. Nobody will miss one.” “You said all the presents.” “I was being ambitious.” “The cake has a dinosaur on it.” “I respect its security team.” “Molly, the children are arriving.” “Then we should leave.” “Without presents?” “With our dignity.” “We never packed that.” “We packed masks, rope, and a sack.” “The sack says ‘Happy Birthday.’” “I thought it was subtle.” “It has balloons printed on it.” “Camouflage balloons.” “A tiny girl just handed me a cupcake.” “Is that a bribe?” “It’s frosting, Bill.” “I know. I’m trying to classify the evidence.” “She says the presents are for the shelter.” “The children made them themselves.” “Then we can’t steal them.” “Correct.” “We can steal a second cupcake.” “That’s still theft.” “It’s a charitable cupcake.” “You just invented that.” “I am innovating.” “The birthday girl is coming over.” “Quick, look natural.” “How?” “Hold the sack closed.”
created_at
"2026-09-23T08:28:24.458Z"
9/23/2026, 8:28:24 AM
model
"openai/gpt-6-luna-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
4392
moderation_latency
(null)
generation_time
21609
tokens_prompt
75
tokens_completion
1170
native_tokens_prompt
66
native_tokens_completion
1807
native_tokens_completion_images
0
native_tokens_reasoning
1516
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.0009101
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790152104-c0n1jrZ2bjaXP8lNYP9s"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790152104-WNSQSfeGFJZCFrwflmp5"
upstream_id
"resp_007a797f0b1868f9016ab38da89b7887d19a224e59a6083d70"
provider_responses
0
endpoint_id
"05e94e02-b9c2-4bbb-ba55-4082ee9ad687"
id
"resp_007a797f0b1868f9016ab38da89b7887d19a224e59a6083d70"
is_byok
false
latency
743
model_permaslug
"openai/gpt-6-luna-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.0009101
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
201
36.79%
Dialogue to Total Word Ratio
Ratio: 100.00%, Deviation: 10.00%
neededClean
false
wordsTotal
201
wordsDialogue
201
68.3935%