NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M36R23HK5TMTRYCMJ06AFESY
Run Details
Model:
GPT-6 Luna (Reasoning, High)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
26.1s
Cost:
$0.001265
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
100.0000%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly studied the invitation through binoculars, though the party was happening next door and the window was open. Bill adjusted his black sweater, which had tiny yellow ducks because the laundry had betrayed him. “Midnight, we enter, lift every present, and vanish before anyone notices.” “Excellent. If questioned, we’re the emergency wrapping committee, obviously.” They crouched behind a hedge while children sang the birthday song with the confidence of a fire alarm. Molly’s plan involved a borrowed cake trolley, a decoy kazoo, and absolutely no physical effort. Bill peeked over the fence. Inside, grandparents guarded the gift pile like seasoned palace sentries, while a toddler wore a paper crown and wielded a sticky spoon. Molly lowered binoculars, conceding the toddler’s tactical superiority. Then the party host emerged carrying a cardboard box marked DONATIONS. Molly’s gaze followed it to a sign: gifts would go to the children’s hospital. The thieves exchanged a look, then quietly rolled the empty trolley toward the cake table instead. Bill’s sweater snagged on the hedge, leaving one yellow duck behind like a tiny calling card. Beneath the moonlit garden fence, they escaped with two slices of cake, a kazoo, and a sudden, inconvenient respect for toddlers.
created_at
"2026-09-23T09:03:02.209Z"
9/23/2026, 9:03:02 AM
model
"openai/gpt-6-luna-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
9542
moderation_latency
(null)
generation_time
26006
tokens_prompt
75
tokens_completion
608
native_tokens_prompt
66
native_tokens_completion
2516
native_tokens_completion_images
0
native_tokens_reasoning
2251
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.0012646
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.21.0; linux; x64))"
http_referer
(null)
request_id
"req-1790154182-YYIDbJ32kGJMjjAGyExM"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790154182-yDGKZbcoNEJ3xHj60Pyo"
upstream_id
"resp_0b4eec47ce97ab51016ab395c65c9487d1a5a8710bad86b67e"
provider_responses
0
endpoint_id
"05e94e02-b9c2-4bbb-ba55-4082ee9ad687"
id
"resp_0b4eec47ce97ab51016ab395c65c9487d1a5a8710bad86b67e"
is_byok
false
latency
1372
model_permaslug
"openai/gpt-6-luna-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.0012646
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
100.00%
Dialogue to Total Word Ratio
Ratio: 10.00%, Deviation: 0.00%
neededClean
false
wordsTotal
200
wordsDialogue
20
100.0000%