NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M36P1ACBTV80DM9RR167SREB
Run Details
Model:
GPT-6 Luna (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 200 words with 50% dialogue
Duration:
30.1s
Cost:
$0.001439
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
27.01%
Dialogue to Total Word Ratio
(details)
63.5028%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 50% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly studied the party through binoculars while Bill crouched behind a hedge wearing a borrowed birthday hat. The garden glittered with balloons, bunting, and suspiciously competent parents. Between them stood a mountain of presents, guarded only by a sleepy golden retriever and a papier-mâché dragon with googly eyes and dignity. “Which present looks most valuable?” Molly whispered. “The enormous one,” Bill said. “It has a silver bow, and a small country inside.” “We take everything,” she replied. “That’s the plan,” Bill said, “though I’d prefer not to fight the dragon.” “It’s cardboard.” “So was my landlord; he evicted me.” Bill produced a clipboard, three disguises, and a map drawn on a pizza box. Molly examined the route: past the punch bowl, beneath the streamers, then out through the side gate. A toddler in a dinosaur costume suddenly pressed his nose against the window with frosting on his chin, grinning. “Abort?” Bill asked. “No,” said Molly. “We’re professionals.” “Then why is the toddler holding our map?” “Because you gave it to him.” “I thought it was the menu.” The child waved, and Bill waved back. “We should leave,” he said. “Agreed. Empty-handed, but with reviews from the smallest security guard.”
created_at
"2026-09-23T08:27:39.281Z"
9/23/2026, 8:27:39 AM
model
"openai/gpt-6-luna-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
5202
moderation_latency
(null)
generation_time
30073
tokens_prompt
75
tokens_completion
586
native_tokens_prompt
66
native_tokens_completion
2865
native_tokens_completion_images
0
native_tokens_reasoning
2579
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.0014391
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790152059-d2e3bW8Cz0SHKBf51xp7"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790152059-i5mAOpxgOvSiJmsRlDDP"
upstream_id
"resp_08d8ddc57ef0d96a016ab38d7c743087d1b22c3efeb734c366"
provider_responses
0
endpoint_id
"05e94e02-b9c2-4bbb-ba55-4082ee9ad687"
id
"resp_08d8ddc57ef0d96a016ab38d7c743087d1b22c3efeb734c366"
is_byok
false
latency
2009
model_permaslug
"openai/gpt-6-luna-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.0014391
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
199
27.01%
Dialogue to Total Word Ratio
Ratio: 39.30%, Deviation: 10.70%
neededClean
false
wordsTotal
201
wordsDialogue
79
63.5028%