NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M36P3TGRY20G9X8Q4WN9JE5F
Run Details
Model:
GPT-6 Luna (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
39.8s
Cost:
$0.001431
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
100.0000%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly and Bill crouched behind a hedge outside the community hall, studying a hand-drawn map that looked suspiciously like a napkin. The party’s paper crowns bobbed beyond the windows; inside, children circled a mountain of presents with the protective intensity of tiny dragons. Bill adjusted his borrowed magician’s cape, which was actually a tablecloth and still smelled faintly of gravy. “Bill, the clown guards presents.” “Excellent; I brought emergency glitter and tap shoes.” “This plan cannot possibly go wrong tonight.” Molly inspected the side entrance. A sign announced NO ADULTS WITHOUT A CHILD, so Bill had brought a teddy bear wearing sunglasses and a fake mustache. Their getaway vehicle, a shopping cart, waited under a blanket nearby, its wheels squeaking like guilty mice. They would slip in during musical chairs, collect every gift, and escape before the birthday cake was cut. The plan relied on timing, stealth, and nobody asking why two adults were carrying a cartful of wrapped toys. Bill practiced a casual smile. It looked less like innocence than a man remembering taxes. A cheer shook the hall. Children raced in sacks. Bill hopped inside, clutching the teddy bear; Molly followed, regretting their plan with complete theatrical dignity.
created_at
"2026-09-23T08:29:01.345Z"
9/23/2026, 8:29:01 AM
model
"openai/gpt-6-luna-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
5610
moderation_latency
(null)
generation_time
39726
tokens_prompt
75
tokens_completion
1114
native_tokens_prompt
66
native_tokens_completion
2849
native_tokens_completion_images
0
native_tokens_reasoning
2588
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.0014311
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790152141-1pKZt8Jw59dM4FhpKt1t"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790152141-PXsxWaJDPLWQQNuddiwl"
upstream_id
"resp_02d99c0f447df843016ab38dcd959087d199e0b2617c81fe07"
provider_responses
0
endpoint_id
"05e94e02-b9c2-4bbb-ba55-4082ee9ad687"
id
"resp_02d99c0f447df843016ab38dcd959087d199e0b2617c81fe07"
is_byok
false
latency
1658
model_permaslug
"openai/gpt-6-luna-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.0014311
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
100.00%
Dialogue to Total Word Ratio
Ratio: 9.95%, Deviation: 0.05%
neededClean
false
wordsTotal
201
wordsDialogue
20
100.0000%