NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M36P0409WTCPNDWYQ7CXB0G1
Run Details
Model:
GPT-6 Luna (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 200 words with 50% dialogue
Duration:
36.8s
Cost:
$0.001828
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
100.0000%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 50% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly unfolded a map across the laundromat dryer, pinning its corners with stolen cupcakes. The diagram showed the party house, a trampoline. Bill studied it upside down, which he claimed was his specialty. Outside, rain tapped the window like burglars. “The party starts at three, and the presents are stacked beneath the balloon arch. We enter disguised as entertainers, collect every shiny box, and leave before anyone notices,” Molly said. “I can juggle, provided nobody expects the balls to come back. You handle the gift table; I’ll distract the parents with my famous disappearing-cake routine,” Bill replied. Their getaway bicycle leaned against a vending machine, decorated with streamers and a nose. Molly checked her watch; Bill checked the cupcakes. One had vanished. He blamed gravity, though chocolate frosting marked his sleeve. “The children may be small, but their security team is fierce: one grandmother, two toddlers, and a dog wearing a sheriff badge. Stay calm, and nobody gets frosting on the getaway coats,” Molly warned. “What if the dog asks for my name? This fake mustache is shedding quite badly,” Bill asked. At the appointed hour, they marched toward the house, practicing waves and pretending the squeaking bicycle was applause.
created_at
"2026-09-23T08:26:59.984Z"
9/23/2026, 8:26:59 AM
model
"openai/gpt-6-luna-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
4300
moderation_latency
(null)
generation_time
36769
tokens_prompt
75
tokens_completion
1241
native_tokens_prompt
66
native_tokens_completion
3643
native_tokens_completion_images
0
native_tokens_reasoning
3385
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.0018281
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790152019-ISBdROmuuZbVXHYDPGza"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790152019-fI83xnrGHTPJDYs1uoo0"
upstream_id
"resp_018e2572ef84305a016ab38d542b4c87d195e521f3a8ebe379"
provider_responses
0
endpoint_id
"05e94e02-b9c2-4bbb-ba55-4082ee9ad687"
id
"resp_018e2572ef84305a016ab38d542b4c87d195e521f3a8ebe379"
is_byok
false
latency
644
model_permaslug
"openai/gpt-6-luna-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.0018281
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
100.00%
Dialogue to Total Word Ratio
Ratio: 50.25%, Deviation: 0.25%
neededClean
false
wordsTotal
201
wordsDialogue
101
100.0000%