NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M36NK16QMJ6Y04ZFGTZTGN35
Run Details
Model:
GPT-6 Luna (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 200 words with 50% dialogue
Duration:
25.2s
Cost:
$0.001318
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
99.44%
Dialogue to Total Word Ratio
(details)
99.7205%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 50% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Their getaway van idled beside the community hall, decorated with painted stars and a suspiciously large trunk. Through the window, children bounced beneath paper planets while parents arranged cupcakes. Molly adjusted her glittering cape; Bill wore a bear costume whose head turned only when he sneezed. A hand-lettered sign announced “NO GIFTS UNTIL AFTER GAMES,” which Molly considered both schedule and challenge. She checked the sack, lock picks, and Bill’s emergency kazoo. He tested the kazoo. A shrill note escaped, startling a pigeon onto the roof. Inside, the host lifted a microphone. Molly smiled, hearing opportunity. Bill swallowed, hearing “juggle.” “The presents are stacked beneath the balloon arch, Bill. We enter disguised as entertainers, take the gifts, and vanish before anyone notices.” “I refuse to juggle. Last time, I caught a bowling pin with my face, and the circus sent me a sympathy coupon.” “Then be the magician. Say abracadabra, produce smoke, and stuff every wrapped dinosaur into your coat.” “What if a child asks for a rabbit? I can’t steal rabbits; they know my parole officer.” “Fine. We’ll take presents only, leave the cake, and act offended if anyone suspects the two adults carrying sacks.” “That seems reasonable, somehow.”
created_at
"2026-09-23T08:19:51.135Z"
9/23/2026, 8:19:51 AM
model
"openai/gpt-6-luna-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
3787
moderation_latency
(null)
generation_time
25035
tokens_prompt
75
tokens_completion
591
native_tokens_prompt
66
native_tokens_completion
2622
native_tokens_completion_images
0
native_tokens_reasoning
2350
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.0013176
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790151591-gxjQpO0FpyCgQxpXCjtf"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790151591-427P8CZLRvvv4PpaNBgC"
upstream_id
"resp_0393237f6fae4e41016ab38ba73e2487d182830c0c6c71c924"
provider_responses
0
endpoint_id
"05e94e02-b9c2-4bbb-ba55-4082ee9ad687"
id
"resp_0393237f6fae4e41016ab38ba73e2487d182830c0c6c71c924"
is_byok
false
latency
360
model_permaslug
"openai/gpt-6-luna-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.0013176
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
99.44%
Dialogue to Total Word Ratio
Ratio: 52.74%, Deviation: 2.74%
neededClean
false
wordsTotal
201
wordsDialogue
106
99.7205%