NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M36R5W1CRBZDM3SGH2RN5SR0
Run Details
Model:
GPT-6 Luna (Reasoning, High)
Test:
Dialogue tags
Scenario:
Write 200 words with 50% dialogue
Duration:
1m 16s
Cost:
$0.004804
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
100.0000%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 50% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly spread the invitation across a diner table. “We enter through the kitchen when the magician pulls out the rabbit,” Molly said. Bill studied the map upside down, pointing toward the aquarium. “And then we grab every present before the children notice?” Bill asked. Outside, rain rattled windows like burglars. “Precisely. You distract them with your kazoo,” Molly replied. Their target was a party in the community hall, where children would receive presents. “My kazoo sounds like a goose being audited,” Bill said. Molly had packed a sack, disguises, and a decorative conscience. “Excellent. Fear will keep them seated,” Molly insisted. “What if they ask why we’re stealing their gifts?” Bill asked. They whispered over coffee, plotting entry during the magician’s trick. “We say Santa outsourced delivery to two suspicious professionals,” Molly said. “Do professionals usually wear paper crowns?” Bill asked. Neither had considered that children noticed suspicious adults carrying sacks. “Only when undercover. Put yours on backward,” Molly replied. Still Molly drew an arrow toward the door. “I refuse to rob children without benefits,” Bill said. “Fine. You can keep the cake. Bring the empty sack,” Molly said. “Deal, but I’m taking the frosting as hazard pay,” Bill promised.
created_at
"2026-09-23T09:05:05.592Z"
9/23/2026, 9:05:05 AM
model
"openai/gpt-6-luna-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
21657
moderation_latency
(null)
generation_time
75844
tokens_prompt
75
tokens_completion
1818
native_tokens_prompt
66
native_tokens_completion
9594
native_tokens_completion_images
0
native_tokens_reasoning
9322
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.0048036
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.21.0; linux; x64))"
http_referer
(null)
request_id
"req-1790154305-ez1xwtekcHXDUUU3GSEH"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790154305-ebjWey9HIaiywxQdK3jq"
upstream_id
"resp_0f15845c7d973856016ab39641b36c87d18a5cbaacac309d24"
provider_responses
0
endpoint_id
"05e94e02-b9c2-4bbb-ba55-4082ee9ad687"
id
"resp_0f15845c7d973856016ab39641b36c87d18a5cbaacac309d24"
is_byok
false
latency
2374
model_permaslug
"openai/gpt-6-luna-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.0048036
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
100.00%
Dialogue to Total Word Ratio
Ratio: 50.00%, Deviation: 0.00%
neededClean
false
wordsTotal
200
wordsDialogue
100
100.0000%