NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M36R2MSGCGQTN68TVMPMQ90C
Run Details
Model:
GPT-6 Luna (Reasoning, High)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
33.3s
Cost:
$0.001661
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
100.0000%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
After sunset Molly and Bill crouched behind the hedge outside the community hall, where a banner announced NED’S SIXTH BIRTHDAY in letters so cheerful they seemed to threaten legal action. Through the window they surveyed their target: a mountain of wrapped presents beneath a paper moon. Molly unfolded a blueprint drawn on a napkin; Bill held the wrong end up. “Bill, the bouncy castle has a guard.” “Excellent. I brought stickers and a tiny disguise.” “He looks highly professionally trained.” They waited until the children surged toward the cake, following a magician whose rabbit had escaped into the punch bowl. Bill slipped through the side door in a borrowed clown wig, carrying a sack embroidered with the words DEFINITELY NOT PRESENTS. Molly remained outside, monitoring the guard, who was now licking frosting from his sleeve and regarding the sticker sheet with grave suspicion. Inside, Bill approached the gift pile. He lifted one box, heard a faint rattle, and froze. A handwritten tag promised dinosaur eggs for Ned. He set it down with the tenderness of a bomb technician and chose a safer-looking package shaped like a violin. The guard caught him; children cheered, and Molly traded sack for cake peacefully.
created_at
"2026-09-23T09:03:19.865Z"
9/23/2026, 9:03:19 AM
model
"openai/gpt-6-luna-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
28180
moderation_latency
(null)
generation_time
33262
tokens_prompt
75
tokens_completion
432
native_tokens_prompt
66
native_tokens_completion
3309
native_tokens_completion_images
0
native_tokens_reasoning
3059
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.0016611
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.21.0; linux; x64))"
http_referer
(null)
request_id
"req-1790154199-rsPbytFXSQ6gFSehDALk"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790154199-e5wepgTzZPCOvTGZnij6"
upstream_id
"resp_092bb0a06b2d5284016ab395d7fd8887d1999041f70d243996"
provider_responses
0
endpoint_id
"05e94e02-b9c2-4bbb-ba55-4082ee9ad687"
id
"resp_092bb0a06b2d5284016ab395d7fd8887d1999041f70d243996"
is_byok
false
latency
1562
model_permaslug
"openai/gpt-6-luna-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.0016611
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
100.00%
Dialogue to Total Word Ratio
Ratio: 9.95%, Deviation: 0.05%
neededClean
false
wordsTotal
201
wordsDialogue
20
100.0000%