NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M36RETN8XS6FSB759TE276T2
Run Details
Model:
GPT-6 Luna (Reasoning, High)
Test:
Dialogue tags
Scenario:
Write 500 words with 50% dialogue
Duration:
1m 28s
Cost:
$0.005452
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
99.95%
Dialogue to Total Word Ratio
(details)
99.9746%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 500 word scene that has 50% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly unfolded the invitation on a café table, smoothing a frosting stain as though it were a map to buried treasure and trouble. “The presents sit beneath a balloon arch. We stroll in, lift the table, and vanish with dignity before anyone notices, or our excellent timing fails,” Molly said. Bill peered over his coffee, where a tiny plastic pirate floated face-down, apparently regretting his association with them and their business plan already. “I object to tasteful. Last time you chose disguises, I spent Tuesday dressed as a municipal duck with a permit and emotional baggage—and feathers,” Bill said. Molly glanced toward the party hall, where children sang tunelessly with the confidence of a poorly rehearsed choir, audible through the wall again. “A duck is memorable. We need forgettable. I propose party hats, cheerful smiles, and a convincing story about being distant relatives of the cake maker,” Molly said. Bill considered this, then examined his own reflection in a spoon; it looked like a man whose distant relatives had changed their numbers. “What if the children ask our names? I can’t say ‘Bill and Molly, professional present enthusiasts.’ That sounds suspiciously accurate and painfully bad branding, too,” Bill said. Molly opened her notebook to a page headed Operation Tinsel, decorated with stars, a cupcake, and a regrettable sketch of Bill wearing antlers. “We’re not stealing anything yet. We’re discussing a snack-related misunderstanding involving shiny packages and several ethical blind spots in a theoretical room, for now,” Molly said. A squeak came from the notebook’s spine. The toy pirate, rescued from coffee, was wedged there like a tiny informant, full of secrets. “The pirate is judging us. I recognize that look; my mother made it whenever I described my career as temporary and not especially lucrative, either,” Bill said. Molly stared at the party room’s open doorway. A child in a paper crown guarded the gifts with a seriousness reserved for monuments. “That child has better security instincts than we do. Also, he’s holding a juice box like it contains classified government secrets and our future together,” Molly said. Bill tugged his jacket closed, revealing a party invitation pinned inside; he had dressed for the occasion and forgotten it completely until now. “Maybe we should return the invitation and leave. I hear morality is popular at parties, especially among people who own the presents and invitation lists,” Bill said. A burst of laughter rolled from the hall. The child offered them a cupcake through the doorway, disarming both criminals with sweet sprinkles. “Fine. We’ll return the invitation, accept one cupcake, and make a generous donation to the birthday child’s present pile, under an adult’s watchful eye nearby,” Molly said. Bill carried the cupcake carefully, as if it were evidence. Molly followed, planning a career involving legal snacks and honest bookkeeping for children. “If anyone asks, we’re the entertainment. Our act is called Two Adults Making One Sensible Decision, and it has a thrilling ending with no getaway,” Bill said.
created_at
"2026-09-23T09:09:59.091Z"
9/23/2026, 9:09:59 AM
model
"openai/gpt-6-luna-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
4181
moderation_latency
(null)
generation_time
87669
tokens_prompt
75
tokens_completion
2247
native_tokens_prompt
66
native_tokens_completion
10890
native_tokens_completion_images
0
native_tokens_reasoning
10265
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.0054516
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.21.0; linux; x64))"
http_referer
(null)
request_id
"req-1790154599-RXnTW2k9ZPclADBb76Hb"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790154599-6QgYzT5Sz1TfdirU5Sbt"
upstream_id
"resp_055f5457fcfe8761016ab397673d1487d199b714f32f3863b5"
provider_responses
0
endpoint_id
"05e94e02-b9c2-4bbb-ba55-4082ee9ad687"
id
"resp_055f5457fcfe8761016ab397673d1487d199b714f32f3863b5"
is_byok
false
latency
775
model_permaslug
"openai/gpt-6-luna-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.0054516
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
499
99.95%
Dialogue to Total Word Ratio
Ratio: 49.90%, Deviation: 0.10%
neededClean
false
wordsTotal
501
wordsDialogue
250
99.9746%