NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M37HZ35RA4ZBGV7VDSEXQ5GP
Run Details
Model:
GPT-6 Sol (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 500 words with 50% dialogue
Duration:
25.5s
Cost:
$0.013252
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
90.48%
Matches word count
(details)
0.15%
Dialogue to Total Word Ratio
(details)
45.3169%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 500 word scene that has 50% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
The community hall smelled of icing, floor polish, and the particular terror of a man wearing a rented mascot costume. From behind a hedge shaped like a disappointed poodle, Molly watched parents carry bright packages through the doors. Bill crouched beside her, studying a crayon map spread over his knees. “Security?” Molly asked. “Three parents, one face painter, and a seven-year-old with a clipboard. The clipboard is our biggest problem.” “That’s absurd. We’re professionals. We once stole an aquarium without spilling a fish.” “He made his father sign a waiver to use the toilet.” Molly leaned over the map. It showed a square marked PRESENTS, a rectangle marked CAKE, and, inexplicably, a dragon. Bill had borrowed it from the hall’s noticeboard, where it had been pinned beneath a reminder to wash hands. Through the window, Molly could see a mountain of wrapped boxes. “Where’s the entrance?” she asked. “Here. We go in as entertainers, gather every present in a sack, and leave during the magic show.” “Who’s doing the magic?” “We are.” “Bill, your last trick ended with a pigeon living in my coat for six weeks.” “He was a loyal assistant.” Inside, a shriek rose above the music. It was followed by applause, suggesting either a successful game or an injury nobody wished to discuss. Molly examined Bill’s disguise: a paper crown, a bow tie, and a vest bearing the words FUN UNCLE. He had no nieces or nephews, a detail he considered irrelevant to the profession. “You look like a man banned from three circuses,” she said. “Two. And the second ban was a misunderstanding about the trapeze.” “What am I supposed to be?” “The birthday princess.” Molly looked down at her black jacket and muddy boots. “The princess has had a difficult winter.” “You’re undercover. Children love a backstory.” The mascot stumbled out of the hall. Its enormous rabbit head turned toward the hedge, then lifted clear of the costume. Underneath was a tired woman with a walkie-talkie and the unmistakable gaze of someone counting both children and exits. Bill folded the map so fast the dragon tore in half. “We could try the kitchen,” he whispered. “The kitchen has a window.” “And a cake. We could hide behind it.” “That cake is smaller than you.” “I can crouch. It’s my strongest skill.” Molly was about to explain that an adult crouching behind a cake remained visible when the clipboard child appeared at the window. He pointed at them, checked something on his list, and held up two fingers. Bill slowly held up two in return. “What does that mean?” Molly asked. “I think we’re expected.” The hall door opened. Beside it stood a handwritten sign: VOLUNTEERS NEEDED TO CARRY PRESENTS TO CHILDREN’S WARD. Molly stared at the mountain of gifts, then at Bill. He tucked the torn map into his vest. “Well,” he said, “we did plan to take all the presents.” “And leave with them?” “Eventually. The hospital has a front door.” Molly stepped out from behind the hedge. “Fine. But you’re explaining the dragon.”
created_at
"2026-09-23T16:35:46.499Z"
9/23/2026, 4:35:46 PM
model
"openai/gpt-6-sol-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
13432
moderation_latency
(null)
generation_time
25436
tokens_prompt
75
tokens_completion
937
native_tokens_prompt
66
native_tokens_completion
1312
native_tokens_completion_images
0
native_tokens_reasoning
644
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.013252
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790181346-xS4kWsutjzwkaWi4Xrjc"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790181346-5kRoyojmmSC9iJSr7bZs"
upstream_id
"resp_0188b611978e8021016ab3ffe2b14487d19fd51baf90dd007d"
provider_responses
0
endpoint_id
"1cfc7d9d-4404-4b8e-9ee4-58ae45c9dcd4"
id
"resp_0188b611978e8021016ab3ffe2b14487d19fd51baf90dd007d"
is_byok
false
latency
1267
model_permaslug
"openai/gpt-6-sol-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.013252
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
90.48%
Matches word count
n/a
neededClean
false
words
510
0.15%
Dialogue to Total Word Ratio
Ratio: 38.60%, Deviation: 11.40%
neededClean
false
wordsTotal
513
wordsDialogue
198
45.3169%