NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M37HAEZ31F2SFVQSMDAJJ9GC
Run Details
Model:
GPT-6 Sol (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 200 words with 50% dialogue
Duration:
24.9s
Cost:
$0.020062
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
100.0000%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 50% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Across the street, Molly and Bill watched six shrieking children orbit a mountain of presents beneath a glittering cardboard castle. “We enter disguised as magicians, then make everything disappear quietly.” “Your last magic trick set my trousers on fire. Twice.” Bill tugged his borrowed cape straight. It had once been a tablecloth, and still smelled faintly of gravy and defeat. “Fine. We pose as balloon animals. Nobody suspects balloon animals.” “Nobody invites balloon animals inside either, unless they’re already inflated.” Molly pointed to the invitation taped beside the gate. Apparently, entry required one joke, one costume, and absolutely no parents. “Excellent. I have jokes, and you look permanently unsupervised anyway.” “That’s my face. What’s the plan once we’re actually inside?” A child appeared at the gate, wearing a paper crown and the expression of someone auditing tax returns for fun. “My joke: why did the present cross the road, Bill?” “Because you forgot to check for traffic again, Molly? Seriously?” The child considered this, then silently lifted the gate latch. Behind him, a dozen parents cheered, grateful for incoming entertainment. “We’re in. Grab the presents while I distract everyone magnificently.” “Molly, they’ve hired us. We just burgled ourselves employment instead.”
created_at
"2026-09-23T16:24:30.648Z"
9/23/2026, 4:24:30 PM
model
"openai/gpt-6-sol-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
22677
moderation_latency
(null)
generation_time
24612
tokens_prompt
75
tokens_completion
358
native_tokens_prompt
66
native_tokens_completion
1993
native_tokens_completion_images
0
native_tokens_reasoning
1720
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.020062
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790180670-kfzqVyOx4CiUiX9R6GSk"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790180670-KvAWgMEOep4mBVFISoB0"
upstream_id
"resp_049e767c1ca718b0016ab3fd3ed8f887d19e85e816e6cbc41f"
provider_responses
0
endpoint_id
"1cfc7d9d-4404-4b8e-9ee4-58ae45c9dcd4"
id
"resp_049e767c1ca718b0016ab3fd3ed8f887d19e85e816e6cbc41f"
is_byok
false
latency
427
model_permaslug
"openai/gpt-6-sol-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.020062
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
100.00%
Dialogue to Total Word Ratio
Ratio: 50.00%, Deviation: 0.00%
neededClean
false
wordsTotal
200
wordsDialogue
100
100.0000%