NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M37HGS0V6N56BJ21635QKTXQ
Run Details
Model:
GPT-6 Sol (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 200 words with 50% dialogue
Duration:
20.8s
Cost:
$0.013252
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
99.98%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
99.9905%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 50% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly spread a crayon map across the diner table carefully. “We enter disguised as magicians, then vanish with the presents.” Bill leaned closer, scattering napkins like startled white pigeons everywhere. “Can you actually do magic, or just wear suspicious capes?” Outside, a clown wrestled balloons beside the party entrance noisily. “I can make a balloon poodle look deeply disappointed, Bill.” Molly circled the snack table with a stolen pencil stub. “The cake arrives at three; everyone watches the candles then.” Bill examined his cape, which smelled faintly of pickles today. “What if the birthday child notices our enormous gift sack?” A waitress brought fries and regarded their blueprint with alarm. “Tell them it’s a surprise, because technically it certainly is.” Molly grinned, revealing glitter stuck stubbornly between her teeth again. “We’ll leave replacement gifts: our business cards and one kazoo.” Bill drew an escape route through the bouncy castle carefully. “The surprise being that their presents have entirely disappeared, yes?” Their waitress quietly turned the map around; it was upside-down. “Our business cards say professional burglars, Molly. That’s terrible marketing.” The birthday child appeared behind them, wearing a police badge. Bill whispered, “That sounds less magical and more like my childhood, Molly.”
created_at
"2026-09-23T16:27:57.348Z"
9/23/2026, 4:27:57 PM
model
"openai/gpt-6-sol-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
17484
moderation_latency
(null)
generation_time
20739
tokens_prompt
75
tokens_completion
360
native_tokens_prompt
66
native_tokens_completion
1312
native_tokens_completion_images
0
native_tokens_reasoning
1034
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.013252
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790180877-FlFP8DOh7zZF7KQ6vV5H"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790180877-8AcP2CyG3Qdk13kv9L1B"
upstream_id
"resp_0d99156956980b66016ab3fe0df63887d197e1ce703a7ceb2e"
provider_responses
0
endpoint_id
"1cfc7d9d-4404-4b8e-9ee4-58ae45c9dcd4"
id
"resp_0d99156956980b66016ab3fe0df63887d197e1ce703a7ceb2e"
is_byok
false
latency
1557
model_permaslug
"openai/gpt-6-sol-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.013252
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
99.98%
Matches word count
n/a
neededClean
false
words
202
100.00%
Dialogue to Total Word Ratio
Ratio: 49.26%, Deviation: 0.74%
neededClean
false
wordsTotal
203
wordsDialogue
100
99.9905%