NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M37HDJ9FRSFVJ3W02A3V4S4W
Run Details
Model:
GPT-6 Sol (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
22.4s
Cost:
$0.010932
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
99.98%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
99.9920%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly spread a crayon map across the café table. Its blue squiggle represented the garden, while a yellow blob marked the children's party. Bill studied the perimeter through binoculars made from two toilet rolls. "We enter as magicians, grab the presents, and vanish." Bill lowered the binoculars. "What if the birthday boy asks us to produce a rabbit?" Molly tapped the blob with a sugar packet. They had no rabbit, but Bill owned a cardigan with exceptionally deep pockets. Inside were three balloons, a bent spoon, and a parking ticket he insisted was lucky. Across the street, children shrieked around a mountain of presents. Their parents watched from folding chairs, armed with cameras and the alertness of people expecting someone else to supervise. The thieves rehearsed their entrance behind a hedge. Molly flourished her cape, which was actually a shower curtain. Bill attempted a bow and dislodged the hedge's resident hedgehog. It rolled down the pavement like an indignant pincushion. At the gate, the birthday boy spotted them immediately. He handed Molly a paper crown and Bill a clipboard. Before either could object, they were assigned to carry the presents indoors, sing loudly, and wash thirty sticky plates.
created_at
"2026-09-23T16:26:12.154Z"
9/23/2026, 4:26:12 PM
model
"openai/gpt-6-sol-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
6450
moderation_latency
(null)
generation_time
22383
tokens_prompt
75
tokens_completion
598
native_tokens_prompt
66
native_tokens_completion
1080
native_tokens_completion_images
0
native_tokens_reasoning
823
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.010932
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790180772-QzZUym7IE7WTQ6ABVwut"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790180772-K13cHg2JzPP7KK2AZexo"
upstream_id
"resp_00596d9c7ac75af4016ab3fda4511487d1a1edeedd438a02e7"
provider_responses
0
endpoint_id
"1cfc7d9d-4404-4b8e-9ee4-58ae45c9dcd4"
id
"resp_00596d9c7ac75af4016ab3fda4511487d1a1edeedd438a02e7"
is_byok
false
latency
522
model_permaslug
"openai/gpt-6-sol-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.010932
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
99.98%
Matches word count
n/a
neededClean
false
words
198
100.00%
Dialogue to Total Word Ratio
Ratio: 10.10%, Deviation: 0.10%
neededClean
false
wordsTotal
198
wordsDialogue
20
99.9920%