NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M37HSQH3KGM10WAQWT7PNYXV
Run Details
Model:
GPT-6 Sol
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
6.8s
Cost:
$0.002902
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
86.38%
Matches word count
(details)
99.96%
Dialogue to Total Word Ratio
(details)
93.1723%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly studied the floor plan of Rainbow Castle Fun Zone while Bill sharpened a plastic spoon. Their target, according to the invitation pinned beneath Molly’s elbow, was a mountain of presents beside the birthday throne. Their obstacle was a room full of children, two parents, and a magician with suspiciously alert eyes. “We enter as entertainers,” Molly said. “I refuse to juggle children,” Bill replied. “Balloons, Bill. You juggle balloons.” Bill tried. One balloon escaped, bounced off the ceiling, and settled behind a sofa. Molly added “practice” to the plan. At three o’clock, they arrived wearing capes made from shower curtains. Bill carried a sack disguised as a giant party favor. Molly announced that she could make presents disappear. This earned applause from the children and a long look from the magician. Her first trick involved whisking a wrapped box beneath her cape. Unfortunately, it was bolted to the birthday throne as part of the decorations. The throne lurched, the cape snagged, and Bill dropped his sack. Out rolled twenty plastic spoons. The birthday child gasped. “Are those for the cake?” Bill looked at Molly. Molly looked at the magician. “Yes,” she said. “We brought extras.” The magician handed them plates. For the next hour, they served cake under close supervision.
created_at
"2026-09-23T16:32:50.736Z"
9/23/2026, 4:32:50 PM
model
"openai/gpt-6-sol-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
2435
moderation_latency
(null)
generation_time
6775
tokens_prompt
75
tokens_completion
346
native_tokens_prompt
66
native_tokens_completion
277
native_tokens_completion_images
0
native_tokens_reasoning
0
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.002902
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790181170-8bZi8mt4TcZi4t7SKTv5"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790181170-8usvxAo4gygSooehPeRZ"
upstream_id
"resp_0ab934d951ea3464016ab3ff32e1cc87d1935df0d1a7400fcf"
provider_responses
0
endpoint_id
"1cfc7d9d-4404-4b8e-9ee4-58ae45c9dcd4"
id
"resp_0ab934d951ea3464016ab3ff32e1cc87d1935df0d1a7400fcf"
is_byok
false
latency
804
model_permaslug
"openai/gpt-6-sol-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.002902
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
86.38%
Matches word count
n/a
neededClean
false
words
211
99.96%
Dialogue to Total Word Ratio
Ratio: 11.37%, Deviation: 1.37%
neededClean
false
wordsTotal
211
wordsDialogue
24
93.1723%