NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M37HQH6HCR6JM31SNGKSNJB5
Run Details
Model:
GPT-6 Sol (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
20.6s
Cost:
$0.011462
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
100.0000%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly studied the invitation as if it were a map to buried treasure. The treasure: twenty-three wrapped presents beneath a smiling giraffe. Across the kitchen table, Bill examined a clown costume like a condemned man. “The cake is our diversion.” He lifted a sequined sleeve. “I refuse to wear glitter.” Outside, balloons bobbed over the neighboring garden fence. An army of children shrieked with the wild authority of people who had never paid rent. Molly had spent all morning observing the festivities through binoculars, which Bill maintained was unnecessary because the party was next door. “Then be the grumpy clown.” “I already brought the nose.” He produced a red foam bulb from his pocket. It squeaked. Both thieves froze. Through the window, the birthday girl looked up, waved, and pointed at their kitchen. Molly waved back. Bill ducked behind the table, red nose glowing above the edge like a distress beacon. The doorbell rang. On the step stood the birthday girl, holding two paper crowns. Behind her waited six friends and a woman with a camera. By sunset, Molly and Bill had stolen nothing. They had, however, been hired for three more parties, and Bill’s clown name was Mister Unfortunate.
created_at
"2026-09-23T16:31:38.716Z"
9/23/2026, 4:31:38 PM
model
"openai/gpt-6-sol-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
6784
moderation_latency
(null)
generation_time
20522
tokens_prompt
75
tokens_completion
583
native_tokens_prompt
66
native_tokens_completion
1133
native_tokens_completion_images
0
native_tokens_reasoning
875
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.011462
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790181098-Njr9CisZb0IihMuV7TgP"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790181098-1m3Z4vYOqJlZWInuTN59"
upstream_id
"resp_0b116b9d7f36c55a016ab3feeae1e887d1ad0b2f6cd1f4fe27"
provider_responses
0
endpoint_id
"1cfc7d9d-4404-4b8e-9ee4-58ae45c9dcd4"
id
"resp_0b116b9d7f36c55a016ab3feeae1e887d1ad0b2f6cd1f4fe27"
is_byok
false
latency
775
model_permaslug
"openai/gpt-6-sol-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.011462
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
100.00%
Dialogue to Total Word Ratio
Ratio: 9.95%, Deviation: 0.05%
neededClean
false
wordsTotal
201
wordsDialogue
20
100.0000%