NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M37HKKNW5TYH41Q28N1J43E4
Run Details
Model:
GPT-6 Sol (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
22.2s
Cost:
$0.014112
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
99.92%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
99.9595%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly spread the invitation across the dashboard, pinning it beneath a paper crown stolen from a burger restaurant. At three, the living room would contain twenty children, one exhausted magician, and a mountain of presents. “We enter disguised as entertainers, then quietly collect every present.” Bill considered their getaway vehicle: an ice cream van with no ice cream, a broken horn, and THE FUN BEGINS HERE painted on both sides. “Excellent, except neither of us knows any magic tricks whatsoever.” Molly produced a tablecloth, two plastic wands, and a rabbit. The rabbit produced a larger rabbit without assistance. Bill stared. Neither animal belonged to them. At the house, the birthday girl opened the door before they could knock. She wore a sheriff's badge and held a clipboard. Behind her stood twelve children in sunglasses. Molly lifted a wand. A glittering burst erupted from its tip, followed by a wet popping noise. Every present in the room vanished. For one triumphant second, Molly grinned. Then the birthday girl ticked a box on her clipboard, and both rabbits reappeared inside the van, surrounded by presents. The van drove away. Bill checked his pockets. The rabbits had taken the keys.
created_at
"2026-09-23T16:29:30.18Z"
9/23/2026, 4:29:30 PM
model
"openai/gpt-6-sol-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
5786
moderation_latency
(null)
generation_time
22164
tokens_prompt
75
tokens_completion
897
native_tokens_prompt
66
native_tokens_completion
1398
native_tokens_completion_images
0
native_tokens_reasoning
1143
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.014112
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790180970-lAI9g1w4TiQw6jySxMVk"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790180970-crJ78oKHlsoYPxzHefbG"
upstream_id
"resp_085d0af839c316f3016ab3fe6a4d8087d19a408e697c8e6448"
provider_responses
0
endpoint_id
"1cfc7d9d-4404-4b8e-9ee4-58ae45c9dcd4"
id
"resp_085d0af839c316f3016ab3fe6a4d8087d19a408e697c8e6448"
is_byok
false
latency
616
model_permaslug
"openai/gpt-6-sol-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.014112
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
99.92%
Matches word count
n/a
neededClean
false
words
197
100.00%
Dialogue to Total Word Ratio
Ratio: 10.15%, Deviation: 0.15%
neededClean
false
wordsTotal
197
wordsDialogue
20
99.9595%