NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M37HGT5X0XNAHZQJSVNT2YFK
Run Details
Model:
GPT-6 Sol (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
23.8s
Cost:
$0.016172
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
36.79%
Dialogue to Total Word Ratio
(details)
68.3940%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
“Bill, remind me why we’re robbing a children’s party.” “Because grown-ups label their gifts. Children just scream and point.” “And the security?” “One grandmother, two parents, and a magician who keeps losing his rabbit.” “What’s our disguise?” “You wear the dinosaur costume. I’ll arrive as the cake inspector.” “There’s an emergency cake inspector?” “There will be if I speak confidently and carry this clipboard.” “Fine. When do we take the presents?” “During musical chairs. I’ll cut the music, everybody panics, and you wheel out the loot.” “Children don’t panic when music stops, Bill. They sit down.” “Exactly. Seated witnesses.” “What about the grandmother?” “I’ve prepared a distraction: a coupon for twenty percent off yarn.” “That’s insulting. She knitted my wedding suit.” “Then she’ll recognize you.” “Not beneath a dinosaur head.” “She recognized you through a snowman costume.” “Because I said hello, Grandma.” “Then don’t speak, Molly.” “Bill, that rabbit is following us.” “Ignore it.” “It has your clipboard.” “All right. New plan: we chase the rabbit, recover the clipboard, and postpone the presents.” “Why?” “Because he’s stolen them.” “The rabbit?” “Look at his getaway wagon.” “He’s better at this than we are.” “Obviously. He has a magician for an accomplice.”
created_at
"2026-09-23T16:27:58.534Z"
9/23/2026, 4:27:58 PM
model
"openai/gpt-6-sol-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
21797
moderation_latency
(null)
generation_time
23769
tokens_prompt
75
tokens_completion
398
native_tokens_prompt
66
native_tokens_completion
1604
native_tokens_completion_images
0
native_tokens_reasoning
1311
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.016172
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790180878-XuHhTWXL2tyOj84Efy8j"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790180878-ntaHE1NzahkkAWvw4wy0"
upstream_id
"resp_0dc8ee3cfebee902016ab3fe0ea17087d1aa6ba9e02dd56e33"
provider_responses
0
endpoint_id
"1cfc7d9d-4404-4b8e-9ee4-58ae45c9dcd4"
id
"resp_0dc8ee3cfebee902016ab3fe0ea17087d1aa6ba9e02dd56e33"
is_byok
false
latency
810
model_permaslug
"openai/gpt-6-sol-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.016172
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
36.79%
Dialogue to Total Word Ratio
Ratio: 100.00%, Deviation: 10.00%
neededClean
false
wordsTotal
201
wordsDialogue
201
68.3940%