NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M37HFG3AKVPDP1AYWP0FF1ZT
Run Details
Model:
GPT-6 Sol
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
7.5s
Cost:
$0.002792
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
87.84%
Dialogue to Total Word Ratio
(details)
93.9218%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly spread the birthday invitation across the café table like a map of a fortified castle. A cartoon dinosaur grinned beneath the words: SATURDAY, TWO O’CLOCK, BRING YOUR OWN SOCKS. “Why socks?” Bill asked. “Bounce house,” Molly said. “Stay focused.” Bill studied the guest list Molly had copied from a parent’s group chat. Twelve children, fourteen adults, one magician, and an uncle described only as “competitive.” The presents would sit beneath a crepe-paper arch in the living room. Between them and the front door lay a kitchen full of parents making small talk. Bill produced two tiny party hats from his coat. One had glittery stars; the other said BIRTHDAY GIRL in pink sequins. “Disguises,” he whispered. Molly considered the hats, then Bill’s beard, then the invitation again. They had planned to pose as entertainers, but neither could make balloon animals. Last week Bill had tried to twist a dachshund and produced something that frightened his landlord’s terrier. She folded the invitation and slid it back across the table. The party was at her niece’s house. She knew exactly where the presents would be. What she did not know was how to explain Bill wearing the birthday girl’s hat.
created_at
"2026-09-23T16:27:15.442Z"
9/23/2026, 4:27:15 PM
model
"openai/gpt-6-sol-20260922"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
3194
moderation_latency
(null)
generation_time
7410
tokens_prompt
75
tokens_completion
319
native_tokens_prompt
66
native_tokens_completion
266
native_tokens_completion_images
0
native_tokens_reasoning
0
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.002792
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790180835-yrpdGI4p9AM86RJH8rDH"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790180835-qvy5lDPzrx8DBWjE5Zj7"
upstream_id
"resp_0b3a84a7c5e18f3d016ab3fde3901487d1b2a0bdb39163469c"
provider_responses
0
endpoint_id
"1cfc7d9d-4404-4b8e-9ee4-58ae45c9dcd4"
id
"resp_0b3a84a7c5e18f3d016ab3fde3901487d1b2a0bdb39163469c"
is_byok
false
latency
988
model_permaslug
"openai/gpt-6-sol-20260922"
provider_name
"OpenAI"
status
200
total_cost
0.002792
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
199
87.84%
Dialogue to Total Word Ratio
Ratio: 4.00%, Deviation: 6.00%
neededClean
false
wordsTotal
200
wordsDialogue
8
93.9218%