NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-10-05-01M46B52EE7XD1NY6N3Y1GJX4F
Run Details
Model:
GPT-6.1 Sol (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 500 words with 70% dialogue
Duration:
54.3s
Cost:
$0.025462
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
99.92%
Matches word count
(details)
99.84%
Dialogue to Total Word Ratio
(details)
99.8776%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 500 word scene that has 70% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly spread a birthday invitation across the table. Bill weighted it with a stolen paperweight shaped like a policeman. Outside, rain tapped the window with the discretion of a snitch. “We are not robbing a child. We are intercepting a shipment of luxury goods delivered by emotionally vulnerable adults. That is entirely different, legally speaking, provided nobody asks a lawyer or the birthday child directly.” “Right, but the invitation says Oliver is turning six. What luxury goods does a six year old get? Last year my nephew received a wooden train and spent three hours eating the little cardboard station.” Bill consulted the invitation. Its glitter had migrated onto his fingers, giving him the hands of a magician recently questioned about several unexplained disappearances. Molly confiscated his emergency party whistle. “Oliver’s father owns three yachts. The presents will be expensive. We enter as entertainers, gather everything during the cake ceremony, and leave before anybody notices. Children are famously distracted by fire and concentrated sugar products.” “Entertainers? Molly, my only talent is looking innocent while carrying a suspiciously heavy television. Also, I refuse to dress as a clown. My father was a clown. We still owe money on his funeral shoes.” Molly produced two rabbit costumes from beneath the table. One had a missing ear. The other wore an expression of permanent disappointment remarkably similar to Bill’s own professional resting face. “Nobody said clowns. We are educational rabbits. You discuss woodland habitats while I supervise the gift table. If anyone challenges us, explain that presents frighten rabbits. We remove them for bunny welfare.” “Children ask questions. What if somebody asks where rabbits sleep? Or why rabbits need a van? Or why one rabbit has my face and a tattoo saying DEBORAH FOREVER, with Deborah crossed out three times?” The rain stopped. Somewhere downstairs, a child laughed. Bill flinched. Molly turned the invitation over and discovered another line of print. Her confidence developed a small but distinctly audible leak. “There is a complication. Oliver has requested no presents. Guests are bringing donations for the animal shelter. Apparently he wants every abandoned dog to have a blanket. Honestly, that feels unnecessarily aggressive for someone six.” “So our grand criminal enterprise is stealing blankets from homeless puppies? I have standards, Molly. Low standards, certainly, but actual standards. Besides, dogs remember faces. I cannot spend my retirement being identified by angry dachshunds.” Molly folded the invitation. Bill picked up the rabbit head. Their careers had survived alarms, police, and betrayal, but apparently neither possessed a strategy for confronting a morally superior child. “Fine. We go, perform, and leave the presents alone. Then we steal something respectable afterward, like a banker’s watch. But we need a gift ourselves. You cannot attend a birthday empty handed. That looks suspicious.” “I have the policeman paperweight. Oliver might enjoy it. Unless you think that sends the wrong message. Also, for the record, I am keeping the rabbit costume. Deborah always said I never committed to anything.”
created_at
"2026-10-05T15:33:09.717Z"
10/5/2026, 3:33:09 PM
model
"openai/gpt-6.1-sol-20260929"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
47524
moderation_latency
(null)
generation_time
54263
tokens_prompt
75
tokens_completion
835
native_tokens_prompt
66
native_tokens_completion
2533
native_tokens_completion_images
0
native_tokens_reasoning
1911
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.025462
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.21.0; linux; x64))"
http_referer
(null)
request_id
"req-1791214389-URdYBCCn517q9UTZ6HP4"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1791214389-YRzsCKbI0rpCUGZbUg8T"
upstream_id
"resp_067df8bbf6f0db2f016ac3c335d52487d1999bed081d65dfaf"
provider_responses
0
endpoint_id
"dfb9d5c3-63f9-4263-8b78-4fa97dcb3d5e"
id
"resp_067df8bbf6f0db2f016ac3c335d52487d1999bed081d65dfaf"
is_byok
false
latency
308
model_permaslug
"openai/gpt-6.1-sol-20260929"
provider_name
"OpenAI"
status
200
total_cost
0.025462
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
99.92%
Matches word count
n/a
neededClean
false
words
497
99.84%
Dialogue to Total Word Ratio
Ratio: 69.82%, Deviation: 0.18%
neededClean
false
wordsTotal
497
wordsDialogue
347
99.8776%