NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-10-05-01M46BGEHH461NSA52SWDP4R63
Run Details
Model:
GPT-6.1 Sol (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
30.0s
Cost:
$0.010692
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
99.9995%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly spread a birthday invitation across the table. Bill eyed the glitter like it might testify against them in court. “Right, Bill. We enter as entertainers, locate the presents, and leave before anybody notices we are terrible entertainers.” “I can juggle.” “You can drop three oranges while screaming. That is different.” “Children appreciate realism.” “We need disguises.” “I brought a dinosaur costume.” “That is a bathrobe with forks taped down the back.” “Budget dinosaur.” “Molly, are we really stealing every present?” “Every last one.” “Even the educational ones?” “We have principles, Bill. No child deserves fractions on their birthday.” “What about the pony?” “There is a pony?” “According to the invitation.” “That complicates the getaway bag.” “Also, the birthday girl is my niece.” “Your niece?” “She asked for burglars. Apparently pirates were booked.” “So we are hired?” “Technically. Payment is cake.” “And the presents?” “Props. We steal them, she catches us, everyone cheers.” “Molly, why are you crying?” “I spent six weeks learning balloon animals for an actual crime.” “You learned one.” “It is a snake.” “It is an uninflated balloon.” “A sleeping snake.” “Come on. Our criminal empire awaits.” “Fine. But I am stealing extra frosting.” “Obviously, boss.”
created_at
"2026-10-05T15:39:22.552Z"
10/5/2026, 3:39:22 PM
model
"openai/gpt-6.1-sol-20260929"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
10337
moderation_latency
(null)
generation_time
29915
tokens_prompt
75
tokens_completion
794
native_tokens_prompt
66
native_tokens_completion
1056
native_tokens_completion_images
0
native_tokens_reasoning
775
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.010692
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.21.0; linux; x64))"
http_referer
(null)
request_id
"req-1791214762-EZWdXcS74uUsVTEaRD7Z"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1791214762-Ehtumm51IQQjOuErOMWh"
upstream_id
"resp_02f17419ef5ac39a016ac3c4aaa98487d1a03ca39d093d4833"
provider_responses
0
endpoint_id
"dfb9d5c3-63f9-4263-8b78-4fa97dcb3d5e"
id
"resp_02f17419ef5ac39a016ac3c4aaa98487d1a03ca39d093d4833"
is_byok
false
latency
352
model_permaslug
"openai/gpt-6.1-sol-20260929"
provider_name
"OpenAI"
status
200
total_cost
0.010692
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
199
100.00%
Dialogue to Total Word Ratio
Ratio: 89.95%, Deviation: 0.05%
neededClean
false
wordsTotal
199
wordsDialogue
179
99.9995%