NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-10-01M25DD4XSS4AG9JSY4JNW6KCK
Run Details
Model:
DeepSeek V4.1 Flash (Reasoning, High)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
17.1s
Cost:
$0.002728
Provider:
DeepSeek
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
100.0000%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly crouched behind an inflatable castle, adjusting her tiny clown wig and checking the party map. Bill, dressed as a magician, studied the gift table through plastic binoculars. Inside, a swarm of five-year-olds sang off-key, fueled by sugar and vengeance. "Bill, the clown suit chafes." "Shut up, Molly." "Balloons are tripwires." "Presents first." "Toddler armed with juice box." "Keep grinning." Bill lowered the binoculars. The plan was simple. When the magician began his boring rabbit trick, Molly would cut the lights, Bill would sweep every wrapped box into a laundry sack, and they would escape through the bouncy castle’s emergency flap. The getaway tricycle waited by the curb, its streamers fluttering. Molly worried about fingerprints, but Bill insisted glitter covered everything. A small girl in a tiara asked if Bill could make her brother disappear. He considered it. Molly shook her head. They had a schedule, a sack, and no moral compass. The DJ announced cake, creating perfect chaos. The cake distraction would last three minutes, tops, if the clown stopped crying. They moved toward the presents, careful not to step on LEGO or sleeping toddlers. This was either the softest score of their careers or the most humiliating arrest.
created_at
"2026-09-10T10:21:53.469Z"
9/10/2026, 10:21:53 AM
model
"deepseek/deepseek-v4.1-flash-20260910"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
765
moderation_latency
(null)
generation_time
17033
tokens_prompt
75
tokens_completion
4183
native_tokens_prompt
93
native_tokens_completion
4523
native_tokens_completion_images
(null)
native_tokens_reasoning
4245
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"stop"
service_tier
(null)
usage
0.00272775
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1789035713-bBH9OMNnaO1yCyJTR5mF"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1789035713-QytbSFz6CWwONpoQi64M"
upstream_id
"577e1dfb-a516-46b5-9e95-54f63ca59839"
provider_responses
0
endpoint_id
"fd104bc8-535a-4d31-a847-a8e402e86a02"
id
"577e1dfb-a516-46b5-9e95-54f63ca59839"
is_byok
false
latency
420
model_permaslug
"deepseek/deepseek-v4.1-flash-20260910"
provider_name
"DeepSeek"
status
200
total_cost
0.00272775
cache_discount
(null)
upstream_inference_cost
0
provider_name
"DeepSeek"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
100.00%
Dialogue to Total Word Ratio
Ratio: 9.85%, Deviation: 0.15%
neededClean
false
wordsTotal
203
wordsDialogue
20
100.0000%