NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-10-01M25DP3FKAGNT67GPA64W56PT
Run Details
Model:
DeepSeek V4.1 Flash (Reasoning, High)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
12.0s
Cost:
$0.001526
Provider:
DeepSeek
Evaluation summary
Score
Evaluator
Details Link
99.98%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
99.9920%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly unfolded a crayon-smeared map of the Patterson house across a sticky food-court table. Bill polished a juice box he claimed was a lockpick. Their target was not a vault or a diamond, but the teetering mountain of presents at little Timmy's fifth birthday party. The venue featured a bouncy castle, a face-painting witch, and at least fourteen toddlers with the grip strength of tiny loan sharks. Molly planned to enter as a piƱata safety inspector. Bill would pose as a magician whose only trick was making gifts disappear. The window was narrow: after cake, before parents started filming. The bouncy castle would provide chaos. The party favor table would provide cover. The birthday boy, however, apparently bit. "What if the kids cry?" Bill asked. "Crying is noisy cover," Molly said. "Grab the wrapped ones. Leave the cards." "And the cake?" "Decoy." Bill worried about the llama. Molly worried about the clown. The clown knew Bill from a previous job. That changed everything. They rehearsed their escape through the petting zoo, past a goat with criminal instincts, and into a waiting minivan shaped like a cupcake. It was, Bill admitted, the most terrifying heist they had ever planned.
created_at
"2026-09-10T10:26:46.904Z"
9/10/2026, 10:26:46 AM
model
"deepseek/deepseek-v4.1-flash-20260910"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
732
moderation_latency
(null)
generation_time
11971
tokens_prompt
75
tokens_completion
2434
native_tokens_prompt
93
native_tokens_completion
2520
native_tokens_completion_images
(null)
native_tokens_reasoning
2255
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"stop"
service_tier
(null)
usage
0.00152595
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1789036006-MxNHVVwsXbwp71OKbEAu"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1789036006-vtpwhjXIolvY3d2V7iBl"
upstream_id
"147cd888-a18c-43a2-b2f3-b103dc145720"
provider_responses
0
endpoint_id
"fd104bc8-535a-4d31-a847-a8e402e86a02"
id
"147cd888-a18c-43a2-b2f3-b103dc145720"
is_byok
false
latency
330
model_permaslug
"deepseek/deepseek-v4.1-flash-20260910"
provider_name
"DeepSeek"
status
200
total_cost
0.00152595
cache_discount
(null)
upstream_inference_cost
0
provider_name
"DeepSeek"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
99.98%
Matches word count
n/a
neededClean
false
words
198
100.00%
Dialogue to Total Word Ratio
Ratio: 9.95%, Deviation: 0.05%
neededClean
false
wordsTotal
201
wordsDialogue
20
99.9920%