NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-10-01M25DYFBD6PENH3PM65W40ASY
Run Details
Model:
DeepSeek V4.1 Flash (Reasoning, High)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
12.3s
Cost:
$0.001792
Provider:
DeepSeek
Evaluation summary
Score
Evaluator
Details Link
99.92%
Matches word count
(details)
36.79%
Dialogue to Total Word Ratio
(details)
68.3535%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
"Bill, we hit the birthday party at three, when the magician does the rabbit thing." "Molly, I still say we go as clowns. Nobody suspects a clown." "Everyone suspects a clown, Bill. That's their whole deal." "Fine. We dress as gift inspectors. Clipboard, hi-vis vest, tragic little hats." "Better. Now, the presents go in the bouncy castle. We deflate it, drag it to the van." "The van with the ice cream jingle?" "No, the van with the puppy rescue stickers. We're not monsters." "Good. What about the kids?" "Bill, we're thieves, not monsters. Distract them with cake." "Whose cake?" "Their cake. We're liberating it too." "That's a war crime, Molly." "It's a party favor. Now, the piƱata. It's full of mini candy and, according to my source, a spare key to the utility closet." "Your source is a six-year-old named Timmy?" "Timmy is reliable. He takes naps for the mob." "Okay. Signal?" "I yell pin the tail on the donkey. You grab the pile." "And if a parent asks?" "We say we're from the birthday police." "Perfect. I'll wear a name tag that says responsible adult." "That's the most terrifying sentence you've ever said." "I practice in the mirror." "Bring the fake mustache."
created_at
"2026-09-10T10:31:21.202Z"
9/10/2026, 10:31:21 AM
model
"deepseek/deepseek-v4.1-flash-20260910"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
560
moderation_latency
(null)
generation_time
12202
tokens_prompt
75
tokens_completion
2641
native_tokens_prompt
93
native_tokens_completion
2963
native_tokens_completion_images
(null)
native_tokens_reasoning
2663
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"stop"
service_tier
(null)
usage
0.00179175
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1789036281-RJQaJHWKQu3wc79up2Ob"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1789036281-wqYnZbR4onI2PYiMJVIZ"
upstream_id
"e35ed80c-527b-4b98-bac2-f09041a06801"
provider_responses
0
endpoint_id
"fd104bc8-535a-4d31-a847-a8e402e86a02"
id
"e35ed80c-527b-4b98-bac2-f09041a06801"
is_byok
false
latency
337
model_permaslug
"deepseek/deepseek-v4.1-flash-20260910"
provider_name
"DeepSeek"
status
200
total_cost
0.00179175
cache_discount
(null)
upstream_inference_cost
0
provider_name
"DeepSeek"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
99.92%
Matches word count
n/a
neededClean
false
words
203
36.79%
Dialogue to Total Word Ratio
Ratio: 100.00%, Deviation: 10.00%
neededClean
false
wordsTotal
206
wordsDialogue
206
68.3535%