NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-10-01M25DE89FH8PTKKQZ4MF97AKT
Run Details
Model:
DeepSeek V4.1 Flash (Reasoning, High)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
47.9s
Cost:
$0.006373
Provider:
DeepInfra
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
36.79%
Dialogue to Total Word Ratio
(details)
68.3940%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
"Bill, the clown suit is non-negotiable." "It smells like a wet hamster, Molly." "That's method acting." "We enter as balloon artists?" "We enter as backup magician and assistant, then sneak upstairs." "Upstairs? The presents are in the bouncy castle, Molly." "Since when, Bill?" "Since I saw a piƱata shaped like a safe." "Fine. You distract the kids with a fake rabbit." "You distract them, Molly. I'll crawl through the ball pit." "The ball pit is a sensory hazard." "Exactly. No one suspects a crying clown in a ball pit." "Then we grab every gift and exit through the petting zoo." "The petting zoo has a llama, Bill." "Llama is the getaway driver." "Bill, that's insane." "Insane is our brand. Do we wrap the presents after?" "No, we re-gift them at another party. It's recycling." "Genius. What about cake?" "We steal that too. It's evidence." "Wait, what if the parents recognize us?" "From the last baby shower? We wore masks." "Dinosaur masks. Very distinctive." "We'll wear different dinosaurs." "Velociraptors?" "Triceratops. Less suspicious." "Good. And the cake?" "We take the cake, the goodie bags, and the llama's dignity." "Bill, you're a monster." "A monster with a plan. Let's steal some birthday joy."
created_at
"2026-09-10T10:22:29.685Z"
9/10/2026, 10:22:29 AM
model
"deepseek/deepseek-v4.1-flash-20260910"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
187
moderation_latency
(null)
generation_time
47914
tokens_prompt
75
tokens_completion
4851
native_tokens_prompt
93
native_tokens_completion
5288
native_tokens_completion_images
(null)
native_tokens_reasoning
4979
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"stop"
service_tier
(null)
usage
0.0063735
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1789035749-zKNZrEHOeelM4Z7f9h4G"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1789035749-8C4MbOBOf1vIUPN6HbBe"
upstream_id
"chatcmpl-RNmdPezRSOz5LdC0LZOt9yYm"
provider_responses
0
endpoint_id
"519706ce-5bef-4de6-a755-71cc3c3ac632"
id
"chatcmpl-RNmdPezRSOz5LdC0LZOt9yYm"
is_byok
false
latency
75
model_permaslug
"deepseek/deepseek-v4.1-flash-20260910"
provider_name
"DeepInfra"
status
200
total_cost
0.0063735
cache_discount
(null)
upstream_inference_cost
0
provider_name
"DeepInfra"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
36.79%
Dialogue to Total Word Ratio
Ratio: 100.00%, Deviation: 10.00%
neededClean
false
wordsTotal
201
wordsDialogue
201
68.3940%