NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-03-01M1KA41QKMKJ9APK8EAPFJT75
Run Details
Model:
DeepSeek V4 Pro 0813 (Reasoning, High)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
10.8s
Cost:
$0.001575
Provider:
NextBit
Evaluation summary
Score
Evaluator
Details Link
90.48%
Matches word count
(details)
36.79%
Dialogue to Total Word Ratio
(details)
63.6358%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
"Bill, we're stealing from a six-year-old's birthday party. That's a new low." "Low? Molly, it's genius. Tiny guests, tiny security." "The bouncy castle isn't a vault. It's inflatable." "Exactly. Everyone's distracted by sugar and pin-the-tail. We waltz in as party clowns." "You can't juggle, Bill." "I can honk a nose. That's ninety percent of clowning." "And what? Stuff a pony into a gift bag?" "No, we grab the pile by the piƱata. I'll create a diversion." "Please don't say you're punching the magician." "I'm releasing a dozen hamsters into the ball pit." "Where do you get a dozen hamsters at this hour?" "I know a guy. You handle the loot." "Me? I'm allergic to glitter and small children." "Wear gloves. And a wig." "This is absurd. The presents are mostly Legos and picture books." "Not the one from Aunt Carol. I saw it. It's heavy. Could be a tablet." "So we're risking jail for one box?" "No, we're risking jail for all the boxes. Think big, Molly." "Fine. But if a child asks me to make a balloon animal, I'm out." "Deal. Now, how do you feel about face paint?"
created_at
"2026-09-03T09:38:09.794Z"
9/3/2026, 9:38:09 AM
model
"deepseek/deepseek-v4-pro-20260813"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
2713
moderation_latency
(null)
generation_time
10772
tokens_prompt
75
tokens_completion
350
native_tokens_prompt
146
native_tokens_completion
349
native_tokens_completion_images
(null)
native_tokens_reasoning
61
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"stop"
service_tier
(null)
usage
0.00157476
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.19.0; linux; x64))"
http_referer
(null)
request_id
"req-1788428289-XLesTSoy6XsNZySwJ5GQ"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1788428289-P2NsBPekX2qfd6YnJ07n"
upstream_id
"chatcmpl-nb3032ad496198417895cc6ff4"
provider_responses
0
endpoint_id
"ca5f0ff9-8764-4809-b5fa-968219bdfc4e"
id
"chatcmpl-nb3032ad496198417895cc6ff4"
is_byok
false
latency
2712
model_permaslug
"deepseek/deepseek-v4-pro-20260813"
provider_name
"NextBit"
status
200
total_cost
0.00157476
cache_discount
(null)
upstream_inference_cost
0
provider_name
"NextBit"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
90.48%
Matches word count
n/a
neededClean
false
words
190
36.79%
Dialogue to Total Word Ratio
Ratio: 100.00%, Deviation: 10.00%
neededClean
false
wordsTotal
194
wordsDialogue
194
63.6358%