NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-03-01M1KBH3JECRCJWC865NZXW422
Run Details
Model:
DeepSeek V4 Pro 0813 (Reasoning, High)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
39.2s
Cost:
$0.025043
Provider:
BaseTen
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
100.0000%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly slid the glitter-glue blueprint across the sticky card table, her elbow knocking a half-empty juice box onto the floor. Bill squinted at the crayon arrows marking the birthday party's gift table, which sat perilously close to the pinata and the hired pony. He tapped the bouncy castle with a stolen crayon and a nervous giggle. "We go in as clowns." Molly shook her head and pointed at the security mother sipping sugar-free coffee by the cake. "Too obvious." Bill frowned and adjusted his fake mustache. "Mimes?" Molly shuddered, remembering the silent incident in the ball pit. "Worse." She traced a route past the pinata, through the craft corner, and under the snack table. Bill nodded very slowly, his eyes gleaming. "Balloon animals?" Molly considered it, then drew a tiny poodle on the napkin. "Acceptable." Bill asked, "What about the pony?" Molly said, "No hooves." Bill sighed and stared at the gift table. "Cupcake tower?" Molly nodded. Bill grinned and pocketed a whoopee cushion. They quietly rehearsed the distraction: Bill would inflate a poodle, Molly would crawl under the snack table, and both would stuff all loot into a diaper bag. The plan was flawless except for the juice boxes.
created_at
"2026-09-03T10:02:46.242Z"
9/3/2026, 10:02:46 AM
model
"deepseek/deepseek-v4-pro-20260813"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
353
moderation_latency
(null)
generation_time
39142
tokens_prompt
75
tokens_completion
5648
native_tokens_prompt
147
native_tokens_completion
6275
native_tokens_completion_images
(null)
native_tokens_reasoning
5995
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"stop"
service_tier
(null)
usage
0.02504304
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.19.0; linux; x64))"
http_referer
(null)
request_id
"req-1788429766-cLxnzzmVPO61l44Y12Ba"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1788429766-hHyEEIKZ9lKjEw9etezs"
upstream_id
"chatcmpl-91f152f41e044f958bc5367db39544be"
provider_responses
0
endpoint_id
"85bb5a4a-983d-4d11-b97b-c7e5b348375f"
id
"chatcmpl-91f152f41e044f958bc5367db39544be"
is_byok
false
latency
352
model_permaslug
"deepseek/deepseek-v4-pro-20260813"
provider_name
"BaseTen"
status
200
total_cost
0.02504304
cache_discount
(null)
upstream_inference_cost
0
provider_name
"BaseTen"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
100.00%
Dialogue to Total Word Ratio
Ratio: 9.85%, Deviation: 0.15%
neededClean
false
wordsTotal
203
wordsDialogue
20
100.0000%