NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-03-01M1KCZC0BZ7HH2QVJGVHVN36P
Run Details
Model:
DeepSeek V4 Pro 0813 (Reasoning, High)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
53.0s
Cost:
$0.016073
Provider:
Phala
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
100.0000%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly spread a crudely drawn map of the bouncy castle across the sticky picnic table. Bill leaned in, his burglar mask already smelling faintly of cupcake frosting. A crayon X marked the present pile near the pinata. "Security's a clown," Molly whispered. "He's paid in balloon animals." Bill nodded, producing a tiny squirt gun filled with chocolate milk. "So we strike during musical statues." Molly grinned. "Pin the tail on the guard." Their ridiculous plan involved a stolen ice cream van, a decoy magician, and exactly forty-seven whoopee cushions. The birthday boy's mother would be distracted by a staged tantrum over the last slice of cake. Bill practiced his stealth walk among scattered Lego bricks, wincing with every sharp, plastic step. Molly studied the guest list, noting which children were likely to scream loudest. They agreed the bouncy castle was both an entry point and a potentially fatal liability. The real challenge, Bill whispered, was not tripping over a juice box fortress. Escape would require blending into a conga line and exiting through the inflatable slide. The presents would vanish one by one during the loud, sticky chaos of the candy scramble. It was ridiculous, deeply undignified, and absolutely foolproof.
created_at
"2026-09-03T10:28:02.209Z"
9/3/2026, 10:28:02 AM
model
"deepseek/deepseek-v4-pro-20260813"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
1096
moderation_latency
(null)
generation_time
52902
tokens_prompt
75
tokens_completion
3834
native_tokens_prompt
146
native_tokens_completion
3676
native_tokens_completion_images
(null)
native_tokens_reasoning
3409
native_tokens_cached
128
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"stop"
service_tier
(null)
usage
0.01607266
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.19.0; linux; x64))"
http_referer
(null)
request_id
"req-1788431282-CxBf6eUvqj2HcwSM8H0V"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1788431282-Bu698fEgeHxJqJTG9kGd"
upstream_id
"req_6f079ee8e20587b717f057d0f0c0e67e"
provider_responses
0
endpoint_id
"bb1fb528-2000-460c-a65e-b82dc347c019"
id
"req_6f079ee8e20587b717f057d0f0c0e67e"
is_byok
false
latency
398
model_permaslug
"deepseek/deepseek-v4-pro-20260813"
provider_name
"Phala"
status
200
total_cost
0.01607266
cache_discount
0.0001664
upstream_inference_cost
0
provider_name
"Phala"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
100.00%
Dialogue to Total Word Ratio
Ratio: 9.95%, Deviation: 0.05%
neededClean
false
wordsTotal
201
wordsDialogue
20
100.0000%