NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-10-07-01M4AT2DK5QT5A83RV4STYWVRK
Run Details
Model:
Mistral Large 4.0 (Reasoning)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
53.7s
Cost:
$0.010314
Provider:
Mistral
Evaluation summary
Score
Evaluator
Details Link
98.71%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
99.3562%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
They sat in the van. Molly adjusted her fake mustache. "You're absolutely certain the birthday girl doesn't have a security detail?" Bill squinted through binoculars. "It's a six-year-old, Molly. The only security is a bouncy castle and a T-rex piñata." "That T-rex could knock us senseless with one swipe." "It's papier-mâché. I checked thoroughly." "How do you know?" "I tasted it." Molly gagged. "You actually licked a piñata?" "For science. Also, I was extremely thirsty." "Bill, focus. We slip through the doggy door at 2:15, snatch the loot bags, and vanish before the cake ceremony." "What if the clown is actually armed?" "Then we negotiate." "With what? My devastating charm?" "You're holding a taser." "That's strictly for emergencies." "The clown is an emergency, Molly. He's wearing stilts." Bill dropped his binoculars. "Stilts change everything. We need a much larger duffel bag." "And a ladder." "Or our own stilts." "Don't be ridiculous." "Fine. But if there's a balloon animal shaped like a firearm, I'm keeping it." "Bill, this is a children's party." "Exactly. The balloon animals are weapons-grade." "Are you sure those aren't just regular clown balloons?" "I tested one. It squeaked like a rat."
created_at
"2026-10-07T09:10:49.196Z"
10/7/2026, 9:10:49 AM
model
"mistralai/mistral-large-4-0-20261006"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
638
moderation_latency
(null)
generation_time
53665
tokens_prompt
75
tokens_completion
4216
native_tokens_prompt
67
native_tokens_completion
4913
native_tokens_completion_images
(null)
native_tokens_reasoning
3893
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"stop"
service_tier
"default"
usage
0.01031373
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.21.0; linux; x64))"
http_referer
(null)
request_id
"req-1791364249-Q3SYYGryy4IHDhpopQwB"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1791364249-xNgz45nuYHPdviV6PLQU"
upstream_id
"5cf041d9f3624c1eb81dde954a8848b2"
provider_responses
0
endpoint_id
"19b3b5d7-23f6-4b63-8109-64a6cc37d996"
id
"5cf041d9f3624c1eb81dde954a8848b2"
is_byok
false
latency
621
model_permaslug
"mistralai/mistral-large-4-0-20261006"
provider_name
"Mistral"
status
200
total_cost
0.01031373
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Mistral"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
98.71%
Matches word count
n/a
neededClean
false
words
194
100.00%
Dialogue to Total Word Ratio
Ratio: 89.90%, Deviation: 0.10%
neededClean
false
wordsTotal
198
wordsDialogue
178
99.3562%