NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-07-25-01KYD667QG3R75KHED912SA5VB
Run Details
Model:
MoonshotAI: Kimi K3 (Reasoning, High)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
42.7s
Cost:
$0.021014
Provider:
Moonshot AI
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
98.46%
Dialogue to Total Word Ratio
(details)
99.2276%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
"Molly, why are we wearing clown shoes to a burglary?" Bill whispered. "Because the invitation said costumes, and I commit to a theme." "It's a children's party." "Exactly. No one suspects a giraffe with lock picks." "You're a duck." "Then quack quieter. Now, when the magician yells 'abracadabra,' we grab the gift table." "That's thirty presents, Molly." "Thirty-one. I brought a tote labeled definitely not stolen loot." "Bill, stop eating the cake bait." "It said 'eat me.'" "That's for the birthday kid." "His wish was probably for us to learn boundaries." "Too late. Ducklings, formation. You distract with balloon animals; I sweep the pile. Keep your wings down." Molly hissed. "If we trip over the clown shoes, I'm telling prison it was community theater." Bill nodded solemnly with jazz hands and emotional deniability. "My balloon animals look like evidence." "Perfect, confuse the detectives." "What if a child sees us?" "Offer them a juice box and a life of plausible deniability." "That's bribery." "That's hospitality with pockets. Ready?" "Molly, the piƱata is judging me." "Good. Let it witness history. On three, we waddle." "One." "Two." "Wait, I left my moral compass in the van." "Bill, we never had one." "Oh. Three."
created_at
"2026-07-25T17:46:55.095Z"
7/25/2026, 5:46:55 PM
model
"moonshotai/kimi-k3-20260715"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
2342
moderation_latency
(null)
generation_time
42554
tokens_prompt
75
tokens_completion
1412
native_tokens_prompt
147
native_tokens_completion
1398
native_tokens_completion_images
(null)
native_tokens_reasoning
1085
native_tokens_cached
147
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"stop"
service_tier
(null)
usage
0.0210141
router
(null)
provider_responses
0
endpoint_id
"650c9b12-7b6d-444d-8d92-df8daef84d7a"
id
"chatcmpl-6a64f690b588261998aba716"
is_byok
false
latency
2342
model_permaslug
"moonshotai/kimi-k3-20260715"
provider_name
"Moonshot AI"
status
200
user_agent
"langchainjs-openai/1.0.0 ((node/v24.18.0; linux; x64))"
http_referer
(null)
request_id
"req-1785001615-uARNdK3exyCfBRniGTTC"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1785001615-dgof4w1eVLXYTsIFiPPG"
upstream_id
"chatcmpl-6a64f690b588261998aba716"
total_cost
0.0210141
cache_discount
0.0003969
upstream_inference_cost
0
provider_name
"Moonshot AI"
response_cache_source_id
(null)
data_region
"global"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
98.46%
Dialogue to Total Word Ratio
Ratio: 93.53%, Deviation: 3.53%
neededClean
false
wordsTotal
201
wordsDialogue
188
99.2276%