NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-02-01M1GWD8E1RX41BBYHBTVH813H
Run Details
Model:
Z.AI GLM 5.3 Flash (Reasoning, Max)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
1m 18s
Cost:
$0.000834
Provider:
Z.AI
Evaluation summary
Score
Evaluator
Details Link
99.92%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
99.9586%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly had survived jewel heists, bank jobs, and one deeply regrettable llama incident, but nothing prepared her for Bill's latest scheme, unveiled like the Mona Lisa. "It's a children's birthday party, Bill." Bill nodded proudly, unrolling a blueprint drawn entirely in crayon. His reconnaissance was thorough: balloon arch at the entrance, bouncy castle on the lawn, and a gift table guarded only by a laminated sign reading NO TOUCHING, which he considered an insult to professionals everywhere. "The sign said no touching. I'm taking it personally." The plan was elegant. They'd infiltrate through the bouncy castle, blend into the chaos of sugared children, and exit with the loot stuffed inside Happy Birthday gift bags. Molly listed the flaws: the security clown, an escape route blocked by a stroller conga line, the biological reality that children scream. "Kids bite." Bill had prepared for everything. His dossier on child behavior drew from forty minutes of nature documentaries, and his disguise included novelty hats plus enough frosting smeared across his face to pass as a professional cake taster. Molly studied the crayon blueprint, the grinning man. Somewhere, a classy jewel thief was spinning in her grave. "Fine. But I'm driving."
created_at
"2026-09-02T11:00:02.635Z"
9/2/2026, 11:00:02 AM
model
"z-ai/glm-5.3-flash-20260826"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
6025
moderation_latency
(null)
generation_time
78149
tokens_prompt
75
tokens_completion
3426
native_tokens_prompt
73
native_tokens_completion
3331
native_tokens_completion_images
(null)
native_tokens_reasoning
3072
native_tokens_cached
64
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"stop"
service_tier
(null)
usage
0.000834385
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.19.0; linux; x64))"
http_referer
(null)
request_id
"req-1788346802-MHwzsn5WV7CGOTs2eRvA"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1788346802-m4UnJIjesW0pMBde42us"
upstream_id
"20260902190002936742b11a804ceb"
provider_responses
0
endpoint_id
"8e9fe48b-2f91-41c3-a8a7-e4a93a8c4ff0"
id
"20260902190002936742b11a804ceb"
is_byok
false
latency
6025
model_permaslug
"z-ai/glm-5.3-flash-20260826"
provider_name
"Z.AI"
status
200
total_cost
0.000834385
cache_discount
0.00000768
upstream_inference_cost
0
provider_name
"Z.AI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
99.92%
Matches word count
n/a
neededClean
false
words
197
100.00%
Dialogue to Total Word Ratio
Ratio: 10.66%, Deviation: 0.66%
neededClean
false
wordsTotal
197
wordsDialogue
21
99.9586%