NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-02-01M1HH7S6A72BFQW1K9F0NKX7Z
Run Details
Model:
Qwen 3.8 Flash (Reasoning)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
1m 50s
Cost:
$0.003628
Provider:
Alibaba
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
100.0000%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly eyed the backyard, streamers snapping. Bill scanned the gift table, grinning greedily. They whispered like villains planning chaos together. "Bill, did you see that pony-shaped piñata?" "Only the one full of candy? No, wait, the one filled with tiny drones." "Tiny drones?" "They unlock gifts. Adults never notice." "What about the children?" "Kids scream. Perfect distraction." "That is evil, Molly." "It is comedy. Also profitable. Now, when the magician says presto, we—" "We?" "We enter with a cart." "What cart?" "Inflatables. I stole one." "Already?" "I am committed. We hide inside, then roll to the presents." "The presents are behind a cake." "Perfect. Sugar obscures judgment. We take the wrapped boxes first." "What about the expensive electronics?" "Those get stuffed into a piñata shell. No one suspects festive technology." "And afterward?" "We exit before the clown gets offended." "The clown?" "He finally owes me for the broken balloon animal." "He is our lookout. Very persuasive makeup." "Molly, this is insane." "Insane works. Bring the glitter bag." "What will I do with glitter?" "Throw it everywhere. Children adore chaos. We need more of it than they can handle." "Should we apologize?" "Not yet. We apologize at the end, in handcuffs."
created_at
"2026-09-02T17:04:30.116Z"
9/2/2026, 5:04:30 PM
model
"qwen/qwen3.8-flash-20260826"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
698
moderation_latency
(null)
generation_time
83584
tokens_prompt
75
tokens_completion
6535
native_tokens_prompt
127
native_tokens_completion
7679
native_tokens_completion_images
(null)
native_tokens_reasoning
7374
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"stop"
service_tier
(null)
usage
0.00362818
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.19.0; linux; x64))"
http_referer
(null)
request_id
"req-1788368670-U8o3qyXmTLDQt4r7znOP"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1788368670-cGCyXO2P3OQNp4D2b5qF"
upstream_id
"chatcmpl-387ee1e6-0c60-902a-b7a9-14f20aadfb8b"
provider_responses
0
endpoint_id
"84b1e4a7-5aed-4464-818f-1994f0b4ee18"
id
"chatcmpl-387ee1e6-0c60-902a-b7a9-14f20aadfb8b"
is_byok
false
latency
697
model_permaslug
"qwen/qwen3.8-flash-20260826"
provider_name
"Alibaba"
status
200
total_cost
0.00362818
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Alibaba"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
100.00%
Dialogue to Total Word Ratio
Ratio: 90.05%, Deviation: 0.05%
neededClean
false
wordsTotal
201
wordsDialogue
181
100.0000%