NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-02-01M1HJ47C3YQ62WJFTA9W7TS62
Run Details
Model:
Qwen 3.8 Flash (Reasoning)
Test:
Dialogue tags
Scenario:
Write 200 words with 50% dialogue
Duration:
1m 51s
Cost:
$0.004448
Provider:
Alibaba
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
100.0000%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 50% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly adjusted her thief costume beside the bouncy castle, her gloved finger tracing the window where many wrapped presents waited inside beneath streamers. "We wait until the clown falls asleep. Then we take the gift pile, not the cupcakes. That tiny crumb incident remains shameful for any master," Molly said. Bill squinted at the invitation on a balloon, then pointed at the sugar crash table where chaos simmered between juice boxes and cake. "Do we knock, roll in, or fake a distant aunt? The toddlers notice everything, and the parents have phones. I would prefer a dramatic entry," Bill said. Molly checked the lockpick bracelet, then pulled a fake mustache from her pocket because disguises made everything more plausible and slightly illegal today. "The mustache is excellent, but it makes us look like we stole a dentist. We need a plan, not a theatrical beard emergency this afternoon," Bill said. The bouncy castle deflated with a sigh, providing conveniently noisy cover for two professionals about to commit crimes against children's gifts before lunch. "Perfect. Bill, you distract the parents with complaints about gluten, while I collect every gift tag and bow. We split afterward by usefulness, not sentiment," Molly said.
created_at
"2026-09-02T17:19:41.608Z"
9/2/2026, 5:19:41 PM
model
"qwen/qwen3.8-flash-20260826"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
3793
moderation_latency
(null)
generation_time
105129
tokens_prompt
75
tokens_completion
8582
native_tokens_prompt
127
native_tokens_completion
9423
native_tokens_completion_images
(null)
native_tokens_reasoning
9169
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"stop"
service_tier
(null)
usage
0.00444786
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.19.0; linux; x64))"
http_referer
(null)
request_id
"req-1788369581-g9nGrYAZwsvUZBBszyMZ"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1788369581-6Wcha7PhcJiefUbyQomm"
upstream_id
"chatcmpl-f2fc6910-c900-9cfd-92ca-80c1231920f9"
provider_responses
0
endpoint_id
"84b1e4a7-5aed-4464-818f-1994f0b4ee18"
id
"chatcmpl-f2fc6910-c900-9cfd-92ca-80c1231920f9"
is_byok
false
latency
3793
model_permaslug
"qwen/qwen3.8-flash-20260826"
provider_name
"Alibaba"
status
200
total_cost
0.00444786
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Alibaba"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
100.00%
Dialogue to Total Word Ratio
Ratio: 50.00%, Deviation: 0.00%
neededClean
false
wordsTotal
200
wordsDialogue
100
100.0000%