NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-02-01M1HHW4QXRABEXC2DQH5K8REB
Run Details
Model:
Qwen 3.8 Flash (Reasoning)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
38.5s
Cost:
$0.000000
Provider:
Alibaba
Evaluation summary
Score
Evaluator
Details Link
0.00%
Matches word count
(details)
85.76%
Dialogue to Total Word Ratio
(details)
42.8815%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly adjusted her balaclava, which had a small unicorn on it because camouflage was for amateurs. Bill patted the empty sack on his back, then checked the lockpick set hidden inside a juice box. They crouched behind a hedge decorated with balloons shaped like smiling animals. The house hummed with sugar, shrieks, and a playlist stuck between songs about friendship. Through the window, a mountain of presents waited beside a bouncy castle, each box ribbed like a cake. Molly: "The clown van is our getaway." Bill: "It smells like cheese and regret." Molly: "We will steal every gift." Bill: "Excellent, obviously." Molly: "Perfect." Then they studied the guest list taped to the door: thirty toddlers, four exhausted adults, one dog named Chairman Meow.
created_at
"2026-09-02T17:15:10.479Z"
9/2/2026, 5:15:10 PM
model
"qwen/qwen3.8-flash-20260826"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
1105
moderation_latency
(null)
generation_time
38392
tokens_prompt
75
tokens_completion
2832
native_tokens_prompt
75
native_tokens_completion
2832
native_tokens_completion_images
(null)
native_tokens_reasoning
2641
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
(null)
native_finish_reason
(null)
service_tier
(null)
usage
0
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.19.0; linux; x64))"
http_referer
(null)
request_id
"req-1788369310-SMIFYs9Rhy4mllKHdqGx"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1788369310-2Oz0J1vjMh7w9naWkQqx"
upstream_id
"chatcmpl-9dd186b8-97e7-911c-bda2-85dafaeb5cb8"
provider_responses
0
endpoint_id
"84b1e4a7-5aed-4464-818f-1994f0b4ee18"
id
"chatcmpl-9dd186b8-97e7-911c-bda2-85dafaeb5cb8"
is_byok
false
latency
1103
model_permaslug
"qwen/qwen3.8-flash-20260826"
provider_name
"Alibaba"
status
200
total_cost
0
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Alibaba"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
0.00%
Matches word count
n/a
neededClean
false
words
123
85.76%
Dialogue to Total Word Ratio
Ratio: 16.26%, Deviation: 6.26%
neededClean
false
wordsTotal
123
wordsDialogue
20
42.8815%