NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-02-01M1HHAJHH7KAH5FHJABC7NYGK
Run Details
Model:
Qwen 3.8 Flash (Reasoning)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
1m 36s
Cost:
$0.003565
Provider:
Alibaba
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
99.9993%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly and Bill planned to raid a party for kids, stealing presents while nearby adults laughed too loudly after sunset tonight. "The clown car can hide our loot, the cake can serve as distraction tonight if we are lucky and clever." "Bill, I can disguise myself as a giant purple llama and carry every box away myself without any squeaky mistake." "No squeaky mistakes. The parents will think the llama has a gift for hoarding glitter before cake time and juice." "Molly, when they open the pinata, we will slip behind the banner and start the heist with style and silence." "Can we steal the stuffed horse too? It has four legs and excellent gift capacity if nobody notices or complains." "Only if the horse becomes a getaway mount. We ride away as birthday villains in sequined pajamas and tiny helmets." "Sequined pajamas are too loud. A cape made of wrapped gifts will make us invisible, smelling frosting and danger too." "Bill, do the presents include toys or just embarrassing socks for confused relatives? We cannot risk stealing boring fabric today." "Every present is valuable, especially if the card says money. We are romantic villains who steal cake and ribbons too."
created_at
"2026-09-02T17:05:34.777Z"
9/2/2026, 5:05:34 PM
model
"qwen/qwen3.8-flash-20260826"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
3785
moderation_latency
(null)
generation_time
95702
tokens_prompt
75
tokens_completion
6880
native_tokens_prompt
127
native_tokens_completion
7544
native_tokens_completion_images
(null)
native_tokens_reasoning
7294
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"stop"
service_tier
(null)
usage
0.00356473
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.19.0; linux; x64))"
http_referer
(null)
request_id
"req-1788368734-ApiVjgbSNptrOYBetH9o"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1788368734-hGsWSXDAAZFrG52Kmu97"
upstream_id
"chatcmpl-b9e247d1-2ac4-96c5-9bc9-cfe48b1b23d2"
provider_responses
0
endpoint_id
"84b1e4a7-5aed-4464-818f-1994f0b4ee18"
id
"chatcmpl-b9e247d1-2ac4-96c5-9bc9-cfe48b1b23d2"
is_byok
false
latency
3785
model_permaslug
"qwen/qwen3.8-flash-20260826"
provider_name
"Alibaba"
status
200
total_cost
0.00356473
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Alibaba"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
201
100.00%
Dialogue to Total Word Ratio
Ratio: 89.55%, Deviation: 0.45%
neededClean
false
wordsTotal
201
wordsDialogue
180
99.9993%