NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-03-01M1M171SEFDTN0R6NTNT6VY0N
Run Details
Model:
Qwen 3.8 27B (Reasoning, XHigh)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
2m 56s
Cost:
$0.025471
Provider:
Phala
Evaluation summary
Score
Evaluator
Details Link
99.98%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
99.9920%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly counted party hats through the keyhole, each one a tiny ridiculous warning. Bill fixed his paper crown. The backyard smelled of lemon cake and parental despair. Inside, a toddler announced the location of every gift. Molly whispered, "Window at nine." Bill replied, "We take everything." They planned it over cereal in a van. The presents wore shark and unicorn paper. Molly would enter through the side window with inflatable gloves. Bill would cut the fence with a butter knife. If alarms sounded, they would freeze, because the alarm system was Crying Uncle. Molly said, "Gloves are balloons." Bill said, "Trap uses gum." Molly said, "Start quiet." Bill said, "Finish fast." Molly said, "No kids." Bill said, "No parents." The door banged open. Four children ran past the hedge, shrieking about piƱata duty. Bill hugged a hydrangea like a man having a private argument with a flower. Molly checked her watch, her gloves, and the suspicious sky. Inside, a grandmother raised a cake knife and declared the presents would be opened at one, not two. Bill exhaled. Molly tightened her crown. The heist was late, but they could fix that with theft, nerve, and zero respect for naptime.
created_at
"2026-09-03T16:21:45.397Z"
9/3/2026, 4:21:45 PM
model
"qwen/qwen3.8-27b-20260814"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
380
moderation_latency
(null)
generation_time
176135
tokens_prompt
75
tokens_completion
6781
native_tokens_prompt
186
native_tokens_completion
8473
native_tokens_completion_images
(null)
native_tokens_reasoning
8193
native_tokens_cached
64
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"stop"
service_tier
(null)
usage
0.025471
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.19.0; linux; x64))"
http_referer
(null)
request_id
"req-1788452505-VFUlzPgLUEL8fubddKan"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1788452505-7ozthO6JDs9F4qN5Rvmj"
upstream_id
"req_b6395005fca6f8d6b4a47188dae662e2"
provider_responses
0
endpoint_id
"cfc4d0ca-df7c-4b5f-9c74-328c41a7cbc3"
id
"req_b6395005fca6f8d6b4a47188dae662e2"
is_byok
false
latency
380
model_permaslug
"qwen/qwen3.8-27b-20260814"
provider_name
"Phala"
status
200
total_cost
0.025471
cache_discount
0.0000224
upstream_inference_cost
0
provider_name
"Phala"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
99.98%
Matches word count
n/a
neededClean
false
words
198
100.00%
Dialogue to Total Word Ratio
Ratio: 10.10%, Deviation: 0.10%
neededClean
false
wordsTotal
198
wordsDialogue
20
99.9920%