NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-01-01M1EKSSHY9Q59GT3PVYSHMHR1
Run Details
Model:
Grok 4.6 (Reasoning, High)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
2m 21s
Cost:
$0.049812
Provider:
xAI
Evaluation summary
Score
Evaluator
Details Link
99.74%
Matches word count
(details)
97.99%
Dialogue to Total Word Ratio
(details)
98.8665%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly leaned in. "Bill, tomorrow we hit the kids' party. Steal all presents during the chaos of games and cake." Bill frowned. "That's terrible, Molly. Kids will be heartbroken." "Heartbroken kids don't concern us. The loot does. Disguise as entertainers. You the clown, me the magician. Distract, then grab everything," Molly said. "Clown makeup itches. Magician better. Pull scarves, pocket gifts," Bill replied. "Scarves yes. Timing crucial. In at start, out before parents notice missing pile," Molly continued. "What if a child spots us taking a toy? We run, no confrontation," Bill warned. "Exactly. Comedy heist style. Slip, trip, presents in bag. Use a big gift bag as cover," Molly planned. "Gift bag brilliant. Fill it, walk out smiling. 'Thanks for inviting us,' we say," Bill added. "Yes. Practice the smile. This score funds our next job. No more small time," Molly declared. "Alright, I'm in. Let's steal those presents comically," Bill agreed. Molly smiled. "We'll laugh about this for years. The time we robbed a piƱata party." Bill nodded. "As long as we don't get caught by a six year old detective." "The presents will fund our retirement. Imagine, all those toys sold," Bill dreamed.
created_at
"2026-09-01T13:51:07.357Z"
9/1/2026, 1:51:07 PM
model
"x-ai/grok-4.6-20260810"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
561
moderation_latency
(null)
generation_time
140483
tokens_prompt
75
tokens_completion
792
native_tokens_prompt
267
native_tokens_completion
8245
native_tokens_completion_images
(null)
native_tokens_reasoning
7961
native_tokens_cached
128
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.049812
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.19.0; linux; x64))"
http_referer
(null)
request_id
"req-1788270667-L8aqtBcg9qdSvoBWsjny"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1788270667-HMmOA8H8ui0kRs6FGtYP"
upstream_id
"ddcfb825-a1b6-9e44-ae30-57972eef0579"
provider_responses
0
endpoint_id
"0d0536e3-7eb1-4acb-8249-c8813365c2d8"
id
"ddcfb825-a1b6-9e44-ae30-57972eef0579"
is_byok
false
latency
498
model_permaslug
"x-ai/grok-4.6-20260810"
provider_name
"xAI"
status
200
total_cost
0.049812
cache_discount
0.000192
upstream_inference_cost
0
provider_name
"xAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
99.74%
Matches word count
n/a
neededClean
false
words
196
97.99%
Dialogue to Total Word Ratio
Ratio: 86.22%, Deviation: 3.78%
neededClean
false
wordsTotal
196
wordsDialogue
169
98.8665%