NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-01-01M1ENGH79A4M7YZQK2WF0VKZM
Run Details
Model:
Grok 4.6 (Reasoning, High)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
1m 57s
Cost:
$0.042204
Provider:
xAI
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
100.0000%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
The dim light of the streetlamp filtered through the grimy window as Molly and Bill sat across from each other, maps and invitations scattered like confetti from a failed party. They were planning the heist of a lifetime, or at least the most childish one: infiltrating a children's birthday celebration to swipe every present in sight. Molly, her eyes gleaming with mischief, outlined how they would pose as distant relatives to gain entry without suspicion. Bill nodded, thinking about the getaway route through the quiet suburban streets after loading their trunk with loot. "Imagine their faces when the gifts disappear. We'll be in and out before the ice cream melts," Molly said. "Let's do this," Bill replied eagerly. The thought of toddlers crying over missing toys brought a chuckle, highlighting the comedic nature of their criminal endeavor. They reviewed the schedule, noting the cake time as the perfect distraction when parents would be busy photographing the event. No security guards, just love and laughter that they intended to exploit for profit. Their van was packed with empty bags ready to be filled. This could be the score that retired them, or the one that made them laughingstocks in the underworld.
created_at
"2026-09-01T14:21:01.041Z"
9/1/2026, 2:21:01 PM
model
"x-ai/grok-4.6-20260810"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
492
moderation_latency
(null)
generation_time
116629
tokens_prompt
75
tokens_completion
1665
native_tokens_prompt
267
native_tokens_completion
7009
native_tokens_completion_images
(null)
native_tokens_reasoning
6774
native_tokens_cached
256
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.042204
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.19.0; linux; x64))"
http_referer
(null)
request_id
"req-1788272461-KRacvL9LiO5t3TAHUY3x"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1788272461-hCsYV4b5uHDs8EfngbYB"
upstream_id
"1f8e3f65-71b5-98bc-97dc-8e62604f42e7"
provider_responses
0
endpoint_id
"0d0536e3-7eb1-4acb-8249-c8813365c2d8"
id
"1f8e3f65-71b5-98bc-97dc-8e62604f42e7"
is_byok
false
latency
101
model_permaslug
"x-ai/grok-4.6-20260810"
provider_name
"xAI"
status
200
total_cost
0.042204
cache_discount
0.000384
upstream_inference_cost
0
provider_name
"xAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
100.00%
Dialogue to Total Word Ratio
Ratio: 10.00%, Deviation: 0.00%
neededClean
false
wordsTotal
200
wordsDialogue
20
100.0000%