NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-01-01M1EMGA78K1XQK0D6QZJACM6W
Run Details
Model:
Grok 4.6 (Reasoning, High)
Test:
Dialogue tags
Scenario:
Write 200 words with 50% dialogue
Duration:
3m 28s
Cost:
$0.077262
Provider:
xAI
Evaluation summary
Score
Evaluator
Details Link
99.74%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
99.8721%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 50% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly and Bill plotted in their secret hideout, a cluttered garage filled with tools and fake IDs. They studied photos of the children's party location, noting the gift table's position near the swing set. The air was thick with tension as they finalized every detail of the upcoming break-in, ensuring no alarms would be triggered by the festive atmosphere. The stakes were high. "Bill, stealing from a kids' party is low even for us, but those presents are too tempting," Molly admitted. Bill grinned. "Come on Molly, it's just stuff. We'll replace them with cheap knockoffs later. The plan is to enter as the magician duo." "Magicians? I can do card tricks. You handle the disappearing act with the gifts." "Exactly. During the show, when all eyes are on the rabbit, we bag the wrapping paper treasures." "What if a child wants a trick involving a present? We improvise and say it's part of the magic." "Remember, no real harm, just the presents. We leave the kids their cake," Bill reminded. They high fived, their comedic heist taking shape amid laughs about potential mishaps with balloons and cake. Excitement built as they imagined the perfect execution of their silly scheme with very eager smiles.
created_at
"2026-09-01T14:03:25.304Z"
9/1/2026, 2:03:25 PM
model
"x-ai/grok-4.6-20260810"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
530
moderation_latency
(null)
generation_time
208138
tokens_prompt
75
tokens_completion
1653
native_tokens_prompt
267
native_tokens_completion
12820
native_tokens_completion_images
(null)
native_tokens_reasoning
12557
native_tokens_cached
128
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.077262
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.19.0; linux; x64))"
http_referer
(null)
request_id
"req-1788271405-Ve1AiOVGox5Z21VyBjty"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1788271405-iLMsAWORMQXgXyG50S5H"
upstream_id
"0296a066-4486-9eff-b2a0-03465809602e"
provider_responses
0
endpoint_id
"0d0536e3-7eb1-4acb-8249-c8813365c2d8"
id
"0296a066-4486-9eff-b2a0-03465809602e"
is_byok
false
latency
156
model_permaslug
"x-ai/grok-4.6-20260810"
provider_name
"xAI"
status
200
total_cost
0.077262
cache_discount
0.000192
upstream_inference_cost
0
provider_name
"xAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
99.74%
Matches word count
n/a
neededClean
false
words
204
100.00%
Dialogue to Total Word Ratio
Ratio: 49.76%, Deviation: 0.24%
neededClean
false
wordsTotal
205
wordsDialogue
102
99.8721%