NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-01-01M1EKRF4DGHQY4622XQZNBG9F
Run Details
Model:
Grok 4.6 (Reasoning, High)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
1m 45s
Cost:
$0.034554
Provider:
xAI
Evaluation summary
Score
Evaluator
Details Link
98.71%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
99.3562%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly and Bill huddled over a map of the suburban neighborhood, their eyes sparkling with the thrill of the upcoming job. The children's party was set to be a goldmine of presents, from high-tech toys to handmade crafts, all ripe for the taking in the most hilarious fashion possible. They planned to pose as entertainers, with Bill as a clumsy magician whose tricks would go spectacularly wrong, drawing attention away from Molly's nimble fingers. The backyard fence offered easy access, and they anticipated navigating through a sea of screaming kids and frazzled parents. Ideas for disguises included giant teddy bear suits that might cause more giggles than suspicion. They laughed at the image of themselves tripping over wrapping paper while stuffing loot into oversized pockets. The heist would be a comedy of errors, ending with a chase involving a piƱata-filled getaway van. Their excitement grew as they pictured the chaos of vanishing gifts amid cheers and laughter. They high-fived excitedly, discussing last-minute tweaks to ensure a smooth yet funny operation. "We'll sneak in the back and steal all presents during the song," Molly said. "I'll make a scene with the cake then," Bill replied.
created_at
"2026-09-01T13:50:24.124Z"
9/1/2026, 1:50:24 PM
model
"x-ai/grok-4.6-20260810"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
719
moderation_latency
(null)
generation_time
104272
tokens_prompt
75
tokens_completion
500
native_tokens_prompt
267
native_tokens_completion
5702
native_tokens_completion_images
(null)
native_tokens_reasoning
5460
native_tokens_cached
128
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.034554
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.19.0; linux; x64))"
http_referer
(null)
request_id
"req-1788270624-JtCivNZLoxgbvCTMEWJo"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1788270624-IucyxYDN3R2NOdZvkbX8"
upstream_id
"5aef8a76-2050-9b0b-8aa7-31b6ce10de7f"
provider_responses
0
endpoint_id
"0d0536e3-7eb1-4acb-8249-c8813365c2d8"
id
"5aef8a76-2050-9b0b-8aa7-31b6ce10de7f"
is_byok
false
latency
535
model_permaslug
"x-ai/grok-4.6-20260810"
provider_name
"xAI"
status
200
total_cost
0.034554
cache_discount
0.000192
upstream_inference_cost
0
provider_name
"xAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
98.71%
Matches word count
n/a
neededClean
false
words
194
100.00%
Dialogue to Total Word Ratio
Ratio: 10.10%, Deviation: 0.10%
neededClean
false
wordsTotal
198
wordsDialogue
20
99.3562%