NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M36S9VXD38P2KTC0HXJZX3TG
Run Details
Model:
Grok 4.7 (Reasoning, High)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
25.5s
Cost:
$0.008210
Provider:
xAI
Evaluation summary
Score
Evaluator
Details Link
98.71%
Matches word count
(details)
78.76%
Dialogue to Total Word Ratio
(details)
88.7360%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly slid the napkin across the table. "Bill, the party is at two. We walk in smiling." "Smiling at what?" Bill said. "I am not a child. I am a man with knees." "You are Uncle Bill from out of town," Molly said. "You admire the balloons. You do not mention your knees." "And we just take the presents?" Bill asked. "The pile with the little cards? I read one. It says Timmy." "Timmy will survive," Molly said. "He has a bouncy castle and three grandmothers." "I am not wrestling a grandmother," Bill said. "Nobody is wrestling anyone," Molly said. "We compliment the cake. We stand near the table. When they sing, we carry the pile to the car like helpful adults." "Helpful adults do not steal scooters," Bill said. "Then leave the scooter," Molly said. "Take the boxes. Leave the piƱata. I have principles." "Your principles are candy," Bill said. "Correct," Molly said. "Also, if the clown stares, do not confess." "Clowns can tell," Bill said. "Clowns cannot arrest you," Molly said. "Probably. Smile, Uncle Bill. And do not say the cake looks moist." Bill groaned. "This is the worst job we have ever planned." "It is a party," Molly said. "Try to look invited."
created_at
"2026-09-23T09:24:45.117Z"
9/23/2026, 9:24:45 AM
model
"x-ai/grok-4.7-20260916"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
567
moderation_latency
(null)
generation_time
25452
tokens_prompt
75
tokens_completion
541
native_tokens_prompt
1303
native_tokens_completion
1564
native_tokens_completion_images
(null)
native_tokens_reasoning
1256
native_tokens_cached
1152
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.0082096
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790155485-kmplGGUKRTpKpKXNcIuR"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790155485-ektZR3mJZpEn8L7yfPFp"
upstream_id
"bb7acf17-1255-9d55-8368-89c1657537d9"
provider_responses
0
endpoint_id
"c942f37a-87e1-4fa8-83d0-8c67fda4fb4a"
id
"bb7acf17-1255-9d55-8368-89c1657537d9"
is_byok
false
latency
534
model_permaslug
"x-ai/grok-4.7-20260916"
provider_name
"xAI"
status
200
total_cost
0.0082096
cache_discount
0.0013824
upstream_inference_cost
0
provider_name
"xAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
98.71%
Matches word count
n/a
neededClean
false
words
206
78.76%
Dialogue to Total Word Ratio
Ratio: 83.01%, Deviation: 6.99%
neededClean
false
wordsTotal
206
wordsDialogue
171
88.7360%