NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-23-01M36SX8098M3PK0J4K99XQZEX
Run Details
Model:
Grok 4.7 (Reasoning, High)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
1m 29s
Cost:
$0.034869
Provider:
xAI
Evaluation summary
Score
Evaluator
Details Link
99.98%
Matches word count
(details)
66.53%
Dialogue to Total Word Ratio
(details)
83.2561%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
"Bill, backyard, three o'clock, presents by the grill," Molly said. "How many?" Bill asked. "Enough for the trunk. We slip the latch, walk in with a balloon, and nobody looks twice." "A balloon." "Camouflage. When the piƱata cracks, every eye goes to the candy. That is our window. You take the left pile, I take the right, we meet at the gate." "What do I say if a mom stops me?" "You say you are with the entertainment. Smile. Do not run until the sidewalk." "And if a kid follows the gifts?" "Hand them one small box and keep walking. Sacrifice a toy, save the haul." "This is unhinged." "This is efficient. My sister's minivan is on the corner. Car seats in the back so we look local." "I hate that this might work." "It works if you do not narrate it. Gloves on. Wrapping paper slides." "Name of the birthday kid?" "Does not matter. We are not staying for cake." "I like cake." "You can like cake in the van. Are you in?" "I am in. If a first-grader tackles me, you owe me dinner." "Done. Now stop shaking and look like you belong at a birthday."
created_at
"2026-09-23T09:35:20.079Z"
9/23/2026, 9:35:20 AM
model
"x-ai/grok-4.7-20260916"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
606
moderation_latency
(null)
generation_time
88727
tokens_prompt
75
tokens_completion
366
native_tokens_prompt
1303
native_tokens_completion
7150
native_tokens_completion_images
(null)
native_tokens_reasoning
6856
native_tokens_cached
1280
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.0348688
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.20.0; linux; x64))"
http_referer
(null)
request_id
"req-1790156120-nAwcV0WX2n87chveGIh1"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790156120-SeN7EY6uUgpAd5RdVV79"
upstream_id
"093c1b3d-31dd-973f-abd2-6886d6318869"
provider_responses
0
endpoint_id
"c942f37a-87e1-4fa8-83d0-8c67fda4fb4a"
id
"093c1b3d-31dd-973f-abd2-6886d6318869"
is_byok
false
latency
206
model_permaslug
"x-ai/grok-4.7-20260916"
provider_name
"xAI"
status
200
total_cost
0.0348688
cache_discount
0.001536
upstream_inference_cost
0
provider_name
"xAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
99.98%
Matches word count
n/a
neededClean
false
words
198
66.53%
Dialogue to Total Word Ratio
Ratio: 97.99%, Deviation: 7.99%
neededClean
false
wordsTotal
199
wordsDialogue
195
83.2561%