NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-07-25-01KYCTMDSSG5Z7TGPNQ0FYQGA5
Run Details
Model:
Gemini 3.6 Flash (Reasoning)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
21.2s
Cost:
$0.039798
Provider:
Google AI Studio
Evaluation summary
Score
Evaluator
Details Link
99.38%
Matches word count
(details)
82.70%
Dialogue to Total Word Ratio
(details)
91.0376%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
"We strike at fourteen hundred hours," Molly whispered, unrolling a blueprint. "That is naptime," Bill said. "Are you serious?" "Dead serious. The target is Timmy's seventh birthday bash. Sector four has a bounce house, sector two has the juice boxes." "And the main loot?" "Stacked on the gift table. Legos, an Xbox, a motorized scooter. High value, extremely easy to fence." "What about security?" "A clown named Sprinkles. Unpredictable, highly aggressive with balloon swords. If he spots you, play dead." "Play dead?" "Or throw glitter in his eyes. Now, your disguise is ready." "Is it the inflatable T-Rex again?" "It is the inflatable T-Rex. It offers perfect tactical cover." "I'll look completely ridiculous." "You'll look like paid entertainment! While you distract the hyperventilating parents with terrible dance moves, I'll slide under the banner, bag the toys, and initiate Protocol Sugar Rush." "Protocol Sugar Rush?" "We smash the giant piƱata to trigger a manic toddler stampede. In the chaos, we vanish with all the presents." "Molly, you are sick." "It's business, Bill." "No, I mean brilliant. Does this tail make my tactical boots look ridiculous?" "Bill, focus on the loot. Is the dinosaur suit zippable?" "Barely. Is it flame-retardant?" "Obviously not." "Right. Hand me the head."
created_at
"2026-07-25T14:24:57.156Z"
7/25/2026, 2:24:57 PM
model
"google/gemini-3.6-flash-20260721"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
1286
moderation_latency
(null)
generation_time
21149
tokens_prompt
75
tokens_completion
1805
native_tokens_prompt
67
native_tokens_completion
5293
native_tokens_completion_images
0
native_tokens_reasoning
4954
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"STOP"
service_tier
"default"
usage
0.039798
router
(null)
provider_responses
0
endpoint_id
"5d6d133d-a953-4781-a835-7ee76f4e1388"
id
"OcdkaruHDqvWjMcPn82XoAY"
is_byok
false
latency
1281
model_permaslug
"google/gemini-3.6-flash-20260721"
provider_name
"Google AI Studio"
status
200
user_agent
"langchainjs-openai/1.0.0 ((node/v24.18.0; linux; x64))"
http_referer
(null)
request_id
"req-1784989497-bRvxxQkH9rMlJ4ngaZos"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1784989497-ckfDFZAlr4wag9GLLuJM"
upstream_id
"OcdkaruHDqvWjMcPn82XoAY"
total_cost
0.039798
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Google AI Studio"
response_cache_source_id
(null)
data_region
"global"
Evaluation details
Result
Evaluator
Details
Meta Data
99.38%
Matches word count
n/a
neededClean
false
words
205
82.70%
Dialogue to Total Word Ratio
Ratio: 96.60%, Deviation: 6.60%
neededClean
false
wordsTotal
206
wordsDialogue
199
91.0376%