NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-07-25-01KYCXWMY7BPVQ5EED0FYT7XP7
Run Details
Model:
Gemini 3.6 Flash (Reasoning)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
16.3s
Cost:
$0.028990
Provider:
Google AI Studio
Evaluation summary
Score
Evaluator
Details Link
99.98%
Matches word count
(details)
95.79%
Dialogue to Total Word Ratio
(details)
97.8863%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly tightened her fake mustache. "Did you secure the getaway vehicle?" "The pink tricycle with the squeaky horn is prepped," Bill said. "Are we truly doing this, Molly? It's a seven-year-old’s birthday party." "Leo’s parents are filthy rich, Bill. That haul includes a motorized mini Ferrari, three pristine Lego castles, and a customized robot dog." "What about the actual guard dog?" "A Goldendoodle named Mr. Fluffs. One slice of stolen pepperoni pizza renders him utterly useless. Now, what’s our entry point?" "The bouncy castle back flap," Bill replied. "I’ll cause a distraction by smashing the pinata early. Once the sugar-crazed toddlers swarm the lawn for cheap candy, you sweep the gift table into our burlap sacks." "Genius. What if the birthday boy spots us?" "Tell him we’re the late-arriving clown entertainment. Can you juggle?" "Only raw eggs, hot soup, and live explosives." "Just stick to terrible card tricks," Bill grunted. "Did you pack the tactical smoke bombs?" "Better. Pocketfuls of pink craft glitter. It blinds them instantly." "Cruel and shiny, but effective. Synchronize watches." "I haven't owned a watch since Denver." "Fine. On my mark, we storm the juice box station." "For the Legos, Bill!" "For the Legos!"
created_at
"2026-07-25T15:21:52.338Z"
7/25/2026, 3:21:52 PM
model
"google/gemini-3.6-flash-20260721"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
1485
moderation_latency
(null)
generation_time
16256
tokens_prompt
75
tokens_completion
1582
native_tokens_prompt
67
native_tokens_completion
3852
native_tokens_completion_images
0
native_tokens_reasoning
3534
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"STOP"
service_tier
"default"
usage
0.0289905
router
(null)
provider_responses
0
endpoint_id
"5d6d133d-a953-4781-a835-7ee76f4e1388"
id
"kNRkaoauF4q4sOIP3qGPkAo"
is_byok
false
latency
1485
model_permaslug
"google/gemini-3.6-flash-20260721"
provider_name
"Google AI Studio"
status
200
user_agent
"langchainjs-openai/1.0.0 ((node/v24.18.0; linux; x64))"
http_referer
(null)
request_id
"req-1784992912-0HzlXBqhKUXvwSoJMUoN"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1784992912-I5UwrLlqYN1XXTeMQroq"
upstream_id
"kNRkaoauF4q4sOIP3qGPkAo"
total_cost
0.0289905
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Google AI Studio"
response_cache_source_id
(null)
data_region
"global"
Evaluation details
Result
Evaluator
Details
Meta Data
99.98%
Matches word count
n/a
neededClean
false
words
198
95.79%
Dialogue to Total Word Ratio
Ratio: 94.55%, Deviation: 4.55%
neededClean
false
wordsTotal
202
wordsDialogue
191
97.8863%