NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-07-25-01KYCYJWM0417BE8R3NKEFQWV0
Run Details
Model:
Thinking Machines Inkling (Reasoning)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
6m 49s
Cost:
$0.022971
Provider:
DeepInfra
Evaluation summary
Score
Evaluator
Details Link
60.28%
Matches word count
(details)
99.99%
Dialogue to Total Word Ratio
(details)
80.1313%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly spread the crinkled blueprints across the picnic table with the ceremonial gravity of a surgeon about to steal a kidney from a very small patient. The target was a five-year-old’s birthday extravaganza at the community center—dinosaur balloons tied with reckless optimism, a bouncy castle reeking of synthetic rubber and parental regret, and a gift table sagging beneath unwrapped Lego architectures and tablets still sealed in boxes promising educational miracles. Bill chewed a licorice stick he’d liberated from an unlocked glove compartment, studying the diagram with the enthusiasm of a man reading tax law in a language he didn’t speak. “The clown starts at two,” he said. Molly tapped the red circle marking the service entrance. “We wear aprons, carry hollow cakes filled with burlap, and vanish while parents weep over sugar.” Bill squinted at the sketched piñata hanging like a colorful threat. “Kids scream loud.” Molly tossed him neon earplugs and a bribed magician’s rabbit, which twitched in her coat pocket. Outside, a minivan sprayed gravel across the empty lot. She checked her watch, grinning at the pure immorality of the enterprise. By three o’clock they’d be drinking stolen champagne in a van crammed with presents nobody deserved more than they did, leaving behind only empty wrapping paper and one very confused bunny.
created_at
"2026-07-25T15:34:01.096Z"
7/25/2026, 3:34:01 PM
model
"thinkingmachines/inkling-20260715"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
997
moderation_latency
(null)
generation_time
408842
tokens_prompt
75
tokens_completion
4912
native_tokens_prompt
75
native_tokens_completion
5660
native_tokens_completion_images
(null)
native_tokens_reasoning
4561
native_tokens_cached
32
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"stop"
service_tier
(null)
usage
0.02297144
router
(null)
provider_responses
0
endpoint_id
"d6b9d68e-115a-49d7-81cd-1ad1c6942b0e"
id
"chatcmpl-Rk9KwqQ5PC2nRgpsVKCdfZ4K"
is_byok
false
latency
221
model_permaslug
"thinkingmachines/inkling-20260715"
provider_name
"DeepInfra"
status
200
user_agent
"langchainjs-openai/1.0.0 ((node/v24.18.0; linux; x64))"
http_referer
(null)
request_id
"req-1784993641-dC7ybRXOppz2UNMMh0op"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1784993641-qws5s4ea408uPZBpedUN"
upstream_id
"chatcmpl-Rk9KwqQ5PC2nRgpsVKCdfZ4K"
total_cost
0.02297144
cache_discount
0.00002656
upstream_inference_cost
0
provider_name
"DeepInfra"
response_cache_source_id
(null)
data_region
"global"
Evaluation details
Result
Evaluator
Details
Meta Data
60.28%
Matches word count
n/a
neededClean
false
words
215
99.99%
Dialogue to Total Word Ratio
Ratio: 11.06%, Deviation: 1.06%
neededClean
false
wordsTotal
217
wordsDialogue
24
80.1313%