NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-07-25-01KYD713882KB4128B2ST1374J
Run Details
Model:
Muse Spark 1.1 (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 200 words with 50% dialogue
Duration:
24.4s
Cost:
$0.018212
Provider:
Meta
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
99.9985%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 50% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
The van smelled like old bologna and bad decisions. "Molly, are we really robbing a six-year-old's birthday party?" Bill whispered. "We are liberating capitalism from children," Molly said. Molly adjusted her clown nose, which squeaked with judgment. "I brought the grappling hook and juice boxes for cover," Bill stared at the fortress, its fortress being a bouncy castle. "That castle is guarded by a very serious dad with a grill spatula," he said. "Spatula Man is nothing. I will distract him with small talk about lawn care," Molly grinned. "What's my job then? Stealing Bluey toys while crying inside?" Bill clutched the empty sack designed for presents and his dignity. "Exactly! You grab gifts, I grab cake. We're monsters, Bill. Efficient monsters," The sprinklers turned on, soaking their getaway Crocs. "If we get caught, we tell them we're entertainers?" "We are the worst entertainers ever," They high-fived, missed, and rolled toward inevitable embarrassment and frosting. "Charge! For plastic dinosaurs and moral bankruptcy!" Molly shouted, charging the bouncy castle like a squeaky, determined gremlin. "I regret everything already!" Bill sobbed, following her into a yard full of screaming children and questionable life choices. "Why?" The banner read welcome to bad decisions.
created_at
"2026-07-25T18:01:35.246Z"
7/25/2026, 6:01:35 PM
model
"meta/muse-spark-1.1-20260709"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
23184
moderation_latency
(null)
generation_time
24333
tokens_prompt
75
tokens_completion
324
native_tokens_prompt
225
native_tokens_completion
4219
native_tokens_completion_images
0
native_tokens_reasoning
3920
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"auto"
usage
0.018212
router
(null)
provider_responses
0
endpoint_id
"b2b9f6f9-8880-41c1-bd0c-867650fd5238"
id
"resp_6a64f9ff24f674dd6e9a49e8"
is_byok
false
latency
341
model_permaslug
"meta/muse-spark-1.1-20260709"
provider_name
"Meta"
status
200
user_agent
"langchainjs-openai/1.0.0 ((node/v24.18.0; linux; x64))"
http_referer
(null)
request_id
"req-1785002495-U9LSEvTirFqd4W1n9c0M"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1785002495-wNyQMGhSAsPsXOkoDSi5"
upstream_id
"resp_6a64f9ff24f674dd6e9a49e8"
total_cost
0.018212
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Meta"
response_cache_source_id
(null)
data_region
"global"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
100.00%
Dialogue to Total Word Ratio
Ratio: 50.74%, Deviation: 0.74%
neededClean
false
wordsTotal
203
wordsDialogue
103
99.9985%