NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-07-25-01KYD7FJX6ZJ54C8CCQBRCRQDE
Run Details
Model:
Muse Spark 1.1 (Reasoning, Minimal)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
17.2s
Cost:
$0.010494
Provider:
Meta
Evaluation summary
Score
Evaluator
Details Link
98.71%
Matches word count
(details)
99.68%
Dialogue to Total Word Ratio
(details)
99.1944%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly peered through the bushes. "Listen, Bill, operation Birthday Burgle starts in ten. You ready?" Bill sighed. "Ready to steal toys from a six-year-old? Molly, we're at rock bottom." "Bottom? We get fifty wrapped presents, no alarms, just juice boxes. Genius!" "Genius? There's twenty kids, two moms, and a clown watching that mountain." "So you distract the clown, I grab the loot with the wagon." "Me? Distract a clown? Last time I nearly got balloon-animaled to death!" "Just tell jokes! You love attention. Say something funny, fall over." "And you? You'll waddle out with a wagon of Paw Patrol screaming victory?" "Exactly. We vanish before cake. Cleanest heist ever, Bill." "Clean? We'll be on a mom's Facebook as the worst magicians ever." "Relax. In and out. What could possibly go wrong with toddlers?" Bill stared. "What if one bites?" "Then you bite back, you big baby. It's a present pile, not Fort Knox." "Molly, I'm telling you, those juice boxes are drugged with sugar. Those kids are feral." "Perfect. Feral kids love me. Now put on the pointy hat." Molly grinned. Bill groaned. "This hat smells like old ham." "Perfect disguise then," Molly whispered.
created_at
"2026-07-25T18:09:30.029Z"
7/25/2026, 6:09:30 PM
model
"meta/muse-spark-1.1-20260709"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
15866
moderation_latency
(null)
generation_time
17074
tokens_prompt
75
tokens_completion
299
native_tokens_prompt
225
native_tokens_completion
2403
native_tokens_completion_images
0
native_tokens_reasoning
2103
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"auto"
usage
0.010494
router
(null)
provider_responses
0
endpoint_id
"b2b9f6f9-8880-41c1-bd0c-867650fd5238"
id
"resp_6a64fbdadf589c5fedf14e7c"
is_byok
false
latency
473
model_permaslug
"meta/muse-spark-1.1-20260709"
provider_name
"Meta"
status
200
user_agent
"langchainjs-openai/1.0.0 ((node/v24.18.0; linux; x64))"
http_referer
(null)
request_id
"req-1785002970-LCJZb9xlNF3tvRntJ9qb"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1785002970-KeSeajswIRbxu4mAd6Ux"
upstream_id
"resp_6a64fbdadf589c5fedf14e7c"
total_cost
0.010494
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Meta"
response_cache_source_id
(null)
data_region
"global"
Evaluation details
Result
Evaluator
Details
Meta Data
98.71%
Matches word count
n/a
neededClean
false
words
194
99.68%
Dialogue to Total Word Ratio
Ratio: 92.39%, Deviation: 2.39%
neededClean
false
wordsTotal
197
wordsDialogue
182
99.1944%