NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-07-25-01KYD5VK7E59Z1HF8E85QSMA0G
Run Details
Model:
Muse Spark 1.1 (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 200 words with 50% dialogue
Duration:
22.5s
Cost:
$0.016410
Provider:
Meta
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
100.0000%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 50% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
The van smelled like old frosting and bad decisions. "We are not robbing a seven-year-old's birthday party, Bill." Molly stared at the suburban house, balloons strangling the mailbox. "We are liberating high-value plastic assets from minors who cannot legally fight us." Bill unfolded a crayon-drawn map with alarming confidence. "You mean stealing Paw Patrol toys from crying children?" She sighed, tapping the steering wheel erratically. "I mean strategic gift redistribution, Molly, with juice boxes." He adjusted his oversized clown nose nervously. "What is the plan, genius?" Molly rolled her eyes so hard they squeaked. "We enter dressed as clowns, grab the pile, run." Bill grinned, revealing lipstick smeared teeth. "You are terrified of clowns and children." Outside, children screamed with sugar-fueled joy. "And you are great with kids, you only stole one stroller before." The present table glittered dangerously near the window. "That was for research and the baby forgave me." Molly checked her watch, already regretting everything. "Fine, we do it, but I keep the cake." Bill opened the door, letting in humid air and distant giggles. She grabbed the squeaky clown shoes reluctantly. "Deal, but no eating frosting in the getaway van." This was their worst idea yet.
created_at
"2026-07-25T17:41:06.515Z"
7/25/2026, 5:41:06 PM
model
"meta/muse-spark-1.1-20260709"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
21212
moderation_latency
(null)
generation_time
22269
tokens_prompt
75
tokens_completion
320
native_tokens_prompt
225
native_tokens_completion
3795
native_tokens_completion_images
0
native_tokens_reasoning
3501
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"auto"
usage
0.01641
router
(null)
provider_responses
0
endpoint_id
"b2b9f6f9-8880-41c1-bd0c-867650fd5238"
id
"resp_6a64f532bef4c18c48f4429a"
is_byok
false
latency
310
model_permaslug
"meta/muse-spark-1.1-20260709"
provider_name
"Meta"
status
200
user_agent
"langchainjs-openai/1.0.0 ((node/v24.18.0; linux; x64))"
http_referer
(null)
request_id
"req-1785001266-0Q7haAuJ2TzdkkMejjGc"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1785001266-HpJFPApOyjx8xiRytoVw"
upstream_id
"resp_6a64f532bef4c18c48f4429a"
total_cost
0.01641
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Meta"
response_cache_source_id
(null)
data_region
"global"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
100.00%
Dialogue to Total Word Ratio
Ratio: 50.24%, Deviation: 0.24%
neededClean
false
wordsTotal
205
wordsDialogue
103
100.0000%