NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-07-25-01KYD62KH21V29CQ0C4N6MAP5X
Run Details
Model:
Muse Spark 1.1 (Reasoning, Minimal)
Test:
Dialogue tags
Scenario:
Write 200 words with 50% dialogue
Duration:
15.9s
Cost:
$0.012559
Provider:
Meta
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
100.0000%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 50% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly spread the crumpled birthday invitation across the getaway car's dashboard like a treasure map. "Bill, this is it. Seven-year-olds, hyperactive, distracted parents, mountain of untouched Lego sets." Bill adjusted his clown nose, which squeaked mournfully. "Molly, we're too old for clown school and too young for federal prison for stealing Bluey merch." The plan was simple, stupid, and involved a bouncy castle diversion. "You create chaos in the ball pit, I grab the present table and we vanish." Outside the suburban house, thirty screaming children chased a harried magician. "That's our cue. Remember, act like hired entertainment, not desperate adult failures." Molly inflated her balloon dog until it resembled a panicked giraffe. "What if they ask us to actually make balloon animals? I only know existential dread." Bill sighed, opened the trunk, and revealed two tiny tuxedos and a sack labeled LOOT. "Relax. Kids love dread. Now let's steal their joy, literally." They high-fived, slipped on oversized shoes, and waddled toward destiny, capitalism, and cake. "If we get caught, we tell them it's performance art about property law." The doorbell rang, sounding suspiciously like a police siren. "Smile, Bill. We're terrible people." He smiled, revealing missing tooth and regret.
created_at
"2026-07-25T17:44:56.104Z"
7/25/2026, 5:44:56 PM
model
"meta/muse-spark-1.1-20260709"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
14951
moderation_latency
(null)
generation_time
15831
tokens_prompt
75
tokens_completion
329
native_tokens_prompt
225
native_tokens_completion
2889
native_tokens_completion_images
0
native_tokens_reasoning
2595
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"auto"
usage
0.0125595
router
(null)
provider_responses
0
endpoint_id
"b2b9f6f9-8880-41c1-bd0c-867650fd5238"
id
"resp_6a64f618170ef23903524f05"
is_byok
false
latency
422
model_permaslug
"meta/muse-spark-1.1-20260709"
provider_name
"Meta"
status
200
user_agent
"langchainjs-openai/1.0.0 ((node/v24.18.0; linux; x64))"
http_referer
(null)
request_id
"req-1785001496-tzEsERjSAh7ubAKO7bC7"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1785001496-SnKILV6QOiJkNnaQNgWa"
upstream_id
"resp_6a64f618170ef23903524f05"
total_cost
0.0125595
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Meta"
response_cache_source_id
(null)
data_region
"global"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
100.00%
Dialogue to Total Word Ratio
Ratio: 50.25%, Deviation: 0.25%
neededClean
false
wordsTotal
203
wordsDialogue
102
100.0000%