NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-07-25-01KYD6T5WHKN8HDAJV1A91GACB
Run Details
Model:
Muse Spark 1.1 (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
15.5s
Cost:
$0.011102
Provider:
Meta
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
99.93%
Dialogue to Total Word Ratio
(details)
99.9629%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly spread the hand-drawn map across the hood of their rusted getaway van, weighing down the corners with stolen juice boxes. The paper showed a suburban house, a bouncy castle labeled DEATH TRAP, and a large X over the living room marked LOOT. Bill stared, his clown disguise already smearing in the heat. "We're stealing from a seven-year-old's birthday?" Molly tapped the X with fierce seriousness. Her plan had always involved low expectations and high sugar locations. "Think bigger. It's twenty seven-year-olds." Bill adjusted his squeaky red nose, which honked mournfully. He had imagined vaults and lasers, not balloon animals and screaming. "That's morally worse somehow." She explained the logistics with the intensity of a general. They would enter through the side gate during pin the tail, infiltrate via the kitchen, and stuff every brightly wrapped box into a laundry bag labeled PINATA. "Unwrapped upstairs. Easy score, right?" In the distance, a magician made terrible rabbit jokes. Bill sighed, realizing this was the only heist where the getaway could be foiled by a Capri-Sun spill. Molly grinned, already picturing their fence paying in LEGO sets and leftover cake frosting. The birthday playlist began loudly, mercilessly looping Baby Shark forever again.
created_at
"2026-07-25T17:57:48.568Z"
7/25/2026, 5:57:48 PM
model
"meta/muse-spark-1.1-20260709"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
14242
moderation_latency
(null)
generation_time
15434
tokens_prompt
75
tokens_completion
328
native_tokens_prompt
225
native_tokens_completion
2546
native_tokens_completion_images
0
native_tokens_reasoning
2254
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"auto"
usage
0.01110175
router
(null)
provider_responses
0
endpoint_id
"b2b9f6f9-8880-41c1-bd0c-867650fd5238"
id
"resp_6a64f91c22bdf9a4f7b2443e"
is_byok
false
latency
420
model_permaslug
"meta/muse-spark-1.1-20260709"
provider_name
"Meta"
status
200
user_agent
"langchainjs-openai/1.0.0 ((node/v24.18.0; linux; x64))"
http_referer
(null)
request_id
"req-1785002268-FFViXYrN7jJi0LoTIGHM"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1785002268-ynvqkCUfCbMam6sDxKsQ"
upstream_id
"resp_6a64f91c22bdf9a4f7b2443e"
total_cost
0.01110175
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Meta"
response_cache_source_id
(null)
data_region
"global"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
99.93%
Dialogue to Total Word Ratio
Ratio: 11.65%, Deviation: 1.65%
neededClean
false
wordsTotal
206
wordsDialogue
24
99.9629%