NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-07-25-01KYD5KAD1XNVF733X3E9K2C9Z
Run Details
Model:
Muse Spark 1.1 (Reasoning, Minimal)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
8.1s
Cost:
$0.004650
Provider:
Meta
Evaluation summary
Score
Evaluator
Details Link
99.98%
Matches word count
(details)
51.96%
Dialogue to Total Word Ratio
(details)
75.9736%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
"Molly, you said big score. You didn't say Peppa Pig themed," Bill hissed. "Focus, Bill. There's twenty kids. Twenty presents. Do the math." "That's like a hundred bucks in Paw Patrol trucks?" "Try two thousand in untouched toys we flip on eBay. It's genius." "Is it? Or is it morally bankrupt? Mrs. Henderson will cry. Little Timmy will cry." "Timmy cries when someone breathes near him. He'll live." "What's the plan? We can't just walk past a clown and a bouncy castle." "We wear the clown costumes. We infiltrate. We say the presents need sanitizing for balloon safety." "Oh brilliant. Hi kids, I'm Grumbles the felonious clown. Hand over Spider-Man." "Exactly. You distract with balloon animals. I grab the loot sack." "What if they want swords and I only know how to make existential dread?" "Then make existential dread, Bill. Make it long and twisty." "Remind me why we aren't robbing a bank like normal thieves?" "Because banks have cameras, guards, and vaults. This backyard has a piƱata and terrible parental supervision." "Okay, but if we get caught by a mom with a juice box, I'm blaming you." "Deal. If we get caught, I'll cry louder than Timmy."
created_at
"2026-07-25T17:36:35.24Z"
7/25/2026, 5:36:35 PM
model
"meta/muse-spark-1.1-20260709"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
6432
moderation_latency
(null)
generation_time
8011
tokens_prompt
75
tokens_completion
294
native_tokens_prompt
225
native_tokens_completion
1028
native_tokens_completion_images
0
native_tokens_reasoning
735
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"auto"
usage
0.00465025
router
(null)
provider_responses
0
endpoint_id
"b2b9f6f9-8880-41c1-bd0c-867650fd5238"
id
"resp_6a64f423006960a4c33343bf"
is_byok
false
latency
284
model_permaslug
"meta/muse-spark-1.1-20260709"
provider_name
"Meta"
status
200
user_agent
"langchainjs-openai/1.0.0 ((node/v24.18.0; linux; x64))"
http_referer
(null)
request_id
"req-1785000995-CzRnjvp25kCmooms71XD"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1785000995-GofXiIOcDxGXYAdKfmR2"
upstream_id
"resp_6a64f423006960a4c33343bf"
total_cost
0.00465025
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Meta"
response_cache_source_id
(null)
data_region
"global"
Evaluation details
Result
Evaluator
Details
Meta Data
99.98%
Matches word count
n/a
neededClean
false
words
198
51.96%
Dialogue to Total Word Ratio
Ratio: 98.99%, Deviation: 8.99%
neededClean
false
wordsTotal
199
wordsDialogue
197
75.9736%