NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-07-25-01KYD5RT0V5PVYFFNMA59P5VZB
Run Details
Model:
Muse Spark 1.1 (Reasoning, Minimal)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
11.4s
Cost:
$0.007880
Provider:
Meta
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
99.95%
Dialogue to Total Word Ratio
(details)
99.9742%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly crouched behind the bounce house, glaring at the clown. "Listen, Molly, this is big. Timmy's mom went overboard. That's like three hundred bucks in LEGO alone." "Bill, we are professional art thieves. We stole the Van Gogh last month. Now you want us to mug toddlers for slime?" "Those limited edition Bluey sets resell faster than diamonds. Plus, juice boxes. Unlimited juice boxes." "Molly, focus. The laser grid is actually just silly string. The guard dog is a goldfish. We walk in, grab the pile, vanish." "Fine. But if I have to sing Happy Birthday again, I'm keeping the cake. And I'm not wearing the clown nose." "Deal. You distract with magic tricks, I'll bag the loot. Remember the code word? Banana." "Banana? Bill, last time you picked that, we got attacked by actual monkeys. Let's use cupcake." Bill adjusted his stolen pirate costume impatiently. "Fine, cupcake. You ready to become the most hated people at a child's party since the sugar-free parents?" "Bill, we already are." "Okay, operation Pinata Disaster is a go. On three. One, two..." "Three! And no pushing the birthday boy into the cake this time!" "I made no promises. Just grab the LEGOs."
created_at
"2026-07-25T17:39:35.074Z"
7/25/2026, 5:39:35 PM
model
"meta/muse-spark-1.1-20260709"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
10124
moderation_latency
(null)
generation_time
11327
tokens_prompt
75
tokens_completion
303
native_tokens_prompt
225
native_tokens_completion
1788
native_tokens_completion_images
0
native_tokens_reasoning
1494
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"auto"
usage
0.00788025
router
(null)
provider_responses
0
endpoint_id
"b2b9f6f9-8880-41c1-bd0c-867650fd5238"
id
"resp_6a64f4d763a0e2aca0ff4c26"
is_byok
false
latency
260
model_permaslug
"meta/muse-spark-1.1-20260709"
provider_name
"Meta"
status
200
user_agent
"langchainjs-openai/1.0.0 ((node/v24.18.0; linux; x64))"
http_referer
(null)
request_id
"req-1785001175-kIwxGKikN3kG0PtFIUWC"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1785001175-emtgQpgBGUZ0UIi5eGYX"
upstream_id
"resp_6a64f4d763a0e2aca0ff4c26"
total_cost
0.00788025
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Meta"
response_cache_source_id
(null)
data_region
"global"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
199
99.95%
Dialogue to Total Word Ratio
Ratio: 91.50%, Deviation: 1.50%
neededClean
false
wordsTotal
200
wordsDialogue
183
99.9742%