NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-03-01M1K4TX92V2DKZZR7J6YXZMPN
Run Details
Model:
Muse Spark 1.2 (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
20.0s
Cost:
$0.013379
Provider:
Meta
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
99.84%
Dialogue to Total Word Ratio
(details)
99.9201%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly and Bill crouched behind the bouncy castle watching the clown. "We do this quick, Bill. In and out before cake." "Quick? Molly, there are twenty kids hopped up on sugar and juice." "Perfect distraction. Parents are too busy stopping them from dying." "And the presents? They're all Paw Patrol and slime." "Slime sells, Bill. Plus that one big box is definitely a bike." "How do we get past the birthday mom? She looks lethal." "We dress as entertainers. You be the magician." "I hate kids. They see through my rabbit trick." "You'll be fine. Just pull the bike out of your hat." "What if they cry?" "Then we cry too and steal faster." "Okay, plan: you distract clown, I grab gifts, we run." "Why do I always distract the scary clown?" "Because you're great with scary people. Remember the bank guard?" "That was different. He wasn't wearing giant shoes and honking." "Trust me, this is easier. Kids love magic until you steal their loot." "What about the parents filming everything?" "We smile, wave, then vanish like responsible adults." "Fine. But I want the bike." "You can have the bike if I get the slime." He sighed, adjusting his mustache.
created_at
"2026-09-03T08:05:47.433Z"
9/3/2026, 8:05:47 AM
model
"meta/muse-spark-1.2-20260805"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
2913
moderation_latency
(null)
generation_time
19990
tokens_prompt
75
tokens_completion
310
native_tokens_prompt
68
native_tokens_completion
3128
native_tokens_completion_images
0
native_tokens_reasoning
2831
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"auto"
usage
0.013379
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.19.0; linux; x64))"
http_referer
(null)
request_id
"req-1788422747-1qP29naHBQCVRNyDqioY"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1788422747-6uAAhqbST7R2zY2psAYw"
upstream_id
"resp_6a992a5b327b9e1e22264a6f"
provider_responses
0
endpoint_id
"ffc8e76e-808a-44d2-bd21-c6ac16e721cb"
id
"resp_6a992a5b327b9e1e22264a6f"
is_byok
false
latency
361
model_permaslug
"meta/muse-spark-1.2-20260805"
provider_name
"Meta"
status
200
total_cost
0.013379
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Meta"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
99.84%
Dialogue to Total Word Ratio
Ratio: 92.00%, Deviation: 2.00%
neededClean
false
wordsTotal
200
wordsDialogue
184
99.9201%