NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-03-01M1K5CEE8C79HPG3578BAJXKA
Run Details
Model:
Muse Spark 1.2 (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
15.0s
Cost:
$0.022474
Provider:
Meta
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
99.9999%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Bill and Molly crouched behind the garden fence watching. "Molly, the plan is simple, we walk in during cake time," "Simple? Bill, we are robbing a birthday party," "We are not robbing, we are redistributing unopened plastic dinosaurs," "You mean stealing Bluey dolls and juice boxes for profit?" "Exactly, low security, high reward, plus free cake for us," "Security is fifteen sugar crazed kids and a terrifying clown," "That clown is the vault guard, I fear him," "Everyone fears the clown, that is why we need disguises," "I brought my old Spider-Man costume, it still fits mostly," "Bill, you are forty, you cannot wear Spider-Man anymore," "The kids will think I am cool dad energy, trust me," "You are creepy uncle energy, wear the magician cape," "Fine, but I get the giant present with the roaring dinosaur," "Only if you distract the mom with enthusiastic small talk," "I can do that, I love talking about Paw Patrol lore," "Bill, you have never seen Paw Patrol in your life," "I have seen one episode, Chase is the police dog, right?" "Wow, you will definitely fool five year olds then," They stared at the bounce house in tense silence together now anxiously.
created_at
"2026-09-03T08:15:22.068Z"
9/3/2026, 8:15:22 AM
model
"meta/muse-spark-1.2-20260805"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
3661
moderation_latency
(null)
generation_time
14919
tokens_prompt
75
tokens_completion
408
native_tokens_prompt
68
native_tokens_completion
5268
native_tokens_completion_images
0
native_tokens_reasoning
4880
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"auto"
usage
0.022474
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.19.0; linux; x64))"
http_referer
(null)
request_id
"req-1788423322-ICpnKX5FOJmMR0UgeWUs"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1788423322-IWD1loQDPVL6sbyZzCG7"
upstream_id
"resp_6a992c9a5d74d37edc6242c6"
provider_responses
0
endpoint_id
"ffc8e76e-808a-44d2-bd21-c6ac16e721cb"
id
"resp_6a992c9a5d74d37edc6242c6"
is_byok
false
latency
177
model_permaslug
"meta/muse-spark-1.2-20260805"
provider_name
"Meta"
status
200
total_cost
0.022474
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Meta"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
100.00%
Dialogue to Total Word Ratio
Ratio: 89.60%, Deviation: 0.40%
neededClean
false
wordsTotal
202
wordsDialogue
181
99.9999%