NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-07-25-01KYD7H4HV1QCYDCGQGGSRCZCP
Run Details
Model:
Muse Spark 1.1 (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write unattributed dialogue
Duration:
9.6s
Cost:
$0.006599
Provider:
Meta
Evaluation summary
Score
Evaluator
Details Link
100.00%
Count dialogue tags
(details)
Model input
User/Human
You are an expert novel writer in the cozy mystery genre. Write a scene of about 200 words, with dialogue between two characters. The dialogue should be unattributed, meaning the speaker is not explicitly mentioned. The scene should be set in a small town bakery. Only return the story, nothing else.
Model response
Messages
Meta
AI
The bell over Honeycomb Bakery jingled, letting in a gust of October air and the smell of wet leaves. The old radiator hissed behind the display case full of crooked apple danishes. "You're two hours late for the cinnamon rolls. I saved you one, though it might have gone a bit stale out of spite." "I was busy solving a murder, thank you very much. Stale spite-roll is perfect." "The only thing murdered here this morning was my sourdough starter. Look, it's bubbling like it's got a grudge." "Did Deputy Frank finally dust for prints on Mrs. Abernathy's missing cat trophy?" "He dusted the whole counter. Now everything tastes like fingerprint powder and regret. Here, taste this." "Is that... lavender?" "It was supposed to be vanilla. I got the jars mixed up after you left your true-crime podcast blasting all night." "Well, the victim in episode three was poisoned with lavender scones. Very on brand." "Wonderful. So if anyone else dies in this town, I'm officially a suspect." "Only if you keep baking like this. On second thought, better make it two rolls. I'll need the energy for the autopsy report."
created_at
"2026-07-25T18:10:20.875Z"
7/25/2026, 6:10:20 PM
model
"meta/muse-spark-1.1-20260709"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
8142
moderation_latency
(null)
generation_time
9420
tokens_prompt
78
tokens_completion
286
native_tokens_prompt
225
native_tokens_completion
1503
native_tokens_completion_images
0
native_tokens_reasoning
1239
native_tokens_cached
64
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"auto"
usage
0.0065986
router
(null)
provider_responses
0
endpoint_id
"b2b9f6f9-8880-41c1-bd0c-867650fd5238"
id
"resp_6a64fc0da0e3d53f982049fb"
is_byok
false
latency
360
model_permaslug
"meta/muse-spark-1.1-20260709"
provider_name
"Meta"
status
200
user_agent
"langchainjs-openai/1.0.0 ((node/v24.18.0; linux; x64))"
http_referer
(null)
request_id
"req-1785003020-TSCwf1wBdyCHyUYBXwDr"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1785003020-GJ0nUFtdC41QPVvqzm75"
upstream_id
"resp_6a64fc0da0e3d53f982049fb"
total_cost
0.0065986
cache_discount
0.0000704
upstream_inference_cost
0
provider_name
"Meta"
response_cache_source_id
(null)
data_region
"global"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Count dialogue tags
n/a
neededClean
false
dialogueTags
(empty)