NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-07-25-01KYD4WQWSA3HHJ1DB7PD2PB4K
Run Details
Model:
Muse Spark 1.1 (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write unattributed dialogue
Duration:
12.1s
Cost:
$0.008199
Provider:
Meta
Evaluation summary
Score
Evaluator
Details Link
100.00%
Count dialogue tags
(details)
Model input
User/Human
You are an expert novel writer in the cozy mystery genre. Write a scene of about 200 words, with dialogue between two characters. The dialogue should be unattributed, meaning the speaker is not explicitly mentioned. The scene should be set in a small town bakery. Only return the story, nothing else.
Model response
Messages
Meta
AI
The bell over Honeycutt Bakery jingled, bringing in cold air and the smell of wet wool. Rain hammered the awning, fogging the windows and making the kitchen smell even sweeter against the damp. The display case was half empty, and a single lemon tart sat crooked, its powdered sugar disturbed by a fingerprint. "You still have the tart." "Can't sell it now. Evidence, remember?" "Right. Evidence of who can't bake. It's lopsided." "That's not funny. Someone tried to poison the mayor with it." "With a lemon tart? In this town? He'd have done it himself with your coffee." "You saw anyone near it yesterday after three? You were right here frosting those cinnamon rolls." "Only Mrs. Finch. And she kept saying her nephew would inherit everything." "Her nephew is the mayor." "I know." The back door was still unlocked, its draft fluttering the order slips on the counter. "Did you leave it open again?" "You think Mrs. Finch snuck in through the alley? She's eighty." "She runs marathons. And she hates lemon. We should call Deputy Miller before you dust it again."
created_at
"2026-07-25T17:24:15.392Z"
7/25/2026, 5:24:15 PM
model
"meta/muse-spark-1.1-20260709"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
11014
moderation_latency
(null)
generation_time
12063
tokens_prompt
78
tokens_completion
277
native_tokens_prompt
225
native_tokens_completion
1863
native_tokens_completion_images
0
native_tokens_reasoning
1607
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"auto"
usage
0.008199
router
(null)
provider_responses
0
endpoint_id
"b2b9f6f9-8880-41c1-bd0c-867650fd5238"
id
"resp_6a64f13f98e098b4482549f5"
is_byok
false
latency
257
model_permaslug
"meta/muse-spark-1.1-20260709"
provider_name
"Meta"
status
200
user_agent
"langchainjs-openai/1.0.0 ((node/v24.18.0; linux; x64))"
http_referer
(null)
request_id
"req-1785000255-wBVwH8xpWCMez2Xp168i"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1785000255-wlOcPSak8c386wqRlPxz"
upstream_id
"resp_6a64f13f98e098b4482549f5"
total_cost
0.008199
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Meta"
response_cache_source_id
(null)
data_region
"global"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Count dialogue tags
n/a
neededClean
false
dialogueTags
(empty)