NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-07-25-01KYD6JJ2N7Y5VECHKRTW6Y8PH
Run Details
Model:
Muse Spark 1.1 (Reasoning, Minimal)
Test:
Dialogue tags
Scenario:
Write 200 words with 10% dialogue
Duration:
17.1s
Cost:
$0.010116
Provider:
Meta
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
99.9981%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 10% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Bill spread the hand-drawn map across the hood of their getaway van, which still had baby stickers from their last failed job. He pointed a greasy finger at the bouncy castle. Molly stared, arms crossed, wearing a stolen magician's cape that was three sizes too large. The plan was ludicrous even by their standards. They had robbed banks, museums, and once a very aggressive bake sale, but a suburban backyard birthday party felt like a new low. The target house was heavily fortified by juice boxes, screaming toddlers, and one terrifying mother with a clipboard. The presents were piled high beside the trampoline, guarded only by a sleeping golden retriever. Bill had brought grappling hooks, smoke bombs, and a gift bag disguise. Molly had brought morals, barely. "We are stealing from six-year-olds now?" "Think bigger, it's about the piƱata." "Fine, but I'm not wearing clown shoes again." Their escape route involved tripping over the sprinkler, pretending to be entertainers, and sprinting while carrying what was hopefully not just socks and educational board games. Molly sighed, adjusting her cape, wondering when they had become the villains in a very low-budget cartoon. For once she genuinely missed robbing boring old banks together.
created_at
"2026-07-25T17:53:38.907Z"
7/25/2026, 5:53:38 PM
model
"meta/muse-spark-1.1-20260709"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
15853
moderation_latency
(null)
generation_time
16990
tokens_prompt
75
tokens_completion
320
native_tokens_prompt
225
native_tokens_completion
2314
native_tokens_completion_images
0
native_tokens_reasoning
2030
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"auto"
usage
0.01011575
router
(null)
provider_responses
0
endpoint_id
"b2b9f6f9-8880-41c1-bd0c-867650fd5238"
id
"resp_6a64f82636f70eaa98044a3a"
is_byok
false
latency
3239
model_permaslug
"meta/muse-spark-1.1-20260709"
provider_name
"Meta"
status
200
user_agent
"langchainjs-openai/1.0.0 ((node/v24.18.0; linux; x64))"
http_referer
(null)
request_id
"req-1785002018-UJDItIALt562CAzMTbxI"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1785002018-C5PuX6u4BNJo7BdEcwkC"
upstream_id
"resp_6a64f82636f70eaa98044a3a"
total_cost
0.01011575
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Meta"
response_cache_source_id
(null)
data_region
"global"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
100.00%
Dialogue to Total Word Ratio
Ratio: 10.78%, Deviation: 0.78%
neededClean
false
wordsTotal
204
wordsDialogue
22
99.9981%