NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-07-25-01KYD6KNVR4HGVB95E258HVHBM
Run Details
Model:
Muse Spark 1.1 (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 200 words with 50% dialogue
Duration:
21.4s
Cost:
$0.016691
Provider:
Meta
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
99.9997%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 50% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly spread the blueprint across the van dashboard, revealing a bouncy castle, a clown, and a mountain of brightly wrapped boxes. "Bill, this is the score. Twenty kids, zero security, all presents piled by the cake," Molly whispered. Bill stared at the crayon drawing taped beside it. "You stole a kid's treasure map? This is low, even for us," he said. "It's not low, it's logistics. The parents will be distracted by pin the tail," she replied. He sighed, adjusting his ski mask. "So we crash a six-year-old's birthday for LEGO and gift cards?" he asked. "Fifteen gift cards, plus a Nintendo Switch. Timmy's mom overshares on Facebook," she grinned. Molly folded the blueprint decisively. The house smelled faintly of frosting and desperation down the street. "Fine, but if I get stuck in the ball pit, you leave me," Bill muttered. "Deal. You distract the clown, I grab the loot," Molly said. The wind rattled the van doors as children's laughter echoed from the cul-de-sac. Bill pulled his mask down, leaving his nervous mustache exposed. Molly checked her watch, counting down until cake time. "Remember, we split sixty-forty," she added. "Sixty-forty? I'm risking clown trauma here!" "Worth it." "Trust me."
created_at
"2026-07-25T17:54:15.552Z"
7/25/2026, 5:54:15 PM
model
"meta/muse-spark-1.1-20260709"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
20074
moderation_latency
(null)
generation_time
20876
tokens_prompt
75
tokens_completion
312
native_tokens_prompt
225
native_tokens_completion
3861
native_tokens_completion_images
0
native_tokens_reasoning
3559
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"auto"
usage
0.0166905
router
(null)
provider_responses
0
endpoint_id
"b2b9f6f9-8880-41c1-bd0c-867650fd5238"
id
"resp_6a64f8473d4e470375cc4020"
is_byok
false
latency
361
model_permaslug
"meta/muse-spark-1.1-20260709"
provider_name
"Meta"
status
200
user_agent
"langchainjs-openai/1.0.0 ((node/v24.18.0; linux; x64))"
http_referer
(null)
request_id
"req-1785002055-cT6QGqp439vukFXDI83p"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1785002055-MF78vK6fcEDLhvwoM27k"
upstream_id
"resp_6a64f8473d4e470375cc4020"
total_cost
0.0166905
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Meta"
response_cache_source_id
(null)
data_region
"global"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
100.00%
Dialogue to Total Word Ratio
Ratio: 50.49%, Deviation: 0.49%
neededClean
false
wordsTotal
206
wordsDialogue
104
99.9997%