NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-03-01M1K508V1PGW11KVZP83T22V5
Run Details
Model:
Muse Spark 1.2 (Reasoning, Medium)
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
23.3s
Cost:
$0.017518
Provider:
Meta
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
100.00%
Dialogue to Total Word Ratio
(details)
100.0000%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
"Okay, so the plan is simple, we walk in dressed as clowns, grab the presents, and walk out." Molly whispered. "Molly, that is the dumbest plan I've ever heard, and I once tried to rob a police station with a water pistol." Bill hissed. "Oh, like you have a better idea, genius?" she snapped. "Yeah, we pretend to be parents, say 'oops, our kid Timmy invited us,' then just take the pile." he said. "Bill, we don't have a kid, we don't even like kids, and the invitation says 'Ava turns six'." Molly groaned. "Exactly, perfect cover, no one questions tired parents at a screaming kids' party, we just look exhausted and grab gifts." Bill grinned. "Fine, but if we have to sing happy birthday, I'm stealing the cake too." she muttered. "Deal, you distract them with clown shoes anyway, I'll stuff the presents in the 'gifts for Timmy' bag." he whispered. "We don't own clown shoes, Bill." Molly sighed. "Then buy some, they're on sale, plus kids love clowns, we blend right in." Bill added. "And if anyone asks, we just say Timmy has severe allergies and needs all presents for treatment?" "Perfect, totally believable, let's go." "Now?" "Now."
created_at
"2026-09-03T08:08:43.112Z"
9/3/2026, 8:08:43 AM
model
"meta/muse-spark-1.2-20260805"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
2262
moderation_latency
(null)
generation_time
23252
tokens_prompt
75
tokens_completion
302
native_tokens_prompt
68
native_tokens_completion
4102
native_tokens_completion_images
0
native_tokens_reasoning
3789
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"auto"
usage
0.0175185
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.19.0; linux; x64))"
http_referer
(null)
request_id
"req-1788422923-jJdVpFntDzDGUjMj8yUa"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1788422923-0Og1cGiEDItrZKFoNy8J"
upstream_id
"resp_6a992b0b92ebb08dd72c4a1b"
provider_responses
0
endpoint_id
"ffc8e76e-808a-44d2-bd21-c6ac16e721cb"
id
"resp_6a992b0b92ebb08dd72c4a1b"
is_byok
false
latency
337
model_permaslug
"meta/muse-spark-1.2-20260805"
provider_name
"Meta"
status
200
total_cost
0.0175185
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Meta"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
words
200
100.00%
Dialogue to Total Word Ratio
Ratio: 90.00%, Deviation: 0.00%
neededClean
false
wordsTotal
200
wordsDialogue
180
100.0000%