NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-07-25-01KYD5A84F7VKHVM5XGWZ3B9BJ
Run Details
Model:
Laguna XS 2.1
Test:
Dialogue tags
Scenario:
Write 200 words with 90% dialogue
Duration:
3.8s
Cost:
$0.000085
Provider:
Poolside
Evaluation summary
Score
Evaluator
Details Link
93.65%
Matches word count
(details)
84.27%
Dialogue to Total Word Ratio
(details)
88.9598%
Overall
Model input
User/Human
You are a author of comedic heist books. Please write a 200 word scene that has 90% words of dialogue between two thieves named Molly and Bill who are planning to break into a children's party to steal all the presents. Only include the prose of the scene without any markdown or commentary.
Model response
Messages
Meta
AI
Molly leaned against a graffiti-covered dumpster, her voice a conspiratorial whisper. "Okay, Bill, listen up. The Smiths’ backyard bash starts at three. We hit it at three-fifteen." Bill adjusted his hat, squinting at her. "Three-fifteen? What, we’re waiting for the sugar rush to peak?" "Exactly. Parents are busy chasing toddlers, and the presents are in the garage—locked, but not *baby-proofed*. I’ve got a friend with a lockpick set." "A friend who’s also a convicted felon?" Bill raised an eyebrow. "Details, Bill. Details." Molly pulled out a crumpled map. "Garage door code is 4-7-2. I hacked their smart lock app. Easy." Bill snorted. "You *hacked* a smart lock? What’s next, a drone to distract the dog?" "Genius, actually. I’ll deploy the drone—call it ‘Sir Barksalot’—while I slip in. You create a distraction at the gate. Maybe dressed as a delivery guy?" "A *delivery guy* who’s clearly not holding a pizza?" "Exactly! They’ll assume it’s a mix-up. Parents are confused, kids are screaming, and we’re in and out." "And the presents? You can’t just grab a sack and run." Molly grinned. "We’ll use the garage’s utility truck. Fake a breakdown, load up, and vanish before anyone notices." Bill sighed. "And if we get caught?" "Then we blame the dog."
created_at
"2026-07-25T17:31:38.009Z"
7/25/2026, 5:31:38 PM
model
"poolside/laguna-xs-2.1-20260625"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
75
moderation_latency
(null)
generation_time
2794
tokens_prompt
75
tokens_completion
718
native_tokens_prompt
81
native_tokens_completion
673
native_tokens_completion_images
(null)
native_tokens_reasoning
323
native_tokens_cached
16
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"stop"
service_tier
(null)
usage
0.00008514
router
(null)
provider_responses
0
endpoint_id
"089ec178-f6dc-4450-aff5-46f68798ce97"
id
"chatcmpl-4a850408888c4b90b9c96edca99237bd"
is_byok
false
latency
75
model_permaslug
"poolside/laguna-xs-2.1-20260625"
provider_name
"Poolside"
status
200
user_agent
"langchainjs-openai/1.0.0 ((node/v24.18.0; linux; x64))"
http_referer
(null)
request_id
"req-1785000698-3Khdvg0AuzlKIk0iPKoU"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1785000698-T3rxExA2QXwzKYOO33Uw"
upstream_id
"chatcmpl-4a850408888c4b90b9c96edca99237bd"
total_cost
0.00008514
cache_discount
8e-7
upstream_inference_cost
0
provider_name
"Poolside"
response_cache_source_id
(null)
data_region
"global"
Evaluation details
Result
Evaluator
Details
Meta Data
93.65%
Matches word count
n/a
neededClean
false
words
209
84.27%
Dialogue to Total Word Ratio
Ratio: 83.57%, Deviation: 6.43%
neededClean
false
wordsTotal
213
wordsDialogue
178
88.9598%