NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-09-29-01M3P2F1C4FYQWH1RGS6M0V1KF
Run Details
Model:
Claude Sonnet 5.5 (Reasoning)
Test:
N-Length Sentences
Scenario:
Write sentences with 20 words each
Duration:
14.2s
Cost:
$0.019074
Provider:
Claude Platform on AWS
Evaluation summary
Score
Evaluator
Details Link
83.33%
Matches word count
(details)
Model input
User/Human
Please write the opening scene of a murder mystery where each sentence has exactly 20 words. Please make sure that each sentence has the same length. Only return the opening itself.
Model response
Messages
Meta
AI
The fog rolled over Harrowgate Manor like a burial shroud, and somewhere inside, a clock struck midnight with grim finality. Inspector Lena Marsh stepped from her car, pulled her collar tight, and studied the crooked silhouette of the sleeping house. Constable Pike met her at the door, his face pale as candle wax beneath the dripping porch lamp's yellow glow. "Lord Aldous Harrowgate is dead, ma'am," he whispered, "found in the locked library with a silver dagger in his back." Lena frowned, because locked rooms rarely happened by accident, and murderers rarely trusted anyone else to keep their dark secrets. Inside the drawing room, six guests huddled near the dying fire, each one watching the others with glittering, suspicious eyes. One of them had killed tonight, Lena realized, and before the sun rose, she would surely learn exactly which one.
created_at
"2026-09-29T07:53:28.204Z"
9/29/2026, 7:53:28 AM
model
"anthropic/claude-sonnet-5.5-20260928"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
3760
moderation_latency
(null)
generation_time
14204
tokens_prompt
46
tokens_completion
457
native_tokens_prompt
62
native_tokens_completion
1895
native_tokens_completion_images
(null)
native_tokens_reasoning
1592
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"end_turn"
service_tier
"default"
usage
0.019074
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.21.0; linux; x64))"
http_referer
(null)
request_id
"req-1790668408-iX0kruXigdamvOKAq4FR"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1790668408-Svp3fRcRDp5zD8BRcJB0"
upstream_id
"msg_011CfXNNQBY96XfpgpweGNGh"
provider_responses
0
endpoint_id
"99aaad94-923b-4fc1-b763-271ed5486f7a"
id
"msg_011CfXNNQBY96XfpgpweGNGh"
is_byok
false
latency
606
model_permaslug
"anthropic/claude-sonnet-5.5-20260928"
provider_name
"Claude Platform on AWS"
status
200
total_cost
0.019074
cache_discount
(null)
upstream_inference_cost
0
provider_name
"Claude Platform on AWS"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
83.33%
Matches word count
n/a
neededClean
false
sentences
6
wordCounts
0
20
1
20
2
20
3
40
4
20
5
20