NC Bench
Overview
About
Tests
Categories
Models
Vendors
Compare
Run: 2026-10-05-01M46B7JAGNF7SWA669W8EGVB9
Run Details
Model:
GPT-6.1 Sol (Reasoning, Medium)
Test:
N-Length Sentences
Scenario:
Write sentences with 20 words each
Duration:
19.0s
Cost:
$0.007912
Provider:
OpenAI
Evaluation summary
Score
Evaluator
Details Link
100.00%
Matches word count
(details)
Model input
User/Human
Please write the opening scene of a murder mystery where each sentence has exactly 20 words. Please make sure that each sentence has the same length. Only return the opening itself.
Model response
Messages
Meta
AI
Rain hammered the windows of Blackthorn Manor when Inspector Vale arrived, carrying a telegram that smelled faintly of burnt roses. The butler opened the door before Vale knocked, his white gloves stained with something darker than the evening could explain. Upstairs, a woman screamed once, then stopped so abruptly that the silence seemed less like peace than another violent act. Lord Blackthorn lay across his study desk, a silver letter opener beneath his hand and fresh snow upon his shoes. There had been no snow for weeks, and every window in the study was locked firmly against the autumn rain. Beside the body stood six untouched glasses, though the telegram had insisted that only five guests would attend tonight's dinner. Vale unfolded the message again, noticing now that its final line was written in ink unlike anything above the signature. It read, with dreadful simplicity, that the inspector must arrive before midnight, unless he wished to become the seventh guest.
created_at
"2026-10-05T15:34:31.519Z"
10/5/2026, 3:34:31 PM
model
"openai/gpt-6.1-sol-20260929"
app_id
182717
external_user
(null)
streamed
true
cancelled
false
latency
10128
moderation_latency
(null)
generation_time
18988
tokens_prompt
46
tokens_completion
829
native_tokens_prompt
41
native_tokens_completion
783
native_tokens_completion_images
0
native_tokens_reasoning
591
native_tokens_cached
0
num_media_prompt
(null)
num_input_audio_prompt
(null)
num_media_completion
0
num_search_results
(null)
num_fetches
(null)
web_search_engine
(null)
origin
"https://ncbench.com/"
is_byok
false
finish_reason
"stop"
native_finish_reason
"completed"
service_tier
"default"
usage
0.007912
router
(null)
user_agent
"langchainjs-openai/1.0.0 ((node/v24.21.0; linux; x64))"
http_referer
(null)
request_id
"req-1791214471-1CwnH2LzssSFErUdtP6B"
session_id
(null)
preset_id
(null)
api_type
"completions"
id
"gen-1791214471-yZjqQrhd1iuAFJjfUiIR"
upstream_id
"resp_07fd08d983739d76016ac3c387a59887d19c962b0d82f502ce"
provider_responses
0
endpoint_id
"dfb9d5c3-63f9-4263-8b78-4fa97dcb3d5e"
id
"resp_07fd08d983739d76016ac3c387a59887d19c962b0d82f502ce"
is_byok
false
latency
373
model_permaslug
"openai/gpt-6.1-sol-20260929"
provider_name
"OpenAI"
status
200
total_cost
0.007912
cache_discount
(null)
upstream_inference_cost
0
provider_name
"OpenAI"
response_cache_source_id
(null)
data_region
"global"
workspace_id
"97e315e5-d303-487d-83c1-83180e8a13d4"
Evaluation details
Result
Evaluator
Details
Meta Data
100.00%
Matches word count
n/a
neededClean
false
sentences
8
wordCounts
0
20
1
20
2
20
3
20
4
20
5
20
6
20
7
20