Vendors

Model creators/vendors and how their models compare across the benchmark.

Vendor Models Avg Score Best Score ▼Best Model
OpenAI3385.76%95.15%GPT-5.6 Sol (Reasoning)
Anthropic1990.25%95.06%Claude Opus 4.6 (Reasoning)
Google2484.94%94.90%Gemini 3.6 Flash (Reasoning)
Qwen1686.77%94.55%Qwen3.7 Max
xAI687.69%94.12%Grok 4.5 (Reasoning, High)
meta292.62%93.81%Muse Spark 1.1 (Reasoning, Medium)
Z.AI988.14%93.74%Z.AI GLM 5.1
MoonshotAI492.12%92.99%MoonshotAI: Kimi K3 (Reasoning, High)
thinkingmachines282.37%90.78%Thinking Machines Inkling (Reasoning)
minimax387.80%90.45%MiniMax M3
bytedance-seed482.62%89.59%ByteDance Seed 1.6
DeepSeek983.63%89.28%DeepSeek V4 Pro (Reasoning)
aion-labs386.38%88.78%Aion 3.0
xiaomi285.00%86.05%Xiaomi MIMO v2.5 Pro
Mistral AI1272.18%84.29%Mistral Large 3
inception181.99%81.99%Inception Mercury 2
NVIDIA278.10%81.69%Nemotron 3 Super
Nous Research275.27%80.80%Hermes 3 405B
poolside277.70%79.98%Laguna S 2.1
Writer178.11%78.11%Writer: Palmyra X5
Meta177.41%77.41%Llama 3.1 70B
TheDrummer172.68%72.68%Cydonia 24B V4.1
Microsoft171.45%71.45%WizardLM 2 8x22b
arcee-ai167.68%67.68%Arcee AI: Trinity Mini
Cohere167.04%67.04%Cohere Command R+ (Aug. 2024)