xAI

Comparing 8 models from xAI.

Model Total ▼ Released Context CoTTooling Creative Writing Language Utility Reasoning Text Editing Rule Following Hallucination
Grok 4.5 (Reasoning, High)94.12%Jul 8, 26500k✓99.99%87.43%99.82%86.69%94.45%98.73%86.79%99.06%
Grok 4.6 (Reasoning, High)92.76%Aug 12, 26500k✓100.00%88.97%95.04%83.88%95.61%99.17%81.19%98.25%
Grok 4.7 (Reasoning, High)92.62%Sep 21, 26500k✓100.00%91.29%97.50%91.12%96.71%98.17%70.68%95.48%
Grok 4.3 (Reasoning)90.99%Apr 30, 261m✓98.47%85.11%97.50%92.94%75.64%97.64%82.80%97.86%
Grok 4.5 (Reasoning, Low)90.94%Jul 8, 26500k✓99.33%87.45%94.50%82.63%89.59%98.54%76.44%99.01%
Grok 4.20 (Reasoning)90.87%Mar 31, 262m✓100.00%86.25%96.61%92.61%74.28%98.83%82.04%96.30%
Grok 4.2081.21%Mar 31, 262m–95.00%83.44%78.86%84.11%72.46%95.63%59.71%80.48%
Grok 4.378.00%Apr 30, 261m–91.21%84.51%84.74%66.41%70.17%90.19%49.02%87.79%
Model Performance
Cost vs Performance

Compares total benchmark cost against overall score for xAI models. Quadrant lines are drawn at the median values.

2 low-scoring outliers hidden: Grok 4.20 (81.2%), Grok 4.3 (78.0%).

Cost Breakdown

Total benchmark cost per model, broken down by input, reasoning, and output tokens. Toggle between USD and token views.