xAI
Comparing 8 models from xAI.
| Model | Total ▼ | Released | Context | CoT | Tooling | Creative Writing | Language | Utility | Reasoning | Text Editing | Rule Following | Hallucination |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Grok 4.5 (Reasoning, High) | 94.12% | Jul 8, 26 | 500k | ✓ | 99.99% | 87.43% | 99.82% | 86.69% | 94.45% | 98.73% | 86.79% | 99.06% |
| Grok 4.6 (Reasoning, High) | 92.76% | Aug 12, 26 | 500k | ✓ | 100.00% | 88.97% | 95.04% | 83.88% | 95.61% | 99.17% | 81.19% | 98.25% |
| Grok 4.7 (Reasoning, High) | 92.62% | Sep 21, 26 | 500k | ✓ | 100.00% | 91.29% | 97.50% | 91.12% | 96.71% | 98.17% | 70.68% | 95.48% |
| Grok 4.3 (Reasoning) | 90.99% | Apr 30, 26 | 1m | ✓ | 98.47% | 85.11% | 97.50% | 92.94% | 75.64% | 97.64% | 82.80% | 97.86% |
| Grok 4.5 (Reasoning, Low) | 90.94% | Jul 8, 26 | 500k | ✓ | 99.33% | 87.45% | 94.50% | 82.63% | 89.59% | 98.54% | 76.44% | 99.01% |
| Grok 4.20 (Reasoning) | 90.87% | Mar 31, 26 | 2m | ✓ | 100.00% | 86.25% | 96.61% | 92.61% | 74.28% | 98.83% | 82.04% | 96.30% |
| Grok 4.20 | 81.21% | Mar 31, 26 | 2m | – | 95.00% | 83.44% | 78.86% | 84.11% | 72.46% | 95.63% | 59.71% | 80.48% |
| Grok 4.3 | 78.00% | Apr 30, 26 | 1m | – | 91.21% | 84.51% | 84.74% | 66.41% | 70.17% | 90.19% | 49.02% | 87.79% |
Model Performance
Cost vs Performance
Compares total benchmark cost against overall score for xAI models. Quadrant lines are drawn at the median values.
2 low-scoring outliers hidden: Grok 4.20 (81.2%), Grok 4.3 (78.0%).
Cost Breakdown
Total benchmark cost per model, broken down by input, reasoning, and output tokens. Toggle between USD and token views.