Research · Grok 4.6 · AA Elo
2026 leaderboards
Last updated: 2026-08-12
Recent GDPval metrics on frontier releases (through Aug 2026)
❓ How have later models—especially Grok 4.6—moved the GDPval scoreboard, and how should those scores relate to the 2025 paper’s human win rates and productivity math?
Three different “GDPval” instruments
Do not collapse these into one number:
| Instrument | Who | What it measures | How graded | Role in this package |
|---|---|---|---|---|
| OpenAI GDPval (paper v1) | OpenAI, Sep–Oct 2025 | Expert deliverables on gold/full sets | Human occupational experts, pairwise vs human gold | Source of quality baseline + Try-1× cost math |
| GDPval-AA v2 | Artificial Analysis | Same task family / dataset lineage, agentic harness | Elo from model-vs-model (and human baseline = 1000) | Live public leaderboard for recent releases |
| Snorkel GDPval+ (related) | Snorkel | Expanded professional tasks / rubrics | Pass rates on expert criteria | Secondary signal; not identical to paper w |
Hard rule: AA Elo and OpenAI’s 39% GPT-5 win rate are not interchangeable. Elo 1750 does not mean “175% of human.” On AA’s scale, Human Baseline = 1000; scores above 1000 mean the system beats that baseline under AA’s pairwise/agent setup. The paper’s w is the fraction of times a model deliverable beat a specific human gold under expert blinding.
GDPval-AA v2 snapshot (observed 2026-08-12)
Source: Artificial Analysis leaderboard page and Grok 4.6 analysis article (same day as Grok 4.6 public push). Leaderboard table Elo (high-effort configs where listed):
| Rank | Model (config) | Creator | GDPval-AA v2 Elo | Release (AA) |
|---|---|---|---|---|
| 1 | Claude Opus 5 (Adaptive, Max Effort) | Anthropic | 1849 | Jul 2026 |
| 2 | Claude Opus 5 (Adaptive, Xhigh) | Anthropic | 1817 | Jul 2026 |
| 3 | Grok 4.6 (high) | SpaceXAI / xAI stack | 1749 (leaderboard) / 1753 (AA article & xAI table) | Aug 2026 |
| 4 | Claude Fable 5 (Max, Opus 4.8 fallback) | Anthropic | 1741 | Jun 2026 |
| 5 | Qwen3.8 Max | Alibaba | 1737 | Aug 2026 |
| 7 | GPT-5.6 Sol (max) | OpenAI | 1728 | Jul 2026 |
| 21 | Grok 4.5 (high) | SpaceXAI | 1526 | Jul 2026 |
| — | Human Baseline | — | 1000 | — |
| 85 | GPT-5 (high) | OpenAI | 1083 | Aug 2025 |
Slight Elo disagreements (1749 vs 1753 for Grok 4.6) appear between the live table and the same-day analysis/vendor table—report both; treat as ~1750 class.
Grok 4.6 in focus
Confirmed (primary / AA):
- Public release framing: 12 Aug 2026 (x.ai/news/grok-4-6); marketed for long-running agents and interactive/visual work artifacts.
- GDPval-AA v2: ~1753 Elo in xAI’s published comparison table and AA’s analysis writeup; #3 on the live AA v2 table behind two Opus 5 effort settings.
- Jump vs Grok 4.5 high: 1526 → ~1753 (~+227 Elo on AA’s scale in one generation).
- Beats GPT-5.6 Sol max (1728) and Claude Fable 5 max (1741) on the xAI/AA comparison rows used at launch (Fable still within noise on some CI statements).
- AA also places Grok 4.6 on cost–performance Pareto for agentic work: headline $2 / $6 per M input/output tokens (unchanged from 4.5); AA measured ~$0.84 per GDPval-AA task in their harness—far below Opus 5 / Sol list prices they cite ($5/$25 and $5/$30).
- AA-Briefcase (private long-horizon knowledge-work Elo): Grok 4.6 1577 vs Grok 4.5 1313, Fable 5 1574, GPT-5.6 Sol max 1502; AA notes ~53 turns / ~0.5B input tokens avg vs ~103 turns / ~2.0B for Opus 5 max—efficiency story, not only peak score.
- Original OpenAI paper (2025) already included Grok 4 in the human pairwise set; it was not the quality leader then (Claude Opus 4.1 led aesthetics; GPT-5 led accuracy). The 4.5→4.6 arc is a later agentic-leaderboard story.
Credible but protocol-bound:
- AA’s claim that Grok 4.6 is “behind only Claude Opus 5” on GDPval-AA with CIs overlapping Fable/Qwen—true inside AA’s harness, not a re-run of OpenAI’s 2025 human gold pairwise study.
- Vendor-collected rows on multi-bench tables can differ from AA’s own reruns; prefer AA page + AA article over secondary roundups when they conflict.
Weak / do not use for ROI math:
- Social posts claiming Grok 4.6 is “#1 on GDPval” without naming AA Elo vs human pairwise.
- Any conversion that turns Elo 1753 into a Drop-in replacement for paper w = 0.39 without a new human study.
What this does to the productivity story
The 2025 paper’s Try-1× math still uses human win rate w and gold-set \(H_T, R_T\). GDPval-AA Elo is a leading indicator that agentic quality on the same task family kept climbing through mid-2026:
- Directionally, models far above Human Baseline 1000 (Opus 5, Grok 4.6, Sol max) are the class of systems for which oversight-compatible ROI is more plausible, not less—higher effective ship rates shrink the (1−w)·HT redo term.
- You still cannot skip local w and RT. Leaderboard Elo does not measure your review checklist minutes or your occupation mix.
- Cost per task on the leaderboard (AA’s $0.84 for Grok 4.6) sits in the same order of magnitude as the paper’s sub-dollar API costs—still rounding error next to ~$361 human labor proxies. The binding enterprise cost remains reviewer time, not tokens.
- Harness still matters. AA and xAI emphasize multi-turn agents, tool use, and self-checking. That reinforces the paper’s scaffolding result: chat-only deployments under-sample GDPval-class capability.
- Snorkel’s GDPval+ caution: even strong 2026 models can show mean pass rates under one-third of expert rubric criteria on harder expanded sets (e.g. Grok 4.5 at ~29% mean pass on ~2,000 GDPval+ tasks in Snorkel’s July 2026 writeup). “Leaderboard Elo high” and “passes every expert criterion” are different bars. Reviewers stay mandatory.
Trend update
Trend: Public GDPval-family scores for frontier models continued rising from late 2025 through Aug 2026 on third-party agentic harnesses.
Pattern: Stepwise generational jumps (e.g. Grok 4.5→4.6 +200 Elo) plus dense effort-tier ladders (Opus 5 low→max).
Classification: Too data-poor to call exponential on Elo (short window, changing harnesses); best label is rapid stepwise climb on agentic preference/Elo metrics, with Anthropic holding peak Elo and SpaceXAI/xAI reclaiming Pareto cost-efficiency on the same family.
Falsifier: A year of flat AA top-10 Elo with stable harness, or independent human pairwise studies showing win rates stuck near 2025 GPT-5 levels despite Elo gains.
Implications for decision-makers
- When a vendor waves “#1 GDPval,” ask: which instrument, which config, Elo or human w, cost per task, turns/tokens?
- For bakeoffs in 2026, include at least one Opus-class, one Grok 4.6-class, and one Sol-class system under the same internal harness and reviewer rubric—do not import AA rank order as procurement truth.
- Recompute Try-1× with your measured w; use leaderboard only to choose candidates likely to clear the ~28% break-even under paper-like review costs.