Research · Grok 4.6 · AA Elo

2026 leaderboards

Last updated: 2026-08-12

Recent GDPval metrics on frontier releases (through Aug 2026)

How have later models—especially Grok 4.6—moved the GDPval scoreboard, and how should those scores relate to the 2025 paper’s human win rates and productivity math?

Three different “GDPval” instruments

Do not collapse these into one number:

Instrument Who What it measures How graded Role in this package
OpenAI GDPval (paper v1) OpenAI, Sep–Oct 2025 Expert deliverables on gold/full sets Human occupational experts, pairwise vs human gold Source of quality baseline + Try-1× cost math
GDPval-AA v2 Artificial Analysis Same task family / dataset lineage, agentic harness Elo from model-vs-model (and human baseline = 1000) Live public leaderboard for recent releases
Snorkel GDPval+ (related) Snorkel Expanded professional tasks / rubrics Pass rates on expert criteria Secondary signal; not identical to paper w

Hard rule: AA Elo and OpenAI’s 39% GPT-5 win rate are not interchangeable. Elo 1750 does not mean “175% of human.” On AA’s scale, Human Baseline = 1000; scores above 1000 mean the system beats that baseline under AA’s pairwise/agent setup. The paper’s w is the fraction of times a model deliverable beat a specific human gold under expert blinding.

GDPval-AA v2 snapshot (observed 2026-08-12)

Source: Artificial Analysis leaderboard page and Grok 4.6 analysis article (same day as Grok 4.6 public push). Leaderboard table Elo (high-effort configs where listed):

Rank Model (config) Creator GDPval-AA v2 Elo Release (AA)
1 Claude Opus 5 (Adaptive, Max Effort) Anthropic 1849 Jul 2026
2 Claude Opus 5 (Adaptive, Xhigh) Anthropic 1817 Jul 2026
3 Grok 4.6 (high) SpaceXAI / xAI stack 1749 (leaderboard) / 1753 (AA article & xAI table) Aug 2026
4 Claude Fable 5 (Max, Opus 4.8 fallback) Anthropic 1741 Jun 2026
5 Qwen3.8 Max Alibaba 1737 Aug 2026
7 GPT-5.6 Sol (max) OpenAI 1728 Jul 2026
21 Grok 4.5 (high) SpaceXAI 1526 Jul 2026
Human Baseline 1000
85 GPT-5 (high) OpenAI 1083 Aug 2025

Slight Elo disagreements (1749 vs 1753 for Grok 4.6) appear between the live table and the same-day analysis/vendor table—report both; treat as ~1750 class.

Grok 4.6 in focus

Confirmed (primary / AA):

Credible but protocol-bound:

Weak / do not use for ROI math:

What this does to the productivity story

The 2025 paper’s Try-1× math still uses human win rate w and gold-set \(H_T, R_T\). GDPval-AA Elo is a leading indicator that agentic quality on the same task family kept climbing through mid-2026:

  1. Directionally, models far above Human Baseline 1000 (Opus 5, Grok 4.6, Sol max) are the class of systems for which oversight-compatible ROI is more plausible, not less—higher effective ship rates shrink the (1−w)·HT redo term.
  2. You still cannot skip local w and RT. Leaderboard Elo does not measure your review checklist minutes or your occupation mix.
  3. Cost per task on the leaderboard (AA’s $0.84 for Grok 4.6) sits in the same order of magnitude as the paper’s sub-dollar API costs—still rounding error next to ~$361 human labor proxies. The binding enterprise cost remains reviewer time, not tokens.
  4. Harness still matters. AA and xAI emphasize multi-turn agents, tool use, and self-checking. That reinforces the paper’s scaffolding result: chat-only deployments under-sample GDPval-class capability.
  5. Snorkel’s GDPval+ caution: even strong 2026 models can show mean pass rates under one-third of expert rubric criteria on harder expanded sets (e.g. Grok 4.5 at ~29% mean pass on ~2,000 GDPval+ tasks in Snorkel’s July 2026 writeup). “Leaderboard Elo high” and “passes every expert criterion” are different bars. Reviewers stay mandatory.

Trend update

Trend: Public GDPval-family scores for frontier models continued rising from late 2025 through Aug 2026 on third-party agentic harnesses.
Pattern: Stepwise generational jumps (e.g. Grok 4.5→4.6 +200 Elo) plus dense effort-tier ladders (Opus 5 low→max).
Classification: Too data-poor to call exponential on Elo (short window, changing harnesses); best label is rapid stepwise climb on agentic preference/Elo metrics, with Anthropic holding peak Elo and SpaceXAI/xAI reclaiming Pareto cost-efficiency on the same family.
Falsifier: A year of flat AA top-10 Elo with stable harness, or independent human pairwise studies showing win rates stuck near 2025 GPT-5 levels despite Elo gains.

Implications for decision-makers

← ImplicationsResearch references →