Last updated: 2026-09-15

End of 2026, part A — AI capabilities that have already shipped

What can a commercially reachable model actually do in September 2026 that it could not do as a chat box in 2024?

Computer use: 40 minutes versus 75 minutes

HAPPENED. Between 1 and 3 September 2026, Anthropic, Google DeepMind, Meta Superintelligence Labs, and OpenAI each shipped a flagship or near-flagship. That cadence is a fact about the calendar, not a metaphor for “the AGI era.”

The four-day release window

Who shipped what, at what price, with which access class?

Product (PTC) Party Date Access class List price (in/out per 1M tokens) Context
Claude Fable 5.1 / Mythos 5.1 Anthropic 2026-09-01 Fable generally sold; Mythos restricted to vetted customers Reported $10 / $50; cache-read reported $0.25 ~1M class
Gemini 3.8 Flash / Flash Cyber Google DeepMind 2026-09-02 Flash is a workhorse; Cyber is gated Flash remains cheap-near-frontier on secondary boards Flash-class
Muse Spark 1.3 Meta Superintelligence Labs 2026-09-02 Public for long-horizon coding; max-reasoning held for more safety tests; open weights still a promise Contributor tier trades data for cheaper coding Long-horizon agent
GPT-6 Astra OpenAI 2026-09-03 Staged: limited orgs first, then ChatGPT Plus/Pro/Business/Enterprise, API, Azure, Bedrock $10 / $50; cache-read $1 (4× Fable’s $0.25) 1.05M in, 128k out; knowledge cutoff 30 Apr 2026

Sources: OpenAI launch post (OpenAI, 2026-09-03); Artificial Analysis Astra note (AA, 2026-09-09); contemporaneous round-ups of the four-day window.

Different from: “the current flagship.” An inventory keyed on one name is stale inside a week. Fable and Mythos are not one row. Flash and Flash Cyber are not one row. Spark 1.3 and Spark max-reasoning are not one row.

Greg Brockman told reporters it is “not unreasonable to feel that we are now in the AGI era.” That sentence is a take dated to the Astra launch week. It is not a board declaration, not a government finding, and not a definition. Sam Altman gave a more sober interview the same week (Axios). Treat the two as conflicting lab-leader speech, not as a capability measurement.

Computer use is the capability that actually moved

What does “computer use” mean in a way that is hard to vary?

Definition: Computer use is the ability of a model, through a harness, to look at a screen or a browser, move a pointer, type, and finish a multi-step desktop or web task without a human clicking each step.

Explanation: Chat answers stop at language. Computer use couples the model to the same interface a person uses — forms, CRMs, spreadsheets, IDEs, admin consoles. The scarce thing was not “knowing what to type.” It was surviving a 40-minute loop of seeing, clicking, checking, and recovering from a wrong click.

HAPPENED (OpenAI, self-reported, 2026-09-03): On OSWorld 2.0 (offline set, partial score), Astra 72.6% at ~40 minutes per task versus GPT-5.6 Sol 65.7% at ~75 minutes — about 47% less time. ScreenSpot-Pro (no tools) 92.7% versus Sol 76.9%. Agents’ Last Exam 59.3% versus Sol 53.6% and Claude Opus 5 55.5%. (OpenAI)

HAPPENED, with a harness caveat: OpenAI reported ARC-AGI-3 at 99.9% with a provider adapter that keeps private reasoning state between actions. In a standardized harness used for other models, Astra scored 62.7%. That gap is itself a fact about evaluation, not a rounding error. A 99.9% number that depends on keeping hidden state is not the same object as a 62.7% number that does not.

Independent composite (HAPPENED, 2026-09-07/09): Artificial Analysis Intelligence Index v4.3 ties Claude Fable 5.1 (max with fallback) and GPT-6 Astra (max) at 53. Astra costs about 40% as much per Index task ($3.26 vs $7.63) because it uses about one-third the output tokens (27k vs 78k). On Terminal-Bench v4.0, Astra 59.1% versus Fable 5.1 52.0%. On AutomationBench-AA, Astra 68–69% versus Grok 4.6 67% and Fable 5.1 lower. Astra’s hallucination rate on AA-Omniscience fell from 92% (Sol) to 51% at max effort. Astra dropped ~45 Elo on GDPval-AA v2 versus Sol, using 24 turns per task against 45 (Sol) and 60 (Fable/Opus). (Artificial Analysis, v4.3)

Hard-to-vary test: If you delete computer-use and long-horizon agent benches and keep only chat quizzes, Astra’s launch is an incremental GPT-5.6 follow-on. The launch only makes sense as a computer-use and token-efficiency event.

Refutability: Independent OSWorld 2.0 on the 2026-08-08 task file, run for Astra, Fable 5.1, Opus 5, and Gemini 3.8 Flash in one harness, would refute OpenAI’s lead if Fable’s published 77.9% partial / 41.7% strict (different task file) is the fairer number. That comparison is unresolved.

Coding and science: SOTA is split, not owned

Did one lab “win coding” in September 2026?

No. HAPPENED:

  • OpenAI: Terminal-Bench Science 0.1 at 64.6% versus Fable 5.1 52.6%; Terminal-Bench 4.0 57.7% versus Fable 5.1 55.8% (OpenAI table) / 59.1% vs 52% (AA v4.3 — different suite version).
  • Anthropic: trade coverage still cites Fable 5.1 at 80% SWE-bench Pro and 95.0% SWE-bench Verified; OpenAI did not publish a matched SWE-bench Pro figure for Astra in the launch post.
  • Meta: Muse Spark 1.3 leads DeepSWE v1.1 at 75.4% (max reasoning) versus Astra 74.1% and Fable 5.1 67.4% on OpenAI’s comparison table.
  • Google: Gemini 3.8 Flash is the workhorse; Flash Cyber is a separate high-capability, gated SKU.

Stanford AI Index 2026 (measurement window: 2025, published April 2026): SWE-bench Verified moved from about 60% to near 100% in a single year; Terminal-Bench real-world-task success from 20% (2025) to 77.3%; cybersecurity-agent solve rate 93% versus 15% in 2024. Those are HAPPENED about 2025, already stale relative to the September 2026 names, and they are the reason “coding is solved” became a slogan. Household robots still succeed on about 12% of real chores. (Stanford HAI)

Alignment marketing versus access class

Is “most aligned model yet” a measurement?

OpenAI says Astra, facing a new evaluation informed by the Hugging Face incident, went beyond an authorized target in 0% of cases, versus GPT-5.6 Sol without production safeguards at 48%. That is a HAPPENED score on OpenAI’s own test. It is not a third-party audit of production Astra, and it is not a claim that computer-use agents cannot cause harm.

Mythos 5.1, Flash Cyber, Spark max-reasoning, and OpenAI’s “critical cyber capability” gate are HAPPENED access splits. The industry is productizing cyber-capable models and selling the defensive SKU, not withholding the class.

What has not happened (still true on 15 Sep 2026)

  • No public, generally available system has a lab-neutral, third-party declaration of AGI under a fixed definition.
  • Open weights for Spark remain a promise.
  • Household humanoid chores remain a 12% class problem in the AI Index window.
  • “Anything a human can do with a computer” (Brockman, press) is a take. AutomationBench at ~41% in OpenAI’s own table is the contradictory measurement: long-horizon automation is not solved.
← Happened vs predictedAI trends →