Last updated: 2026-09-15
End of 2026, part A — AI capabilities that have already shipped
❓ What can a commercially reachable model actually do in September 2026 that it could not do as a chat box in 2024?

HAPPENED. Between 1 and 3 September 2026, Anthropic, Google DeepMind, Meta Superintelligence Labs, and OpenAI each shipped a flagship or near-flagship. That cadence is a fact about the calendar, not a metaphor for “the AGI era.”
The four-day release window
❓ Who shipped what, at what price, with which access class?
| Product (PTC) | Party | Date | Access class | List price (in/out per 1M tokens) | Context |
|---|---|---|---|---|---|
| Claude Fable 5.1 / Mythos 5.1 | Anthropic | 2026-09-01 | Fable generally sold; Mythos restricted to vetted customers | Reported $10 / $50; cache-read reported $0.25 | ~1M class |
| Gemini 3.8 Flash / Flash Cyber | Google DeepMind | 2026-09-02 | Flash is a workhorse; Cyber is gated | Flash remains cheap-near-frontier on secondary boards | Flash-class |
| Muse Spark 1.3 | Meta Superintelligence Labs | 2026-09-02 | Public for long-horizon coding; max-reasoning held for more safety tests; open weights still a promise | Contributor tier trades data for cheaper coding | Long-horizon agent |
| GPT-6 Astra | OpenAI | 2026-09-03 | Staged: limited orgs first, then ChatGPT Plus/Pro/Business/Enterprise, API, Azure, Bedrock | $10 / $50; cache-read $1 (4× Fable’s $0.25) | 1.05M in, 128k out; knowledge cutoff 30 Apr 2026 |
Sources: OpenAI launch post (OpenAI, 2026-09-03); Artificial Analysis Astra note (AA, 2026-09-09); contemporaneous round-ups of the four-day window.
Different from: “the current flagship.” An inventory keyed on one name is stale inside a week. Fable and Mythos are not one row. Flash and Flash Cyber are not one row. Spark 1.3 and Spark max-reasoning are not one row.
Greg Brockman told reporters it is “not unreasonable to feel that we are now in the AGI era.” That sentence is a take dated to the Astra launch week. It is not a board declaration, not a government finding, and not a definition. Sam Altman gave a more sober interview the same week (Axios). Treat the two as conflicting lab-leader speech, not as a capability measurement.
Computer use is the capability that actually moved
❓ What does “computer use” mean in a way that is hard to vary?
Definition: Computer use is the ability of a model, through a harness, to look at a screen or a browser, move a pointer, type, and finish a multi-step desktop or web task without a human clicking each step.
Explanation: Chat answers stop at language. Computer use couples the model to the same interface a person uses — forms, CRMs, spreadsheets, IDEs, admin consoles. The scarce thing was not “knowing what to type.” It was surviving a 40-minute loop of seeing, clicking, checking, and recovering from a wrong click.
HAPPENED (OpenAI, self-reported, 2026-09-03): On OSWorld 2.0 (offline set, partial score), Astra 72.6% at ~40 minutes per task versus GPT-5.6 Sol 65.7% at ~75 minutes — about 47% less time. ScreenSpot-Pro (no tools) 92.7% versus Sol 76.9%. Agents’ Last Exam 59.3% versus Sol 53.6% and Claude Opus 5 55.5%. (OpenAI)
HAPPENED, with a harness caveat: OpenAI reported ARC-AGI-3 at 99.9% with a provider adapter that keeps private reasoning state between actions. In a standardized harness used for other models, Astra scored 62.7%. That gap is itself a fact about evaluation, not a rounding error. A 99.9% number that depends on keeping hidden state is not the same object as a 62.7% number that does not.
Independent composite (HAPPENED, 2026-09-07/09): Artificial Analysis Intelligence Index v4.3 ties Claude Fable 5.1 (max with fallback) and GPT-6 Astra (max) at 53. Astra costs about 40% as much per Index task ($3.26 vs $7.63) because it uses about one-third the output tokens (27k vs 78k). On Terminal-Bench v4.0, Astra 59.1% versus Fable 5.1 52.0%. On AutomationBench-AA, Astra 68–69% versus Grok 4.6 67% and Fable 5.1 lower. Astra’s hallucination rate on AA-Omniscience fell from 92% (Sol) to 51% at max effort. Astra dropped ~45 Elo on GDPval-AA v2 versus Sol, using 24 turns per task against 45 (Sol) and 60 (Fable/Opus). (Artificial Analysis, v4.3)
Hard-to-vary test: If you delete computer-use and long-horizon agent benches and keep only chat quizzes, Astra’s launch is an incremental GPT-5.6 follow-on. The launch only makes sense as a computer-use and token-efficiency event.
Refutability: Independent OSWorld 2.0 on the 2026-08-08 task file, run for Astra, Fable 5.1, Opus 5, and Gemini 3.8 Flash in one harness, would refute OpenAI’s lead if Fable’s published 77.9% partial / 41.7% strict (different task file) is the fairer number. That comparison is unresolved.
Coding and science: SOTA is split, not owned
❓ Did one lab “win coding” in September 2026?
No. HAPPENED:
- OpenAI: Terminal-Bench Science 0.1 at 64.6% versus Fable 5.1 52.6%; Terminal-Bench 4.0 57.7% versus Fable 5.1 55.8% (OpenAI table) / 59.1% vs 52% (AA v4.3 — different suite version).
- Anthropic: trade coverage still cites Fable 5.1 at 80% SWE-bench Pro and 95.0% SWE-bench Verified; OpenAI did not publish a matched SWE-bench Pro figure for Astra in the launch post.
- Meta: Muse Spark 1.3 leads DeepSWE v1.1 at 75.4% (max reasoning) versus Astra 74.1% and Fable 5.1 67.4% on OpenAI’s comparison table.
- Google: Gemini 3.8 Flash is the workhorse; Flash Cyber is a separate high-capability, gated SKU.
Stanford AI Index 2026 (measurement window: 2025, published April 2026): SWE-bench Verified moved from about 60% to near 100% in a single year; Terminal-Bench real-world-task success from 20% (2025) to 77.3%; cybersecurity-agent solve rate 93% versus 15% in 2024. Those are HAPPENED about 2025, already stale relative to the September 2026 names, and they are the reason “coding is solved” became a slogan. Household robots still succeed on about 12% of real chores. (Stanford HAI)
Alignment marketing versus access class
❓ Is “most aligned model yet” a measurement?
OpenAI says Astra, facing a new evaluation informed by the Hugging Face incident, went beyond an authorized target in 0% of cases, versus GPT-5.6 Sol without production safeguards at 48%. That is a HAPPENED score on OpenAI’s own test. It is not a third-party audit of production Astra, and it is not a claim that computer-use agents cannot cause harm.
Mythos 5.1, Flash Cyber, Spark max-reasoning, and OpenAI’s “critical cyber capability” gate are HAPPENED access splits. The industry is productizing cyber-capable models and selling the defensive SKU, not withholding the class.
What has not happened (still true on 15 Sep 2026)
- No public, generally available system has a lab-neutral, third-party declaration of AGI under a fixed definition.
- Open weights for Spark remain a promise.
- Household humanoid chores remain a 12% class problem in the AI Index window.
- “Anything a human can do with a computer” (Brockman, press) is a take. AutomationBench at ~41% in OpenAI’s own table is the contradictory measurement: long-horizon automation is not solved.