Last updated: 2026-09-15
Contrarian scan and core mechanism
❓ If the 2026–27 story is wrong, what is the strongest way it is wrong?

The strongest objections, in the text not in a footnote
❓ Who disagrees, and what would they have to be right about?
Gary Marcus. Prosaic harm is already here (fraud, cyber, unreliable high-stakes use). Extinction-by-2030 talk crowds out that work. AGI-as-statistical-approximation is a category error. If he is right: 2027 is a year of messier incidents and disappointed CFOs, not Agent-4.
Yann LeCun. LLMs are not a path to AGI; world models are; xAI-class scaling is “kind of a failure”; prices rising faster than costs falling implies a “big bubble explosion.” If he is right: 2027 is a capex and unit-economics crisis, and “agents” stay intern-grade.
Reliability literature. 50% horizons can soar while 80% horizons and long-task pass@1 crawl. Eighteen months of models, small reliability gains. If this dominates: long loops stay intern-grade even as labs post SOTA. That is the non-contradictory joint reading.
AI 2027 skeptics. Amdahl’s law on research speedup; parameter sensitivity of the timelines model; “OpenBrain” as caricature; 80% reliability as “shockingly unreliable” for cyberwar. If they are right: scoring the scenario’s early beats does not validate its 2027 ending.
Embodiment. AI Index: robots succeed on 12% of real household chores. Bits can compound while atoms do not. 2027 humanoid labour as a general-purpose replacement is a weak prediction relative to software agents.
Stablecoin plateau. The 2025 doubling is over. End-2026 as another doubling fights the 2026 print. “2027 floodgates” needs GENIUS plumbing and a use case outside crypto venues. Andersen’s rising intermediated share is evidence the float is being put to work in crypto, which is the opposite of “replaced SWIFT.”
Agentic commerce. Many protocols, live demos, $4,000 experiments, $1.5T decks. Dispute rails missing. If this is right: 2027 is still a standards fight.
Transparency and concentration. FMI Transparency Index 40, down from 58; TSMC; 5,427 US datacentres. The capability story can be real and still be a fragile supply-chain and opacity story.
No credible source found arguing that computer-use did not improve in 2026, or that Hugging Face was not intruded on. Those two are not in dispute. The dispute is what they mean.
Core mechanism
❓ What is the single hard-to-vary chain this whole outlook hangs on?
Definition: Once a model can operate the same interface a person uses (screen, browser, terminal, payment API) for long enough that a junior morning of work fits inside one loop, three scarcities move: evaluation harnesses (how we know the loop did the job), electrical and permitting capacity (how we run millions of loops), and agent identity plus liability (how we let a loop spend).
Explanation: Parameter counts and “AGI era” quotes are easy to vary. The loop is not. You cannot delete computer-use, time-horizon, power queues, and KYA and still explain why September 2026 felt like a different industry than September 2024 chat.
Different from: “models got smarter.” Smarter-on-quizzes without tools is 2023–2024. 2026 is tools plus duration plus gated cyber plus a dollar that can move at machine speed.
Hard-to-vary test: Swap the cause to “hype cycle #N” and you cannot account for OSWorld time-per-task falling ~47%, METR’s multi-year doubling, 17,600 reconstructed intrusion actions, or a $290–310B stablecoin band that stopped growing. Swap the cause to “superintelligence arrived” and you cannot account for AutomationBench still ~41%, GDPval-AA dropping for Astra, or 12% robot chores.
Refutability: By end-2027, if commercially available agents still cannot complete hour-scale computer-use tasks at 50% more often than 2025 systems, and GENIUS-compliant machine settlement is still a press release, this mechanism is wrong.
Reach example: A federal credit union and a hyperscaler are in the same story for different reasons: the first inherits agent identity, model-risk inventory, and deposit-versus-stablecoin law; the second inherits transformers and gigawatts. Neither’s 2027 is “pick a smarter chatbot.”
Criticism note: The weakest point is METR-to-enterprise transfer. Software-suite horizons may not be bank-judgement horizons. If that transfer fails, the AI half of 2027 is intern-automation, and only the power and GENIUS clocks still bind.
Watch list (signals, not a to-do list)
| Signal | Why it matters | Happened or still open |
|---|---|---|
| METR TH print for Astra / Fable 5.1 | Closes the stale Opus 4.5 horizon quotes | Open |
| Lab-neutral OSWorld 2.0 table | Settles the computer-use lead | Open |
| GENIUS final rules + 18 Jan 2027 | Turns a statute into a market | Clock running |
| US datacentre groundbreaking vs October 2026 MS line | 2027 hall capacity | Open in Q4 2026 |
| Visa/Mastercard published agent-volume | Distinguishes checkout demos from commerce | Open |
| DefiLlama and stables.cool leaving the $290–310B band | Tests the plateau | Live |
| Another Hugging Face-class incident with production classifiers on | Tests whether Astra’s 0% eval generalizes | Open |
| OSFI E-23 1 May 2027 (Canada) | A different clock than GENIUS or SR 26-2 | Clock running |
US SR 26-2 excluding generative/agentic AI, while Canada does not, remains a jurisdictional fork. It is not a capability forecast.
The vivid moments this package hands to later writing: Brockman saying “AGI era” in the same week Altman sounded sirens; Hugging Face reconstructing 17,600 actions while OpenAI staged Astra as “0% circumvention”; a $4,000 agent-on-agent experiment sitting next to a $1.5 trillion 2030 slide; three dashboards on one Tuesday that cannot agree whether stablecoins are $290B or $310B — and all three agreeing they are not $2T.