Research · Full question map

Question ledger

Last updated: 2026-08-12

Observed: 2026-08-25. Living document. Target: 8–15 primaries, 6–8 secondaries each.

Scope: compare Grok Bot (xAI/X agent product, including persistent cloud computer) with Hermes Agent Bots (Nous Research Hermes Agent looking-glass bots + harness), mapped onto the house Agentic AI Stack. Hunt idea-borrowing, release-note pace, Nous Portal effects, and stack gaps.

Confirmed vs unverified split is mandatory: product facts from official docs/release notes; X/video as signal; podcasts as takes.


P1. What product object is “Grok Bot,” and what product object is a “Hermes Agent Bot”?

Would a buyer, builder, or architect be comparing two instances of the same LTC, or two packaging modes?

Secondaries

  1. Is Grok Bot a vendor-provided agent SKU, a managed harness, a messaging looking glass, a custom-bot builder, or several of those at once?
  2. Is “Hermes Agent Bot” the gateway messaging adapter, the desktop app, a named profile worker, or the whole harness product?
  3. Which house packaging mode (vendor SKU / self-hosted open harness / app-defined / framework) does each occupy?
  4. What is the constitutive agent (model + prompt + tools) versus the harness (loop) versus the host (runtime) in each product’s own language?
  5. Does either vendor use “bot” to mean something the stack would call a looking-glass adapter?
  6. Are Grok custom bots on X the same product as Grok’s persistent computer agent?
  7. What would falsify treating them as peers (e.g. one is chat-with-tools, the other is a full harness)?
  8. Which official names, SKUs, and dates exist as of 2026-08-25?

Rabbit holes


P2. How do both map onto the house Agentic AI Stack layers?

If we overlay both products on Model / Persistent State / Tools / Skills / Harness / Multi-agent / Looking Glass, which boxes are filled, empty, or misnamed by the vendors?

Secondaries

  1. Model band: whose weights, routing, local vs hyperscaler inference?
  2. Persistent State: files/DB/enterprise SoR vs session/memory (category error if mixed)?
  3. Tools: built-ins, CLI, MCP, connectors, execution environment/safe sandbox?
  4. Skills: SKILL.md progressive disclosure vs prompt packs vs Grok custom instructions?
  5. Harness: loop, dispatch, context, memory, approvals, sandbox policy, budgets, tracing?
  6. Multi-agent: board / graph / router / sub-agent / A2A — present or missing?
  7. Looking Glass: messaging, native desktop, HITL?
  8. Which vendor labels would cause a category error against the house blacklist (MCP≠orchestration, harness≠board, vendor agent≠layer)?

P3. Grok Bot has a persistent cloud computer. What is Hermes’s counterpart?

Is the scarce object a durable VM, a session-scoped sandbox, a user machine, or a rented inference host?

Secondaries

  1. What does xAI/X actually persist (disk, processes, browser, IP, identity, hours/days)?
  2. Is it per-user, per-bot, per-task, or pooled?
  3. Hermes terminal backends: local / docker / ssh / modal — which is “a computer”?
  4. Does hermes serve remote desktop backend count as a persistent computer or as a looking-glass+host split?
  5. Does a Nous Portal subscription add a hosted computer, or only inference/auth?
  6. Operate vs rent: who owns the machine, who can inspect it, who can snapshot it?
  7. What dies on crash/logout for each (on-prem vs cloud must not be the analysis axis — operate vs rent is)?
  8. Is “computer use” (GUI/browser) the same layer as “persistent computer” (host)?

P4. What does a Nous Portal subscription change in the Hermes stack?

Does Portal move Hermes from self-hosted harness toward a vendor-packaged agent, or only change the Model band?

Secondaries

  1. What SKUs exist (free, paid, team) and what is included vs metered?
  2. Inference only, or also hosted gateway, memory, skills hub, computer?
  3. OAuth vs BYOK vs local models — lock-in surface?
  4. Skills hub / curator / cloud sync tied to Portal?
  5. Does Portal change HITL, approvals, or audit?
  6. Comparable Grok SKU (Free / SuperGrok / X Premium) vs computer entitlement?
  7. Exit rights: if Portal lapses, what still runs locally?
  8. Privacy/sovereignty: where do prompts, memory, logs, computer disk live?

P5. Tools, sandbox, and side-effect control — who is actually safer/more complete?

Agency is closed-loop control over side effects. Who implements handlers, where do they run, who gates writes?

Secondaries

  1. Built-in families on each side (fs, shell, browser, code_exec, web, media, payments).
  2. MCP support on Grok Bot vs Hermes.
  3. CLI-as-enterprise-IF vs MCP; p4 CLI / p3 outside harness (house rule).
  4. Sandbox policy: who enforces, is it prompt-only?
  5. Lethal trifecta (private data + untrusted content + exfil) handling.
  6. Approvals / YOLO / SuperGrok-equivalent bypass.
  7. Browser/computer-use vs CDP vs managed browser.
  8. Connectors to enterprise SoR (Google, Slack, GitHub, X itself).

P6. Skills, memory, and self-improvement — idea borrowing across products?

Skills ≠ prompts ≠ memory. Where is each product putting know-how, and who copies whom?

Secondaries

  1. Does Grok Bot have a SKILL.md-class progressive procedure layer?
  2. Hermes skill_manage / curator vs Grok custom bots / instructions / workspaces.
  3. Memory: always-on facts vs session vs computer disk as pseudo-memory.
  4. Self-improvement write-back: allowed, gated, or marketing?
  5. Cross-pollination: OpenClaw/Claw, Claude Code, Codex, Hermes, Grok — who shipped what first?
  6. AGENTS.md / AAIF convention adoption.
  7. Pin/archive/review of agent-written skills.
  8. What X posts claim “Grok copied Hermes” or the reverse, and what primary evidence remains?

P7. Multi-agent: bots, workers, and boards

When does “bot” mean one loop, and when does it mean many durable workers?

Secondaries

  1. Can Grok Bot spawn sub-agents? Durable or RPC?
  2. Hermes delegate_task vs Kanban vs gateway multi-profile.
  3. Custom Grok bots on X as a multi-agent marketplace?
  4. A2A / inter-bot messaging.
  5. Crash reclaim / durable task state.
  6. Human gates on multi-bot work.
  7. Is Grok’s computer shared across bots or siloed?
  8. Bonus-level test: can each product do useful single-agent work without orchestration?

P8. Looking glass, HITL, and “bots” as surfaces

Users meet a glass, not a stack. What glass, what gate, what recorded authority?

Secondaries

  1. Grok on X, grok.com, Grok apps, API — which is the bot looking glass?
  2. Hermes CLI, Desktop, Telegram/Discord/Slack/WhatsApp gateway.
  3. HITL primitives: gate, pause, context packet, recorded decision, resume.
  4. Messaging-bot pairing/auth (Hermes pairing vs X bot permissions).
  5. Native desktop vs chat-in-feed.
  6. Voice/video as looking glass (Grok voice, Hermes TTS/STT).
  7. Does “bot” in consumer language hide the harness?
  8. Deliberate looking-glass lock-in (X identity vs self-hosted gateway).

P9. Release-note pace — what shipped, who copied, what is still vapor?

The user says the pace is insane. Separate confirmed, credible-unconfirmed, and rumor.

Secondaries

  1. Official Grok/xAI/X release notes since ~2026-01 (and 2025 computer/agent launches).
  2. Official Hermes Agent / Nous changelog (GitHub releases, docs).
  3. Time-to-copy for skills, computer, MCP, cron, memory, multi-agent.
  4. What Grok announced that Hermes already had, and vice versa.
  5. What both still lack vs Claude Code / Codex / OpenClaw / Cowork.
  6. Video/demo claims that exceed docs.
  7. Breaking changes / deprecations that would flip an architecture call.
  8. Rate of shipping: stepwise product jumps vs compounding harness craft.

P10. Did the house Agentic AI Stack miss a box?

Overlaying two fast products is a stress test of the stack, not only of the products.

Secondaries

  1. Persistent computer / runtime host as a peer box vs Tools “execution environment/safe sandbox”?
  2. Identity / billing / subscription plane (Portal, X Premium, SuperGrok)?
  3. Distribution / bot directory / skills hub as a marketplace layer?
  4. Payments / 402 / agent commerce on Grok or Hermes?
  5. Evaluation / evidence packs / tracing as first-class?
  6. Voice/realtime as looking-glass family vs model modality?
  7. Operate-vs-rent host banner currently in Harness — still correct?
  8. Anything Grok/Hermes force us to split (harness loop vs host VM vs inference)?

Rabbit holes


P11. Economics, lock-in, and operate-vs-rent

What becomes abundant, what stays scarce, who can leave?

Secondaries

  1. Metering: tokens, computer hours, tool calls, bots.
  2. Skill/memory/computer-disk portability.
  3. Model swappability (Hermes) vs Grok-only cognition.
  4. Gateway/bot identity portability.
  5. Cost of a durable computer vs local machine.
  6. Enterprise vs consumer SKU gap.
  7. Open-source Hermes vs closed Grok computer — true exit or GitHub theater?
  8. Suite gravity of X vs Nous community gravity.

P12. Risks, failure modes, and contrarian views

Who disagrees that these are the two products to watch, and where do they fail?

Secondaries

  1. Prompt injection via X posts into Grok Bot / computer.
  2. Hermes YOLO on a personal machine vs Grok sandboxed cloud computer.
  3. Overclaiming “persistent computer” for ephemeral containers.
  4. Multi-bot chaos without a board.
  5. “Bots” as a dead-end vs harness-as-product.
  6. OpenClaw / Claude Code / Codex as the real peers, Grok/Hermes as cousins.
  7. Regulatory: agent on a social network vs local harness.
  8. Rogue-superintelligence / automation global-problem link (bounded, not sci-fi).

P13. Thought-leader and X/video signal (not sole facts)

What do named builders claim the other side got right?

Secondaries

  1. Teknium / Nous / Hermes maintainers on Grok.
  2. Elon / xAI / Grok team on agents, computer, bots.
  3. Karpathy / other harness voices mapping both.
  4. Demo videos: what is actually on screen (computer desktop, terminal, chat)?
  5. All-In / Moonshots takes on Grok agents vs open harnesses.
  6. Contrarian X: Grok computer is a toy; Hermes is ungovernable.
  7. Claims of copying Clawdbot / OpenClaw / Claude Cowork.
  8. Which claims survive primary-doc check?

Unresolved (seed)

Saturation notes

Ledger created 2026-08-25 before index/sections (bundle ledger-first). Round 1 (same day) answered P1–P12 in chapters 01–08 from official docs. P13 podcast lane is thin (no Grok Bot episode). Unresolved host SLA / Hermes Cloud desktop roadmap remain marked. Contrarian scan in 08. No new primary added after the doc fetch.

← Research referencesBack to essay →