Research · Full question map
Question ledger
Last updated: 2026-08-12
Observed: 2026-08-25. Living document. Target: 8–15 primaries, 6–8 secondaries each.
Scope: compare Grok Bot (xAI/X agent product, including persistent cloud computer) with Hermes Agent Bots (Nous Research Hermes Agent looking-glass bots + harness), mapped onto the house Agentic AI Stack. Hunt idea-borrowing, release-note pace, Nous Portal effects, and stack gaps.
Confirmed vs unverified split is mandatory: product facts from official docs/release notes; X/video as signal; podcasts as takes.
P1. What product object is “Grok Bot,” and what product object is a “Hermes Agent Bot”?
Would a buyer, builder, or architect be comparing two instances of the same LTC, or two packaging modes?
Secondaries
- Is Grok Bot a vendor-provided agent SKU, a managed harness, a messaging looking glass, a custom-bot builder, or several of those at once?
- Is “Hermes Agent Bot” the gateway messaging adapter, the desktop app, a named profile worker, or the whole harness product?
- Which house packaging mode (vendor SKU / self-hosted open harness / app-defined / framework) does each occupy?
- What is the constitutive agent (model + prompt + tools) versus the harness (loop) versus the host (runtime) in each product’s own language?
- Does either vendor use “bot” to mean something the stack would call a looking-glass adapter?
- Are Grok custom bots on X the same product as Grok’s persistent computer agent?
- What would falsify treating them as peers (e.g. one is chat-with-tools, the other is a full harness)?
- Which official names, SKUs, and dates exist as of 2026-08-25?
Rabbit holes
- X’s historical “Grok bot” vs 2026 “Grok computer / agent” rebrand.
- Hermes “gateway bot” vs “Kanban worker” vs “desktop agent.”
P2. How do both map onto the house Agentic AI Stack layers?
If we overlay both products on Model / Persistent State / Tools / Skills / Harness / Multi-agent / Looking Glass, which boxes are filled, empty, or misnamed by the vendors?
Secondaries
- Model band: whose weights, routing, local vs hyperscaler inference?
- Persistent State: files/DB/enterprise SoR vs session/memory (category error if mixed)?
- Tools: built-ins, CLI, MCP, connectors, execution environment/safe sandbox?
- Skills: SKILL.md progressive disclosure vs prompt packs vs Grok custom instructions?
- Harness: loop, dispatch, context, memory, approvals, sandbox policy, budgets, tracing?
- Multi-agent: board / graph / router / sub-agent / A2A — present or missing?
- Looking Glass: messaging, native desktop, HITL?
- Which vendor labels would cause a category error against the house blacklist (MCP≠orchestration, harness≠board, vendor agent≠layer)?
P3. Grok Bot has a persistent cloud computer. What is Hermes’s counterpart?
Is the scarce object a durable VM, a session-scoped sandbox, a user machine, or a rented inference host?
Secondaries
- What does xAI/X actually persist (disk, processes, browser, IP, identity, hours/days)?
- Is it per-user, per-bot, per-task, or pooled?
- Hermes terminal backends: local / docker / ssh / modal — which is “a computer”?
- Does
hermes serveremote desktop backend count as a persistent computer or as a looking-glass+host split? - Does a Nous Portal subscription add a hosted computer, or only inference/auth?
- Operate vs rent: who owns the machine, who can inspect it, who can snapshot it?
- What dies on crash/logout for each (on-prem vs cloud must not be the analysis axis — operate vs rent is)?
- Is “computer use” (GUI/browser) the same layer as “persistent computer” (host)?
P4. What does a Nous Portal subscription change in the Hermes stack?
Does Portal move Hermes from self-hosted harness toward a vendor-packaged agent, or only change the Model band?
Secondaries
- What SKUs exist (free, paid, team) and what is included vs metered?
- Inference only, or also hosted gateway, memory, skills hub, computer?
- OAuth vs BYOK vs local models — lock-in surface?
- Skills hub / curator / cloud sync tied to Portal?
- Does Portal change HITL, approvals, or audit?
- Comparable Grok SKU (Free / SuperGrok / X Premium) vs computer entitlement?
- Exit rights: if Portal lapses, what still runs locally?
- Privacy/sovereignty: where do prompts, memory, logs, computer disk live?
P5. Tools, sandbox, and side-effect control — who is actually safer/more complete?
Agency is closed-loop control over side effects. Who implements handlers, where do they run, who gates writes?
Secondaries
- Built-in families on each side (fs, shell, browser, code_exec, web, media, payments).
- MCP support on Grok Bot vs Hermes.
- CLI-as-enterprise-IF vs MCP; p4 CLI / p3 outside harness (house rule).
- Sandbox policy: who enforces, is it prompt-only?
- Lethal trifecta (private data + untrusted content + exfil) handling.
- Approvals / YOLO / SuperGrok-equivalent bypass.
- Browser/computer-use vs CDP vs managed browser.
- Connectors to enterprise SoR (Google, Slack, GitHub, X itself).
P6. Skills, memory, and self-improvement — idea borrowing across products?
Skills ≠ prompts ≠ memory. Where is each product putting know-how, and who copies whom?
Secondaries
- Does Grok Bot have a SKILL.md-class progressive procedure layer?
- Hermes
skill_manage/ curator vs Grok custom bots / instructions / workspaces. - Memory: always-on facts vs session vs computer disk as pseudo-memory.
- Self-improvement write-back: allowed, gated, or marketing?
- Cross-pollination: OpenClaw/Claw, Claude Code, Codex, Hermes, Grok — who shipped what first?
- AGENTS.md / AAIF convention adoption.
- Pin/archive/review of agent-written skills.
- What X posts claim “Grok copied Hermes” or the reverse, and what primary evidence remains?
P7. Multi-agent: bots, workers, and boards
When does “bot” mean one loop, and when does it mean many durable workers?
Secondaries
- Can Grok Bot spawn sub-agents? Durable or RPC?
- Hermes
delegate_taskvs Kanban vs gateway multi-profile. - Custom Grok bots on X as a multi-agent marketplace?
- A2A / inter-bot messaging.
- Crash reclaim / durable task state.
- Human gates on multi-bot work.
- Is Grok’s computer shared across bots or siloed?
- Bonus-level test: can each product do useful single-agent work without orchestration?
P8. Looking glass, HITL, and “bots” as surfaces
Users meet a glass, not a stack. What glass, what gate, what recorded authority?
Secondaries
- Grok on X, grok.com, Grok apps, API — which is the bot looking glass?
- Hermes CLI, Desktop, Telegram/Discord/Slack/WhatsApp gateway.
- HITL primitives: gate, pause, context packet, recorded decision, resume.
- Messaging-bot pairing/auth (Hermes pairing vs X bot permissions).
- Native desktop vs chat-in-feed.
- Voice/video as looking glass (Grok voice, Hermes TTS/STT).
- Does “bot” in consumer language hide the harness?
- Deliberate looking-glass lock-in (X identity vs self-hosted gateway).
P9. Release-note pace — what shipped, who copied, what is still vapor?
The user says the pace is insane. Separate confirmed, credible-unconfirmed, and rumor.
Secondaries
- Official Grok/xAI/X release notes since ~2026-01 (and 2025 computer/agent launches).
- Official Hermes Agent / Nous changelog (GitHub releases, docs).
- Time-to-copy for skills, computer, MCP, cron, memory, multi-agent.
- What Grok announced that Hermes already had, and vice versa.
- What both still lack vs Claude Code / Codex / OpenClaw / Cowork.
- Video/demo claims that exceed docs.
- Breaking changes / deprecations that would flip an architecture call.
- Rate of shipping: stepwise product jumps vs compounding harness craft.
P10. Did the house Agentic AI Stack miss a box?
Overlaying two fast products is a stress test of the stack, not only of the products.
Secondaries
- Persistent computer / runtime host as a peer box vs Tools “execution environment/safe sandbox”?
- Identity / billing / subscription plane (Portal, X Premium, SuperGrok)?
- Distribution / bot directory / skills hub as a marketplace layer?
- Payments / 402 / agent commerce on Grok or Hermes?
- Evaluation / evidence packs / tracing as first-class?
- Voice/realtime as looking-glass family vs model modality?
- Operate-vs-rent host banner currently in Harness — still correct?
- Anything Grok/Hermes force us to split (harness loop vs host VM vs inference)?
Rabbit holes
- House rule: host banner in Harness not Model; where-it-runs = solid violet.
- HW-near AI: on-prem vs cloud does not factor; operate vs rent does.
P11. Economics, lock-in, and operate-vs-rent
What becomes abundant, what stays scarce, who can leave?
Secondaries
- Metering: tokens, computer hours, tool calls, bots.
- Skill/memory/computer-disk portability.
- Model swappability (Hermes) vs Grok-only cognition.
- Gateway/bot identity portability.
- Cost of a durable computer vs local machine.
- Enterprise vs consumer SKU gap.
- Open-source Hermes vs closed Grok computer — true exit or GitHub theater?
- Suite gravity of X vs Nous community gravity.
P12. Risks, failure modes, and contrarian views
Who disagrees that these are the two products to watch, and where do they fail?
Secondaries
- Prompt injection via X posts into Grok Bot / computer.
- Hermes YOLO on a personal machine vs Grok sandboxed cloud computer.
- Overclaiming “persistent computer” for ephemeral containers.
- Multi-bot chaos without a board.
- “Bots” as a dead-end vs harness-as-product.
- OpenClaw / Claude Code / Codex as the real peers, Grok/Hermes as cousins.
- Regulatory: agent on a social network vs local harness.
- Rogue-superintelligence / automation global-problem link (bounded, not sci-fi).
P13. Thought-leader and X/video signal (not sole facts)
What do named builders claim the other side got right?
Secondaries
- Teknium / Nous / Hermes maintainers on Grok.
- Elon / xAI / Grok team on agents, computer, bots.
- Karpathy / other harness voices mapping both.
- Demo videos: what is actually on screen (computer desktop, terminal, chat)?
- All-In / Moonshots takes on Grok agents vs open harnesses.
- Contrarian X: Grok computer is a toy; Hermes is ungovernable.
- Claims of copying Clawdbot / OpenClaw / Claude Cowork.
- Which claims survive primary-doc check?
Unresolved (seed)
- Exact Grok persistent-computer SLA, isolation, and disk retention (may be not public).
- Exact Nous Portal SKU matrix vs hosted computer (may be inference-only).
- Whether Grok custom bots share one computer or many.
- Independent eval of either product’s long-horizon reliability.
Saturation notes
Ledger created 2026-08-25 before index/sections (bundle ledger-first). Round 1 (same day) answered P1–P12 in chapters 01–08 from official docs. P13 podcast lane is thin (no Grok Bot episode). Unresolved host SLA / Hermes Cloud desktop roadmap remain marked. Contrarian scan in 08. No new primary added after the doc fetch.