Research · Full question map
Question ledger
Last updated: 2026-09-16
Research date: 2026-09-16. Launch window: 2026-09-15 early access.
This ledger is the master question map. Section files answer the primaries in prose. Status: stabilized after one primary-docs round, independent tests, and a contrarian scan. No new primary was added in the last research round.
Confirmed vs unverified (rumor-heavy rule)
- Confirmed: TypeSafe AI is a San Francisco lab founded in 2024; Diogo Almeida is CEO and an InstructGPT paper co-author; cofounders Erik Gafni (CTO) and Sasha Sheng (COO); $40 million seed led by DCVC announced 15 September 2026; Jev is the first public System One model, early access the same day; HTTP
POST https://api.typesafe.ai/v1/systemonewith model aliasjev-latest; three question types (Choice, Score, Noul); published list price $0.042 per million input tokens and $0 output; closed weights, US-hosted API, waitlist; vendor docs state Jev does not generate text. - Credible but unconfirmed independently: post-money valuation of $200 million (Forbes, a person familiar with the deal); vendor latency band 70–500 ms as a universal production SLA; that list prices are unsubsidized; RLCD internals; parameter count, training compute, and architecture paper (none published); that Jev matches “frontier intelligence” on System One tasks as a general claim.
- Weak / rumor: “zero hallucinations” as a claim about being correct; that System One is a new scaling axis rather than a constrained-output inference style; that local/open weights are coming; that All-In or Moonshots have discussed this SKU (transcripts through 2026-09-11 predate the launch); that an X Community Note has already attached to the launch thread (none found 2026-09-16).
Primary questions
P1. What is Jev, and what product class does TypeSafe claim it belongs to?
- S1.1 What is the one-line official pitch?
- S1.2 What is a System One model, and how is the Kahneman name being used?
- S1.3 What are the three primitives, and what does each return?
- S1.4 What is “state,” and how is it separated from questions?
- S1.5 What is Jev not (chatbot, coder, reasoner, local model)?
- S1.6 Is “System One Model” a new LTC or a packaging of structured-output LLMs?
- S1.7 What does the name Jev refer to?
- RH1. Why three primitives instead of one JSON schema?
P2. What is the claimed architecture, and what is actually public?
- S2.1 What three stack pieces does TypeSafe say it rebuilt?
- S2.2 What does “parallel sampler” mean in the launch post?
- S2.3 Why would skipping autoregressive decoding cut latency and output cost?
- S2.4 What would a prefills-and-one-token LLM baseline already buy?
- S2.5 What is unpublished (weights, params, paper)?
- S2.6 Can you run Jev locally?
- S2.7 What model aliases have appeared (
jev-latest,jev-1.12,jev-1.13.0)? - RH2. Is the Transformer-vs-RNN analogy a mechanism or a slogan?
P3. What is RLCD, and how does it differ from RLHF and RLVR?
- S3.1 What does each acronym optimize for?
- S3.2 What is calibration, in one sentence, and what does it not guarantee?
- S3.3 What is mode dropping, and why does TypeSafe blame RLHF for HITL?
- S3.4 Did Almeida actually co-invent RLHF, or is that marketing inflation?
- S3.5 Is RLCD a published algorithm or a product name?
- S3.6 What would a calibration plot have to show to support the claim?
- S3.7 How does confidence differ from a Noul probability?
- RH3. Does “the bitterest lesson” contradict Sutton, or extend him?
P4. Who built this, with whose money, and at what scale?
- S4.1 Founding year, location, headcount?
- S4.2 Almeida’s OpenAI/Google Brain record, dated?
- S4.3 Who are Gafni and Sheng?
- S4.4 Seed size, lead, other named investors?
- S4.5 Valuation — who said $200 million?
- S4.6 What is the manifesto’s economic bet (TFP)?
- S4.7 What is “We’re building prod, not God” refusing?
- RH4. Why announce from @CompleteSkeptic rather than the company account?
P5. What is actually confirmed about speed and price?
- S5.1 Published list price vs typical frontier input/output?
- S5.2 Vendor latency band and the West Coast laptop caveat?
- S5.3 Where do 193.6× faster and 444.6× cheaper come from?
- S5.4 What did Every measure independently?
- S5.5 What did the noul consistency cookbook print for time/cost vs Haiku and GPT-5.5?
- S5.6 Can TypeSafe prove the price is not subsidized?
- S5.7 What is the comparison class (reasoning frontier vs constrained structured output)?
- RH5. Does free output follow from the architecture, or from a promo?
P6. How good are the judgments, and who labeled “correct”?
- S6.1 What is a “workflow eval,” and what is the reference?
- S6.2 What accuracy numbers did OrcaRouter report from the vendor dashboard?
- S6.3 What did Every’s planted-defect test show versus Claude Fable 5.1?
- S6.4 Why is agreement with Astra/Fable not ground truth?
- S6.5 What bias does TypeSafe itself disclose?
- S6.6 Is there a public leaderboard, or only evals.typesafe.ai?
- S6.7 What would an honest buyer’s calibration check look like?
- RH6. Does skipping public benches (Almeida’s anti-bench essay) help or hide?
P7. What does “can’t hallucinate” actually mean?
- S7.1 What error class is structurally impossible?
- S7.2 What error class remains (wrong-but-typed)?
- S7.3 How does TypeSafe’s 0% plot get its number?
- S7.4 How does this differ from OpenAI Structured Outputs / grammars?
- S7.5 What does BoundaryML’s older critique say about constrained decoding quality?
- S7.6 Can confidence-gated routing turn residual error into HITL?
- S7.7 What happens at Choice cardinality 255?
- RH7. Is “hallucination” being redefined to win a chart?
P8. How do you call it, and what sits in the agentic stack?
- S8.1 Endpoint, auth, SDK languages?
- S8.2 How should questions be decomposed versus stuffed into one prompt?
- S8.3 Where does Jev sit relative to model / tools / skills / harness / HITL?
- S8.4 What does the Hermes skill-suggestion cookbook actually measure?
- S8.5 What does pi-jev do as a coding-agent gate?
- S8.6 Is Jev a harness, a tool, or a model?
- S8.7 Data residency, logging, and enterprise SKU — public or not?
- RH8. Does a decision primitive replace structured tool-calling, or sit beside it?
P9. What are the demos, and what do they not prove?
- S9.1 Doom: input modality, query rate, cost, caveat?
- S9.2 Wikiracing: cardinality and two-stage scoring?
- S9.3 Side-by-side vs GPT-5.6 Terra — disclosed advantages?
- S9.4 Map-reduce / real-time / verify-everything use cases — which are shipped vs sketched?
- S9.5 Why is a non-AI Doom bot a relevant comparison?
- S9.6 Images, audio, tools that write — in or out at launch?
- S9.7 Early-access vs GA?
P10. What is the strongest contrarian case?
- S10.1 Sean Goedecke: prefills already buy most of the speed?
- S10.2 Intelligence cap without test-time compute?
- S10.3 Vendor evals as self-reference to Astra/Fable?
- S10.4 Closed weights + unpublished architecture as a moat story?
- S10.5 JSON-mode / Outlines / DSPy / classifiers as substitutes?
- S10.6 Kahneman System 1 as an error-prone brand?
- S10.7 Who disagrees, and what would change their mind?
- RH10. If a frontier lab ships “Terra-System-One,” what is left?
P11. How should the trends and taxonomy be classified?
- S11.1 Is Jev a PTC, and of which LTC?
- S11.2 Is there an exponential trend here, or a stepwise product launch?
- S11.3 Weak vs strong convergence with agents / structured output / cheap inference?
- S11.4 Jevons Paradox: measured trend or naming bet?
- S11.5 What global problems does cheaper judgment touch (automation/UBI, education) — and which does it not?
- S11.6 What milestones would unlock a LAC of unattended semantic routing?
- S11.7 What stays unresolved after this pass?
P12. Vendor-risk lane (live-source rule)
- S12.1 Legal/governance: closed model, no paper, waitlist API
- S12.2 Cloud concentration: US-hosted only at launch
- S12.3 Model-continuity: aliases
jev-latestvs pinnedjev-1.12 - S12.4 Data/control-plane: state you send is the document
- S12.5 Agent-runtime risk: using Jev as a gate is still your policy
- S12.6 Lock-in via question graphs, evals, and workflow context
- S12.7 What a FRFI would still have to put on an E-23 registry if it used this
Rabbit holes chased
| ID | Trigger | Closure |
|---|---|---|
| RH1 | Three primitives vs one schema | Closed: docs treat Choice/Score/Noul as the entire output surface; JSON is not a fourth primitive. |
| RH2 | Transformer-vs-RNN analogy | Closed as slogan: launch post asserts it; no architecture paper. Speed is observed; mechanism is unpublished. |
| RH3 | Bitterest lesson vs Sutton | Closed: Almeida ranks task > data > compute > algorithms; InstructGPT fig. 31 is the exhibit. |
| RH4 | Personal X launch | Closed as color: Doomers tracked @CompleteSkeptic; not load-bearing. |
| RH5 | Free output | Closed: vendor says output is too cheap to meter because there is no autoregressive decode to bill; sustainability unproven. |
| RH6 | Anti-bench essay | Closed: Almeida argues benches are gamed; TypeSafe publishes workflow evals instead. Independent accuracy still thin. |
| RH7 | Hallucination redefinition | Closed: TypeSafe and critics agree the 0% is schema-match, not correctness. |
| RH10 | Lab replica | Open: no frontier-lab System One SKU as of 2026-09-16. |
Unresolved (with reason)
- Architecture, parameter count, training mix, and RLCD algorithm — unpublished.
- Whether list prices are subsidized — vendor says they cannot prove they are not.
- Production calibration on a buyer’s labeled set — no public reliability diagram.
- Data residency, ZDR, SOC reports, DPA — not on the public docs index reviewed.
- Moonshots / All-In commentary on this SKU — transcripts through 2026-09-11 predate 2026-09-15.
- GA date and open-weight path — not announced.
Saturation note
Primaries P1–P12 have answers or explicit unresolved flags. Contrarian scan integrated in chapter 08 and in P5–P7. Podcast pass is thin-signal, not a card quota.