Research · Blog · X · Notes

Claims audit

Last updated: 2026-09-16

TypeSafe’s launch post says, in so many words, “Extraordinary claims require extraordinary evidence so see below for the receipts.”[1] This chapter is that audit. Every headline number is traced to the sentence that produced it, then to the hedge TypeSafe printed next to it, then to what independent testers and X actually showed in the first 24 hours.

Claim table (as of 16 September 2026)

Which launch claims survive contact with TypeSafe’s own footnotes, and which do not?

Claim as marketed Where it appears What the same vendor page says next Independent check Verdict
20–200× faster; 40–400× cheaper, output free Almeida’s X launch post; launch blog comparison table[1][40] End-to-end 70–500 ms vs 3–329 s for “frontier models”; $0.042 / MTok input, output “too cheap to meter”[1] Every: 777 judgments in <0.7 s, ~$0.0025; ~25× faster and ~1/580th the cost of Claude Fable 5.1 on a 12-passage test[17][19] Speed and list price are real enough to measure. The 200× / 400× band is a range over vendor workflows, not a universal SLA.
193.6× faster, 444.6× cheaper Homepage hero[2] Launch post: “this is where the claims of 193.6x faster, 444.6x cheaper on our home page comes from, and we expect that these are on the higher end of real world gains.” Workflows “made by individuals on our model capabilities team, so some bias could exist.”[1] No third party reproduced this pair of numbers Vendor best-case, self-labeled as the high end.
“Can’t hallucinate” / “Zero Hallucinations” Launch post; homepage[1][2] “Our number is not empirical. Schema matching is guaranteed, thus we can confidently add 0% into the plots.”[1] Docs: calibration does not guarantee an individual answer is correct.[7] Almeida has said a model can still be confidently wrong (Forbes interview paraphrased in secondary coverage).[16] Goedecke: calling a wrong Choice a “mistake” instead of a hallucination is a “semantic dodge.”[18] OrcaRouter: schema-valid and wrong is still wrong.[19] True for type errors. False as a claim about being correct. The 0% chart is a definition, not a measurement.
Similar / frontier intelligence on System One tasks Launch post; DCVC (“frontier-level intelligence at less than 100 ms”)[1][15] Workflow evals use the average of GPT-6 Astra and Claude Fable 5.1 as the reference, not ground truth. “We likely underestimate the relative performance of our model and DeepSeek’s models.” LLM baselines are forced through TypeSafe’s structured wrapper, “slower and more expensive than giving decisions without probabilities.”[1][30] OrcaRouter, reading the vendor dashboard: Jev 67.8% mean vs 74.1% best comparator; loses accuracy, wins cost/latency on all four workflows.[19] Every: caught 6 of 7 planted defects; Fable 5.1 caught 7/7.[17][19] Not “as smart as the frontier.” Vendor’s own plot puts Jev near a strong mid-tier judge, cheaper.
New model architecture + parallel sampler Launch post, homepage[1][2] No paper, no params, details “close to the chest” (secondary paraphrase).[19] Goedecke: prefill + one constrained token on an ordinary LLM already buys most of the speed; he measured 2–3× on Qwen2.5-1.5B vs unprefixed structured output, not 200×.[18] Latent Space notes community pushback that Jev is “not a general language model.”[39] Unpublished. Observed latency is consistent with skipping autoregressive decode. That does not prove a new architecture family.
RLCD is a new training method Launch post, primer[1][8] Named and contrasted with RLHF/RLVR. No loss, no ablations, no public recipe. Goedecke has not seen evidence the probabilities are more than ordinary logits.[18] Product name, not a paper. Calibration as a goal is real and old; whether TypeSafe hit it is unshown.
Output tokens free because architecture, not a promo Launch post, homepage FAQ title “Are these prices temporary or subsidized?”[1][2] “We can’t prove it isn’t subsidized; we’ll need the long-term to prove the sustainability of our pricing (which we expect to go down, not up).”[1] No public cost-of-goods. $40M seed can buy a lot of inference.[15][22] Vendor cannot currently prove the meter is unsubsidized. They said so.
Almeida co-invented ChatGPT X bio and launch post[40][1] InstructGPT paper lists Diogo Almeida among primary authors (starred, with Ouyang, Wu, Jiang, Wainwright, Mishkin, Christiano, Leike, Lowe).[20][34] AI Engineer bio: coauthor of InstructGPT, contributor to ChatGPT, coauthor of the GPT-4 technical report.[25] Paper is primary. “Co-invented ChatGPT” is a product-level slogan on top of a multi-author research paper. InstructGPT/RLHF co-authorship: confirmed. Solo “invented ChatGPT”: marketing inflation.
$200 million valuation Forbes headline[16] “according to a person familiar with the deal”[16] DCVC and FinSMEs confirm $40 million seed led by DCVC, not the post-money number.[15][22] $40M seed: confirmed. $200M: single anonymous source.
Doom at ~10 calls/s ≈ $7/hour as real-time AI X + launch post[1] “The demo is on structured state as a data structure with text, not on images (yet…). A non-AI doom bot could play better.”[1] Goedecke treats the latency as the interesting part, not the game skill.[18] X critics (lead-only) called pixel-less Doom less impressive. A latency demo on a text state, by the vendor’s own nuance box.
Path to an “AI-based economic revolution” / TFP +3% Manifesto footnote[3] Favorite definition: global TFP growth reaching 3% within five years and holding ten. Labeled as a mission, not a measurement. No economist replication. Podcasts through 11 September 2026 predate the launch. A bet, not a fact.

What TypeSafe already conceded

Did the company hedge the extraordinary claims in the same post that made them?

Yes. The launch post’s “Evidence / Technical Results” section is more careful than the X bullets and the homepage hero.[1] In order:

A reader who only saw the X thread saw 20–200× / 40–400× / “can’t hallucinate.” A reader who opened the blog saw the hedges. The gap between those two surfaces is the launch’s main information-design fact.

X in the first day, and Community Notes

Is there a Community Note on the launch posts, and what did the crowd actually argue?

No published X Community Note was found on the main @CompleteSkeptic or @typesafeai launch posts as of this research pass on 16 September 2026. Searches for “community note” plus Jev / TypeSafe / Almeida returned discussion, not a signed Note attached to the 17-million-view thread.[40] Absence of a Note is not endorsement; Notes lag viral posts, especially inside 24 hours.

What did show up on X, ranked by how load-bearing the argument is rather than by likes:

  1. Category correction. Several technical accounts (summarized in Latent Space’s 16 September digest) said Jev is not a general language model and should be read as a constrained decision / classifier-like model that cannot produce free-form text.[39] That matches TypeSafe’s docs. It contradicts any reader who heard “new frontier model” and pictured a ChatGPT competitor.
  2. Hallucination redefinition. Independent of Notes, this is the most repeated technical objection: eliminating invalid strings is not eliminating wrong judgments.[18][19]
  3. Prefill replica. After the announcement, people began trying “prefill the JSON and generate one token” on open models; Goedecke cites one such attempt and his own 2–3× Qwen measurement.[18] That is the closest thing to a community falsification test of the “new architecture” slogan.
  4. Subsidy suspicion. Pricing at $0.042 / MTok with free output, days after a $40M seed, reads to some accounts as adoption pricing. TypeSafe’s own blog already refuses to prove otherwise.[1]
  5. Credential inflation. “Co-invented ChatGPT” travels further than “starred coauthor of InstructGPT.” The paper is public.[20][34]
  6. Employee framing on HN. An account writing as a TypeSafe employee said the point is to put reasoning in code so systems stay inspectable — an alignment story, not a benchmark story.[38]

Doomers’ launch tracker recorded very high save-rate and view count on the founder thread; that is distribution, not validation.[37]

How to read the 0% hallucination chart

What error class is structurally impossible, and which one remains?

Two error classes get one word in English, “hallucination”:

  1. Schema / type error — the model emits a tool name, a key, or a sentence that was never in the allowed set. Jev cannot do this, because the allowed set is the question you sent.[1][7] That is the chart’s 0%.
  2. Wrong-but-typed error — the model picks billing when the ticket is about a crash bug. Jev can do this. Confidence-gated routing is how TypeSafe wants you to catch it.[9] That is not on the 0% chart.

BoundaryML’s older critique of constrained decoding is the prior art for this split: forcing a schema can raise parse success while leaving semantic error untouched, and can even hide it.[29] TypeSafe’s contribution, if RLCD works, would be making the probabilities honest so the remaining errors are detectable. That is the claim that still needs a reliability diagram on a buyer’s labels.

Receipts TypeSafe asked for, scored

If we take “extraordinary claims require extraordinary evidence” as the grading rubric the founder proposed, what grade does the launch get?

The honest one-line from the first 24 hours is the same sentence OrcaRouter printed: speed and cost survive contact with a third party; accuracy sits a notch below the frontier; the sample is too small for a production conclusion.[19]

← Architecture & RLCDCompany & founder →