Research · Notebooks · jaggedness

Cookbooks

Last updated: 2026-09-16

The console URL https://console.typesafe.ai/docs/cookbooks is a login wall (Google / email code) as of 16–17 September 2026. The same catalog is public at https://docs.typesafe.ai/ via llms.txt and the individual cookbook Markdown pages. This chapter is from those public pages, not from a logged-in console session.

What a cookbook is here

❓ Are these product SKUs, or worked notebooks?

They are worked notebooks: one state, a battery of Choice / Score / Noul questions, then code that thresholds, ranks, or composes. Several install typesafe-sdk plus cooksafe from https://pypi.typesafe.ai/. Some pin jev-latest, some pin jev-1.12 or jev-1.13. They are vendor demonstrations, not independent evals. Numbers below are the notebooks’ own.

Catalog (public)

❓ What jobs do the published cookbooks actually run?

Cookbook Job, in the notebook’s own terms
Self-consistency: nouls 14 Nouls × 15 repeats on one auto-insurance claim; mean per-question std dev 0.0102; band 0.30–0.70 → uncertain for a human.
Self-consistency: choices Same repeat idea on a moderation rubric; add an uncertain outcome and compare label agreement vs automatic-action share.
Parallel questions 13-question GDPR Wikipedia briefing (~54k characters). One batched call vs 13 singles: 12.2× cheaper, 10.0× faster, answers unchanged. Pins jev-1.12.
Re-ranking BM25 shortlist of 30 passages × 40 CLERC legal queries, then one TypeSafe question per pair. Top-1 5% → 18%; top-10 38% → 62%.
Line-by-line search GitHub Terms of Service: one request scores 218 line ids (Choice) and a Noul “does the document contain an answer?”
Structure recovery Two requests rebuild Markdown from stripped plain text (unwrap lines; classify heading/list/code/callout).
Function calling Natural-language trading requests → ordinary typed functions; closed-set arguments with confidence (weakest argument is the call’s confidence).
Skill suggestion Hermes catalog, 182 skills. Request 1 ranks all and asks whether a skill is needed; request 2 re-reads top three in full and may reject all. Claims wrong skill-loads drop by more than half.
Knowledge graph entity alignment 450 candidate pairs from two beer catalogues. One Score with three levels: merge / leave unlinked / hand to a curator. No extra threshold.
Classifying RAG passages One request per retrieved passage; code keeps, flags (contradiction), or drops (injection).
Double-checking citations One Choice per citation vs source: supports / says nothing / contradicts; confidence 0.8 auto-accept in the notebook. Planted fabricated quotes caught.
Guardrails for LLMs One request on the way in and the way out: Nouls for hazards (“is this a jailbreak?”) plus a Score for harm. Thresholds in your code: pass / review / block / support.
SDE cascade Two-stage structured-data extraction: mini → verify → reasoning model, aiming at most of the quality of a big reasoner at a fraction of the cost.
Date extraction Model names date parts; code resolves and validates; confidence-gated review. Matches the jaggedness “don’t compare dates in the model.”
Pre-parsed value extraction Regex finds candidate emails/phones/amounts; TypeSafe picks the requested span; code copies it verbatim.
Hierarchical classification Beam search over Choice probabilities on deep taxonomies (patents, retail, biomedical, source code).
Autoresearch feature discovery Loop proposes TypeSafe questions as numeric features for a CatBoost regressor.
Classification using confidence 75 SIC industry groups, one Choice; if confidence is low, report the broader division. Cutoff 0.9 in the notebook split 60 filings in half (confident half right 90%; other half 40%).

Architectural patterns (not notebooks) sit next door: speculative fan-out, confidence-gated routing, composite scoring, intent routing.

What this does to the “automation + audit trail” story

❓ Do the cookbooks add a “why” string, or more typed gates?

They add more typed gates. Guardrails, citation check, RAG passage class, skill suggestion, function calling — every one returns Choice/Score/Noul (and maybe confidence), then your code decides pass/review/block/load. None of them asks Jev to write an explanation. The audit trail is still the questions, the distributions, and the branch your program took.

Skill suggestion is the agentic-stack example in TypeSafe’s own docs: Jev as a tool in front of the harness, not as the harness. The winner’s name becomes one line in the agent’s system prompt.

Jaggedness (jev-1.13, reviewed 2026-09-17)

❓ Where does TypeSafe say Jev is weak?

Public page: Jev 1.13 jaggedness. Last reviewed 2026-09-17. Applies to alias jev-1.13. Summary in their table:

Failure mode Do this instead
Literal reading Write the exact condition; put boundaries in criteria
Math and numbers Keep arithmetic in code; do not interpolate Score levels into a precise magnitude
Date and time comparison Extract parts with Choice; compare in code
Indirection Fewer hops; name the relevant state
Large irrelevant state Filter first; bounded context window (see Models page)
Adversarial content Precise criteria; test; state is not treated as hostile by default
Contradictory instructions vs criteria Align them
Common-sense structural invariants Ask each decision one way; enforce identities in code
Generation Use a generative model

Counting, hex colors vs color names, assembly vs high-level languages, and “which date is first” are called out as unreliable. That is TypeSafe’s own limit list, one day after the stealth launch.

What this does not prove

❓ Can a cookbook replace a buyer’s labeled set?

No. Parallel-questions 12.2× is batch vs unbatched Jev, not Jev vs a frontier LLM. Re-rank 5% → 18% is BM25 vs BM25+Jev on 40 queries. Consistency 0.0102 is one claim, 15 draws. Treat the catalog as a map of intended jobs and as vendor-measured deltas inside those notebooks.

← Vendor riskResearch references →