Research · Notebooks · jaggedness
Cookbooks
Last updated: 2026-09-16
The console URL https://console.typesafe.ai/docs/cookbooks is a login wall (Google / email code) as of 16–17 September 2026. The same catalog is public at https://docs.typesafe.ai/ via llms.txt and the individual cookbook Markdown pages. This chapter is from those public pages, not from a logged-in console session.
What a cookbook is here
❓ Are these product SKUs, or worked notebooks?
They are worked notebooks: one state, a battery of Choice / Score / Noul questions, then code that thresholds, ranks, or composes. Several install typesafe-sdk plus cooksafe from https://pypi.typesafe.ai/. Some pin jev-latest, some pin jev-1.12 or jev-1.13. They are vendor demonstrations, not independent evals. Numbers below are the notebooks’ own.
Catalog (public)
❓ What jobs do the published cookbooks actually run?
| Cookbook | Job, in the notebook’s own terms |
|---|---|
| Self-consistency: nouls | 14 Nouls × 15 repeats on one auto-insurance claim; mean per-question std dev 0.0102; band 0.30–0.70 → uncertain for a human. |
| Self-consistency: choices | Same repeat idea on a moderation rubric; add an uncertain outcome and compare label agreement vs automatic-action share. |
| Parallel questions | 13-question GDPR Wikipedia briefing (~54k characters). One batched call vs 13 singles: 12.2× cheaper, 10.0× faster, answers unchanged. Pins jev-1.12. |
| Re-ranking | BM25 shortlist of 30 passages × 40 CLERC legal queries, then one TypeSafe question per pair. Top-1 5% → 18%; top-10 38% → 62%. |
| Line-by-line search | GitHub Terms of Service: one request scores 218 line ids (Choice) and a Noul “does the document contain an answer?” |
| Structure recovery | Two requests rebuild Markdown from stripped plain text (unwrap lines; classify heading/list/code/callout). |
| Function calling | Natural-language trading requests → ordinary typed functions; closed-set arguments with confidence (weakest argument is the call’s confidence). |
| Skill suggestion | Hermes catalog, 182 skills. Request 1 ranks all and asks whether a skill is needed; request 2 re-reads top three in full and may reject all. Claims wrong skill-loads drop by more than half. |
| Knowledge graph entity alignment | 450 candidate pairs from two beer catalogues. One Score with three levels: merge / leave unlinked / hand to a curator. No extra threshold. |
| Classifying RAG passages | One request per retrieved passage; code keeps, flags (contradiction), or drops (injection). |
| Double-checking citations | One Choice per citation vs source: supports / says nothing / contradicts; confidence 0.8 auto-accept in the notebook. Planted fabricated quotes caught. |
| Guardrails for LLMs | One request on the way in and the way out: Nouls for hazards (“is this a jailbreak?”) plus a Score for harm. Thresholds in your code: pass / review / block / support. |
| SDE cascade | Two-stage structured-data extraction: mini → verify → reasoning model, aiming at most of the quality of a big reasoner at a fraction of the cost. |
| Date extraction | Model names date parts; code resolves and validates; confidence-gated review. Matches the jaggedness “don’t compare dates in the model.” |
| Pre-parsed value extraction | Regex finds candidate emails/phones/amounts; TypeSafe picks the requested span; code copies it verbatim. |
| Hierarchical classification | Beam search over Choice probabilities on deep taxonomies (patents, retail, biomedical, source code). |
| Autoresearch feature discovery | Loop proposes TypeSafe questions as numeric features for a CatBoost regressor. |
| Classification using confidence | 75 SIC industry groups, one Choice; if confidence is low, report the broader division. Cutoff 0.9 in the notebook split 60 filings in half (confident half right 90%; other half 40%). |
Architectural patterns (not notebooks) sit next door: speculative fan-out, confidence-gated routing, composite scoring, intent routing.
What this does to the “automation + audit trail” story
❓ Do the cookbooks add a “why” string, or more typed gates?
They add more typed gates. Guardrails, citation check, RAG passage class, skill suggestion, function calling — every one returns Choice/Score/Noul (and maybe confidence), then your code decides pass/review/block/load. None of them asks Jev to write an explanation. The audit trail is still the questions, the distributions, and the branch your program took.
Skill suggestion is the agentic-stack example in TypeSafe’s own docs: Jev as a tool in front of the harness, not as the harness. The winner’s name becomes one line in the agent’s system prompt.
Jaggedness (jev-1.13, reviewed 2026-09-17)
❓ Where does TypeSafe say Jev is weak?
Public page: Jev 1.13 jaggedness. Last reviewed 2026-09-17. Applies to alias jev-1.13. Summary in their table:
| Failure mode | Do this instead |
|---|---|
| Literal reading | Write the exact condition; put boundaries in criteria |
| Math and numbers | Keep arithmetic in code; do not interpolate Score levels into a precise magnitude |
| Date and time comparison | Extract parts with Choice; compare in code |
| Indirection | Fewer hops; name the relevant state |
| Large irrelevant state | Filter first; bounded context window (see Models page) |
| Adversarial content | Precise criteria; test; state is not treated as hostile by default |
| Contradictory instructions vs criteria | Align them |
| Common-sense structural invariants | Ask each decision one way; enforce identities in code |
| Generation | Use a generative model |
Counting, hex colors vs color names, assembly vs high-level languages, and “which date is first” are called out as unreliable. That is TypeSafe’s own limit list, one day after the stealth launch.
What this does not prove
❓ Can a cookbook replace a buyer’s labeled set?
No. Parallel-questions 12.2× is batch vs unbatched Jev, not Jev vs a frontier LLM. Re-rank 5% → 18% is BM25 vs BM25+Jev on 40 queries. Consistency 0.0102 is one claim, 15 draws. Treat the catalog as a map of intended jobs and as vendor-measured deltas inside those notebooks.