Jev is a decision model, not a language model. You send it state and typed questions. It returns structured answers — a choice, a score, or a yes/no probability — in a single parallel pass. No text generation. No string parsing. The output can never violate your schema because the schema is the output. Built by Diogo Almeida (ex-OpenAI; worked on RLHF, InstructGPT, ChatGPT, GPT-4) after two years in stealth. $40M seed from DCVC. Independent tests show 13x faster and 22x cheaper than Claude Opus 5 on classification. TypeSafe's own claim of 193x faster did not hold up. The "can't hallucinate" marketing is misleading — it can absolutely return wrong answers. But the core insight is real: most AI calls in production software are decision calls, not generation calls, and we've been using the wrong tool.
# The Thesis: Intelligence Without Strings
Count the LLM calls in your agent. Now count how many of them actually need text generation. Route this ticket to billing or engineering? LLM call. Is this tool invocation safe to execute? LLM call. Which model should handle this request? LLM call. Should we escalate to a human? LLM call. In each case, you're paying a model to generate a string, parsing that string, validating it, sometimes retrying when it hallucinates a JSON field — all to get a bounded answer your code could act on directly.
Jev's argument: stop generating strings for decisions that have known answer sets. Instead of asking an LLM "classify this as billing or technical and respond in JSON" and hoping the response is valid, you define the options in code and the model picks one. The output is typed by construction. There is no string to parse. There is no schema to validate. There is nothing to retry.
Slow. Deliberate. Generates text one token at a time. Powerful for reasoning, writing, coding — anything where the output space is unbounded. This is what GPT, Claude, Gemini do.
Fast. Intuitive. Evaluates all options in a single parallel pass. Built for decisions where the answer set is known in advance: routing, classification, scoring, gating, verification.
Named after Daniel Kahneman's Thinking, Fast and Slow. "Jev" comes from William Stanley Jevons — the Jevons paradox: cheaper resource → more consumption. TypeSafe's bet: cheaper intelligence means AI moves into every if statement.
# Three Primitives. One Endpoint.
Jev has one API endpoint: POST /v1/systemone. Every request has a state (the thing being judged), a model, and one or more questions. All questions see the same state, evaluate in parallel, and return under the key you chose. Three question types:
Pick one from up to 255 labelled options. Returns the winner, a probability distribution over all options, and a confidence score.
Rate on an ordered 2-10 level scale you define. Returns a probability-weighted fractional position (can land between levels) and confidence.
Yes or no, returned as a probability from 0 to 1. Calibrated: when Jev says 0.8, outcomes should match ~80% of the time across many predictions.
# What It Looks Like in Practice
The TypeScript SDK infers answer types from your question map. department is typed as a ChoiceAnswer, urgency as ScoreAnswer. No JSON.parse. No zod.safeParse. No retry.
import { TypeSafeClient, choice, score, noul } from "@typesafe-ai/sdk"; const client = new TypeSafeClient(); // One call. Three questions. Single round trip. const r = await client.systemOne({ model: "jev-1.13.0", // pin in production state: { ticket: "Charged twice. I need this fixed today.", plan: "annual", history: customerHistory, }, questions: { dept: choice("Which team handles this?", { billing: "Payment, refund, or subscription", technical: "Product bug or integration", account: "Login or access", }), urgency: score("How urgent?", [ "Can wait", "Needs attention soon", "Needs attention today", ]), wants_refund: noul("Is the customer requesting a refund?"), }, }); // The confidence ladder — Jev decides, YOUR CODE acts if (r.answers.dept.confidence > 0.9) { await route(r.answers.dept.choice); // auto-route } else if (r.answers.dept.confidence > 0.6) { await route(r.answers.dept.choice, { flag: true }); // route + flag } else { await escalate(ticket); // human decides }
typesafe-ai/jev OpenRouter: typesafe/jev-1.13 Cloudflare AI Gateway TanStack AI: decide() LangChain Raw HTTP: POST /v1/systemone # Already in Production
Ten days after launch, Jev is already being wired into real systems. Three cases worth reading:
Put Jev as a gate in front of their agent-powered security operations center. One Jev call per alert (Choice + Nouls) before the expensive agent loop. Result: closed 15-33% of alerts automatically, 98% accuracy verified by human analysts. Their agent normally costs $0.34 and 2 minutes per alert. Jev's gate: fractions of a cent, sub-second.
Replaced per-turn LLM tool selection with a Jev call. Context + available tools go in, probabilities come out. If Jev is confident, the LLM skips reasoning and generates the tool call directly (thinking off, cached prefix reused). If Jev isn't confident, falls back to DeepSeek with full reasoning. Result: 36% lower generative-model spend.
Published a "harness" pattern: Jev as middleware that checks tool calls for risky actions before the tool executes. Same pattern that Claude Code, Codex, and Cursor use for their internal safety classifiers — but now available to anyone. A Noul per tool call: "Is this action dangerous?" Act or block based on the probability.
# The Numbers (Independent, Not Marketing)
TypeSafe claims 40-200x faster and 40-400x cheaper. Independent tests from OpenRouter, AY Automate, and JevBench tell a different but still impressive story. The headline multiples come from comparing against LLMs with full chain-of-thought enabled — apples to oranges.
| Metric | Jev 1.13 | Opus 5 | GPT-5.4 nano | Haiku 4.5 |
|---|---|---|---|---|
| Banking77 (77-way) | 81.0% | 84.4% | — | — |
| Latency p50 | 175ms | 2,266ms | 670ms | 890ms |
| Latency p99 | 353ms | 3,835ms | — | — |
| Cost / 1K decisions | $0.11 | $2.42 | $0.52 | $1.15 |
| Invalid responses | 0 | 0 | 3 | 0 |
The real comparison point is small models, not frontier ones. Against GPT-5.4 nano and Gemini Flash-Lite, Jev is 2-3.6x faster and 4.7-7.5x cheaper — real but not two orders of magnitude. Against frontier (Opus 5), it's 13x faster and 22x cheaper. The accuracy gap is consistent: 3-5 percentage points behind frontier LLMs on classification. The question is whether confidence-gated escalation closes that gap in practice. Vega's 98% accuracy at triage suggests it can.
# Under the Hood: RLCD and the Architecture Mystery
TypeSafe describes two technical innovations: a parallel sampler (the model evaluates all questions at once instead of generating tokens sequentially) and RLCD — Reinforcement Learning for Calibrated Decisions (the training objective).
| Method | Optimizes for | Produces |
|---|---|---|
| RLHF | Human preference — responses raters like | Chat models (ChatGPT, Claude) |
| RLVR | Verifiable rewards — programmatically checkable | Reasoning models (o3, DeepSeek R1) |
| RLCD | "Epistemically honest probabilities" | Decision models (Jev) |
The calibration claim is specific: when Jev says a probability is 0.8, outcomes should be correct ~80% of the time across many predictions. This is a batch property, not a guarantee about any single answer. TypeSafe's own docs are explicit about this distinction.
Makes: Chat models — fluent, helpful, aligned
Ships: ChatGPT, Claude, Gemini
Makes: Reasoning models — chain-of-thought
Ships: o3, DeepSeek R1, QwQ
Makes: Decision models — calibrated, typed
Ships: Jev (only known implementation)
The architecture. The parameter count. The base model. The reward function. The training data. Whether a transformer sits underneath. Whether "single query" means one forward pass. No paper, no model card, no ablation. The CEO told HN the architecture is "close to the chest for now." You are trusting a black box with very strong marketing claims. The parallel sampler is partially verifiable from the outside (latency should scale with state size, not question count). RLCD is not verifiable at all — nobody outside TypeSafe has reproduced it because there is nothing published to reproduce.
# The "Can't Hallucinate" Debate
The HN thread was originally titled "New frontier model 40-400x cheaper and 20-200x faster" — changed within an hour. The most-replied comment put it plainly: Jev can't emit an invalid type, but it can absolutely emit a wrong valid value. Calling a misclassification "not hallucination" is a category argument, not a safety guarantee.
✓ Output always matches your schema
✓ Can't invent JSON fields or return free text
✓ Zero invalid responses in all independent tests
✓ Probabilities are calibrated (batch, not per-call)
✗ Can pick the wrong option with high confidence
✗ 3-5pp less accurate than frontier LLMs
✗ "193x faster" uses cherry-picked baselines
✗ "Can't hallucinate" conflates format with correctness
The founder's HN response was direct: "I don't think it's fair to say a random forest hallucinates in the way LLMs do" — and conceded: "because these models are probabilistic, it's also possible to be confidently wrong." He also confirmed Jev is, at core, "basically a zero-shot classifier" with better calibration. Fair enough. The marketing oversells it. The product is still genuinely useful.
- ● Cannot generate text, code, or explanations. Jev complements LLMs; it does not replace them. If your task needs a string output — a reply, a summary, a diff — Jev is the wrong tool.
- ● No images, audio, or PDFs. Text and JSON state only. No multimodal input. If the decision depends on a screenshot or document, you need an LLM or OCR pipeline first.
- ● Architecture is a black box. No paper, no weights, no parameter count. Proprietary API only. Several open-source reproductions (openjev, jev-on-a-laptop) exist but none replicate RLCD calibration.
- ● Early access. Direct API is waitlisted. OpenRouter, Vercel, and Cloudflare serve it without a waitlist but as a gateway-routed call. No free tier — $0.042/M is cheap but not zero.
- ● Noul vs Boolean naming inconsistency. TypeSafe calls it "noul". Vercel AI SDK maps it to "boolean". The two schemas look close enough to mix and then fail validation. Pick one route per codebase.
Strip away the marketing and the HN drama and Jev is this: a fast, cheap, typed classifier with calibrated confidence scores and an API designed for software, not chat. That sounds boring until you count how many LLM calls in your codebase are doing exactly this job at 10-20x the cost and latency. Vega's SOC closed a third of its alerts with one Jev call each. Nym cut 36% of its generative-model spend. LangChain published the guardrail pattern that Cursor and Claude Code have kept proprietary. The 193x speed claim is marketing fiction — independent tests show 3-13x. The "can't hallucinate" framing is misleading — it can be confidently wrong. But if you build agents and your cost dashboard makes you flinch, the pattern is clear: let Jev triage, let the LLM reason, and put the confidence threshold in your code where it belongs.
GPT-6 Sol & Luna: The 8x Honesty Upgrade
Half the price, 8x less deception. But warning circumvention barely moved.
Enjoyed this?
New episodes Mon, Wed, Sat.