cd ~/series/demystifying-ai
Jev: The Model That Gave Up Talking
EP 35 NEW RELEASE Sep 25, 2026 5 min

Jev: The Model That Gave Up Talking

1,500 HN points. 17 million views on X. The loudest "can't hallucinate" debate of 2026. And under all that noise: a genuinely new primitive for building AI software. Here's what Jev actually is, what it gets right, and where TypeSafe is overselling it.

Share:
// TL;DR

Jev is a decision model, not a language model. You send it state and typed questions. It returns structured answers — a choice, a score, or a yes/no probability — in a single parallel pass. No text generation. No string parsing. The output can never violate your schema because the schema is the output. Built by Diogo Almeida (ex-OpenAI; worked on RLHF, InstructGPT, ChatGPT, GPT-4) after two years in stealth. $40M seed from DCVC. Independent tests show 13x faster and 22x cheaper than Claude Opus 5 on classification. TypeSafe's own claim of 193x faster did not hold up. The "can't hallucinate" marketing is misleading — it can absolutely return wrong answers. But the core insight is real: most AI calls in production software are decision calls, not generation calls, and we've been using the wrong tool.

175ms
Median Latency
OpenRouter independent
$0.042
Per 1M Input
Output: $0 (free)
1,500
HN Points
426 comments, day 1
$40M
Seed (DCVC)
2 years in stealth

# The Thesis: Intelligence Without Strings

Count the LLM calls in your agent. Now count how many of them actually need text generation. Route this ticket to billing or engineering? LLM call. Is this tool invocation safe to execute? LLM call. Which model should handle this request? LLM call. Should we escalate to a human? LLM call. In each case, you're paying a model to generate a string, parsing that string, validating it, sometimes retrying when it hallucinates a JSON field — all to get a bounded answer your code could act on directly.

Jev's argument: stop generating strings for decisions that have known answer sets. Instead of asking an LLM "classify this as billing or technical and respond in JSON" and hoping the response is valid, you define the options in code and the model picks one. The output is typed by construction. There is no string to parse. There is no schema to validate. There is nothing to retry.

TWO PATHS TO "ROUTE THIS TICKET"
THE LLM WAY (~2.3s, $0.0024)
Prompt: "Classify as billing or technical. Respond in JSON."
Generate tokens one by one...
JSON.parse() — hope it's valid
zod.safeParse() — validate schema
Failed? Retry (back to top) ↺
"billing" — finally usable
THE JEV WAY (~175ms, $0.00011)
choice("dept", { billing, technical })
Parallel pass — all options at once
{ choice: "billing", confidence: 0.94 }
Typed. No parse. No validate. No retry.
Steps 3 vs 7
Can fail mid-pipeline No
Confidence score Built in
THE KAHNEMAN SPLIT
System Two (LLMs)

Slow. Deliberate. Generates text one token at a time. Powerful for reasoning, writing, coding — anything where the output space is unbounded. This is what GPT, Claude, Gemini do.

System One (Jev)

Fast. Intuitive. Evaluates all options in a single parallel pass. Built for decisions where the answer set is known in advance: routing, classification, scoring, gating, verification.

Named after Daniel Kahneman's Thinking, Fast and Slow. "Jev" comes from William Stanley Jevons — the Jevons paradox: cheaper resource → more consumption. TypeSafe's bet: cheaper intelligence means AI moves into every if statement.

# Three Primitives. One Endpoint.

Jev has one API endpoint: POST /v1/systemone. Every request has a state (the thing being judged), a model, and one or more questions. All questions see the same state, evaluate in parallel, and return under the key you chose. Three question types:

CHOICE

Pick one from up to 255 labelled options. Returns the winner, a probability distribution over all options, and a confidence score.

Your code: switch/case
SCORE

Rate on an ordered 2-10 level scale you define. Returns a probability-weighted fractional position (can land between levels) and confidence.

Your code: threshold
NOUL

Yes or no, returned as a probability from 0 to 1. Calibrated: when Jev says 0.8, outcomes should match ~80% of the time across many predictions.

Your code: if/else
HOW JEV EVALUATES — ONE STATE, THREE QUESTIONS, ONE PASS
STATE
"Charged twice. I need this fixed today."
+ plan: annual, history: [...customer records]
CHOICE — "dept"
billing 94% technical 4% account 2%
SCORE — "urgency"
2.87 / 3 — "needs attention today"
NOUL — "wants_refund"
0.73
probably yes
TYPED RESPONSE — 175ms TOTAL
{ dept: ChoiceAnswer, urgency: ScoreAnswer, wants_refund: NoulAnswer }
All three questions share the same state. Evaluated in parallel. Single HTTP round-trip.

# What It Looks Like in Practice

The TypeScript SDK infers answer types from your question map. department is typed as a ChoiceAnswer, urgency as ScoreAnswer. No JSON.parse. No zod.safeParse. No retry.

CONFIDENCE-GATED ROUTING — THE CORE PATTERN
import { TypeSafeClient, choice, score, noul }
  from "@typesafe-ai/sdk";

const client = new TypeSafeClient();

// One call. Three questions. Single round trip.
const r = await client.systemOne({
  model: "jev-1.13.0",  // pin in production
  state: {
    ticket: "Charged twice. I need this fixed today.",
    plan: "annual", history: customerHistory,
  },
  questions: {
    dept: choice("Which team handles this?", {
      billing: "Payment, refund, or subscription",
      technical: "Product bug or integration",
      account: "Login or access",
    }),
    urgency: score("How urgent?", [
      "Can wait", "Needs attention soon",
      "Needs attention today",
    ]),
    wants_refund: noul("Is the customer requesting a refund?"),
  },
});

// The confidence ladder — Jev decides, YOUR CODE acts
if (r.answers.dept.confidence > 0.9) {
  await route(r.answers.dept.choice);          // auto-route
} else if (r.answers.dept.confidence > 0.6) {
  await route(r.answers.dept.choice, { flag: true }); // route + flag
} else {
  await escalate(ticket);                      // human decides
}
CONFIDENCE-GATED DECISION FLOW
Jev returns: { choice: "billing", confidence: ? }
> 0.9
AUTO-ROUTE
Jev is confident. Route directly to billing. No human in the loop. Fastest path.
0.6–0.9
ROUTE + FLAG
Moderate confidence. Route to billing but flag for review. Human audits afterward.
< 0.6
ESCALATE TO HUMAN
Jev isn't sure. Skip automation entirely. A human decides. You don't pay for a wrong answer.
The thresholds live in YOUR code, not in Jev. You decide the risk tolerance. Jev provides the probability.
ALSO AVAILABLE VIA
Vercel AI Gateway: typesafe-ai/jev OpenRouter: typesafe/jev-1.13 Cloudflare AI Gateway TanStack AI: decide() LangChain Raw HTTP: POST /v1/systemone

# Already in Production

Ten days after launch, Jev is already being wired into real systems. Three cases worth reading:

Vega — SOC Alert Triage

Put Jev as a gate in front of their agent-powered security operations center. One Jev call per alert (Choice + Nouls) before the expensive agent loop. Result: closed 15-33% of alerts automatically, 98% accuracy verified by human analysts. Their agent normally costs $0.34 and 2 minutes per alert. Jev's gate: fractions of a cent, sub-second.

Nym — Agent Tool Selection

Replaced per-turn LLM tool selection with a Jev call. Context + available tools go in, probabilities come out. If Jev is confident, the LLM skips reasoning and generates the tool call directly (thinking off, cached prefix reused). If Jev isn't confident, falls back to DeepSeek with full reasoning. Result: 36% lower generative-model spend.

LangChain — Agent Harness Pattern

Published a "harness" pattern: Jev as middleware that checks tool calls for risky actions before the tool executes. Same pattern that Claude Code, Codex, and Cursor use for their internal safety classifiers — but now available to anyone. A Noul per tool call: "Is this action dangerous?" Act or block based on the probability.

THE JEV GATE PATTERN — USED BY VEGA, NYM, AND LANGCHAIN
INCOMING REQUEST
alert / ticket / tool call / user message
JEV GATE
~175ms · ~$0.0001
Choice + Noul(s) — single round-trip
confident
DIRECT ACTION
Auto-close alert
Route ticket
Execute tool
⚡ sub-second, fractions of a cent
uncertain
LLM AGENT
Full reasoning
Chain-of-thought
~$0.34, ~2 min
🧠 expensive but only when needed
Vega SOC
15-33% auto-closed
98% accuracy
Nym
36% cost reduction
on generative spend
LangChain
Tool safety gate
before execution

# The Numbers (Independent, Not Marketing)

TypeSafe claims 40-200x faster and 40-400x cheaper. Independent tests from OpenRouter, AY Automate, and JevBench tell a different but still impressive story. The headline multiples come from comparing against LLMs with full chain-of-thought enabled — apples to oranges.

Metric Jev 1.13 Opus 5 GPT-5.4 nano Haiku 4.5
Banking77 (77-way) 81.0% 84.4% — —
Latency p50 175ms 2,266ms 670ms 890ms
Latency p99 353ms 3,835ms — —
Cost / 1K decisions $0.11 $2.42 $0.52 $1.15
Invalid responses 0 0 3 0
Sources: OpenRouter Banking77 (3,080 utterances, Sep 22); AY Automate (791 decisions, Sep 19). Jev's slowest call (1.42s) was faster than Opus's fastest (~1.9s).
LATENCY p50 (LOWER IS BETTER)
Jev 1.13 175ms
GPT-5.4 nano 670ms
Haiku 4.5 890ms
Claude Opus 5 2,266ms
COST PER 1K DECISIONS (LOWER IS BETTER)
Jev 1.13 $0.11
GPT-5.4 nano $0.52
Haiku 4.5 $1.15
Claude Opus 5 $2.42
Jev's slowest call (1.42s p99) was still faster than Opus 5's fastest (~1.9s). Cost includes output tokens — Jev charges $0 for output.

The real comparison point is small models, not frontier ones. Against GPT-5.4 nano and Gemini Flash-Lite, Jev is 2-3.6x faster and 4.7-7.5x cheaper — real but not two orders of magnitude. Against frontier (Opus 5), it's 13x faster and 22x cheaper. The accuracy gap is consistent: 3-5 percentage points behind frontier LLMs on classification. The question is whether confidence-gated escalation closes that gap in practice. Vega's 98% accuracy at triage suggests it can.

# Under the Hood: RLCD and the Architecture Mystery

TypeSafe describes two technical innovations: a parallel sampler (the model evaluates all questions at once instead of generating tokens sequentially) and RLCD — Reinforcement Learning for Calibrated Decisions (the training objective).

Method Optimizes for Produces
RLHF Human preference — responses raters like Chat models (ChatGPT, Claude)
RLVR Verifiable rewards — programmatically checkable Reasoning models (o3, DeepSeek R1)
RLCD "Epistemically honest probabilities" Decision models (Jev)

The calibration claim is specific: when Jev says a probability is 0.8, outcomes should be correct ~80% of the time across many predictions. This is a batch property, not a guarantee about any single answer. TypeSafe's own docs are explicit about this distinction.

THE THREE RL TRAINING PARADIGMS
2020
RLHF
Signal: Human raters pick preferred outputs
Makes: Chat models — fluent, helpful, aligned
Ships: ChatGPT, Claude, Gemini
↓
2024
RLVR
Signal: Programmatic verifiers check answers
Makes: Reasoning models — chain-of-thought
Ships: o3, DeepSeek R1, QwQ
↓
2026
RLCD
Signal: "Epistemically honest probabilities"
Makes: Decision models — calibrated, typed
Ships: Jev (only known implementation)
Each paradigm produces a different kind of model by changing what the reward function optimizes for.
WHAT'S NOT PUBLISHED

The architecture. The parameter count. The base model. The reward function. The training data. Whether a transformer sits underneath. Whether "single query" means one forward pass. No paper, no model card, no ablation. The CEO told HN the architecture is "close to the chest for now." You are trusting a black box with very strong marketing claims. The parallel sampler is partially verifiable from the outside (latency should scale with state size, not question count). RLCD is not verifiable at all — nobody outside TypeSafe has reproduced it because there is nothing published to reproduce.

# The "Can't Hallucinate" Debate

The HN thread was originally titled "New frontier model 40-400x cheaper and 20-200x faster" — changed within an hour. The most-replied comment put it plainly: Jev can't emit an invalid type, but it can absolutely emit a wrong valid value. Calling a misclassification "not hallucination" is a category argument, not a safety guarantee.

WHAT'S TRUE

✓ Output always matches your schema

✓ Can't invent JSON fields or return free text

✓ Zero invalid responses in all independent tests

✓ Probabilities are calibrated (batch, not per-call)

WHAT'S MISLEADING

✗ Can pick the wrong option with high confidence

✗ 3-5pp less accurate than frontier LLMs

✗ "193x faster" uses cherry-picked baselines

✗ "Can't hallucinate" conflates format with correctness

The founder's HN response was direct: "I don't think it's fair to say a random forest hallucinates in the way LLMs do" — and conceded: "because these models are probabilistic, it's also possible to be confidently wrong." He also confirmed Jev is, at core, "basically a zero-shot classifier" with better calibration. Fair enough. The marketing oversells it. The product is still genuinely useful.

LIMITATIONS
  • ● Cannot generate text, code, or explanations. Jev complements LLMs; it does not replace them. If your task needs a string output — a reply, a summary, a diff — Jev is the wrong tool.
  • ● No images, audio, or PDFs. Text and JSON state only. No multimodal input. If the decision depends on a screenshot or document, you need an LLM or OCR pipeline first.
  • ● Architecture is a black box. No paper, no weights, no parameter count. Proprietary API only. Several open-source reproductions (openjev, jev-on-a-laptop) exist but none replicate RLCD calibration.
  • ● Early access. Direct API is waitlisted. OpenRouter, Vercel, and Cloudflare serve it without a waitlist but as a gateway-routed call. No free tier — $0.042/M is cheap but not zero.
  • ● Noul vs Boolean naming inconsistency. TypeSafe calls it "noul". Vercel AI SDK maps it to "boolean". The two schemas look close enough to mix and then fail validation. Pick one route per codebase.
// Bottom Line

Strip away the marketing and the HN drama and Jev is this: a fast, cheap, typed classifier with calibrated confidence scores and an API designed for software, not chat. That sounds boring until you count how many LLM calls in your codebase are doing exactly this job at 10-20x the cost and latency. Vega's SOC closed a third of its alerts with one Jev call each. Nym cut 36% of its generative-model spend. LangChain published the guardrail pattern that Cursor and Claude Code have kept proprietary. The 193x speed claim is marketing fiction — independent tests show 3-13x. The "can't hallucinate" framing is misleading — it can be confidently wrong. But if you build agents and your cost dashboard makes you flinch, the pattern is clear: let Jev triage, let the LLM reason, and put the confidence threshold in your code where it belongs.

NEXT EPISODE
Monday
#36 Upcoming

GPT-6 Sol & Luna: The 8x Honesty Upgrade

Half the price, 8x less deception. But warning circumvention barely moved.

Enjoyed this?

New episodes Mon, Wed, Sat.