Jev — The First Non-Hallucinating AI Model That Doesn't Generate Strings
TypeSafe AI emerged from stealth today with a bold claim: they've built a new class of AI model — called System One Models — that can't hallucinate, runs 100x faster than frontier LLMs, and gives up string generation entirely.
Their first public model, Jev, is a machine-native intelligence designed to slot directly into software as a function call. Unstructured state in, typed probabilistic decisions out. No chat. No prose. No hallucination.
What Makes It Different
The comparison with traditional LLMs is striking:
- Outputs: Type-safe structured values instead of strings. The model can't emit an invalid type — it's mathematically impossible. Your code doesn't need to parse and validate text responses.
- Speed: 70ms–500ms end-to-end, versus 3–329 seconds for frontier LLMs. Two orders of magnitude faster.
- Cost: $0.042 per million input tokens. Output tokens are effectively free — "too cheap to meter."
- Sampling: Parallel instead of sequential. Jev generates all outputs in a single query rather than one token at a time.
- Calibrated confidence: Every output includes a probability. Higher confidence means higher accuracy — consistently. Traditional LLMs are systematically overconfident.
The Architecture Behind It
TypeSafe was founded by Diogo Almeida, who helped build the RLHF methods that became ChatGPT at OpenAI. After leaving, he spent two years building a new stack from scratch:
- RLCD (Reinforcement Learning for Calibrated Decisions) — a training method optimized for accurate probability calibration on System One tasks, instead of RLHF's optimization for human-preferred chat responses.
- Parallel sampler — all outputs computed in one pass, hardware-aware and incredibly efficient.
- No string generation — the critical insight. Strings are powerful but costly. By giving them up, Jev gains deterministic outputs with zero chance of going off the rails.
What It's Actually For
This isn't a chatbot competitor. Jev is built for the boring, critical infrastructure layer of AI-powered software:
- Smart if-statements — classify, route, score, extract, or branch where hand-written logic is too brittle. Surrounding code constrains the model's freedom.
- Real-time applications — 100ms response times mean AI decisions can slot into UX-critical paths.
- Map-reducing over big data — turn petabytes of unstructured data into features and insights at token prices that finally make sense.
- Verification/guardrails — score, judge, and detect jailbreaks of other LLM outputs with calibrated confidence.
The "Can't Hallucinate" Question
HN's community debated this hard. The nuance: Jev can still emit a wrong answer — but it will tell you its confidence. A 0.1 confidence score tells you to disregard the result. That's fundamentally different from an LLM that confidently outputs a wrong token and can't tell you it's uncertain. It's not an omniscient oracle; it's an honest one.
The distinction matters for automation. If a model can do a task 95% of the time but never tells you when it's in the 5%, you can't automate that task. Calibrated uncertainty is the feature that makes AI composable into reliable systems.
Bottom Line
TypeSafe's Jev represents a genuine architectural departure — not a better chat model, but a different category of AI altogether. For pipeline builders, agent orchestrators, and anyone wiring AI into deterministic software: this is the most interesting release of the week. Early access is open now.
Sources:
TypeSafe AI — Introducing System One Models & Jev
HN Discussion (921 pts)