Jev takes shared state and returns typed Choice, Score, or yes-or-no answers with probabilities. It does not draft paragraphs or open-ended chat. Construction software that already runs hundreds of route, score, and gate checks over bids, specs, and agent plans can treat those checks like a cheap utility call instead of a long generation.
TypeSafe left stealth on 15 September 2026 with Jev, its first System One model. System One is TypeSafe's name for fast, structured decision calls built for software. Founder Diogo Almeida, a former OpenAI researcher tied to RLHF work from the ChatGPT era, has argued that automation stalls when the only interface is free-form text. The Register described the audience in plain terms: Jev is built for machines that need a constrained answer.
Price and latency, as published
As of 17 September 2026, TypeSafe's launch materials list Jev at $0.042 per million input tokens with free output. TypeSafe's comparison table puts typical frontier generative input much higher and notes that generative output is often several times the input rate. Jev returns fixed answer distributions, so there are no output tokens to bill. Products that already fire large volumes of yes-or-no, score, and route questions pay for input state only.
TypeSafe quotes roughly 70 to 500 milliseconds end-to-end for System One-shaped queries. Its materials put frontier LLM structured-decision calls at three to 329 seconds, and a broader 40-200x faster band for Jev. Multiple typed questions share one state and run in parallel in a single request. Early-access writing from Every reported hundreds of judgments over dozens of documents returning in under a second for a fraction of a cent. Those are launch-week figures that line up with the latency and price TypeSafe is publishing.
Input tokens: $0.042 / MTok. Output tokens: FREE (too cheap to meter). -- TypeSafe, Introducing System One Models and Jev, 15 September 2026
What TypeSafe means by accuracy
TypeSafe's workflow evals do not score Jev against human-labeled ground truth. They measure agreement with a reference built from the average of GPT-6 Astra and Fable 5.1 on the same typed decision graph. On that yardstick, Jev lands near mid-tier frontier models on TypeSafe's published average while costing far less per case. That is useful if your current judge is already an expensive LLM. It is not the same as proving the answer is right on a construction corpus.
TypeSafe's biggest speed and cost multiples are its own lab highs on workflows its team wrote, and the company says they sit at the upper end of what customers should expect. Treat those poster numbers as a vendor claim, not a construction-site measurement.