Home Fashion Introducing System One Models and Jev
FASHION

Introducing System One Models and Jev

Introducing System One Models and Jev. TypeSafe AIManifestoOur TeamJoin WaitlistTypeSafe AIManifestoOur Team∵ Back Company NewsSep 15, 2026Introducing System One Models & JevDiogo Almeida, founder, TypeSafeModels have been superhuman at chat for years, so where is all the automation?

What happened

After two years in stealth, countless technical challenges, and research breakthroughs… I am beyond excited to announce that today, TypeSafe AI is releasing our first System One Model: a new class of frontier models built to make fast, structured decisions that software can use directly. Fun fact: a similar demo was what convinced us to go all-in in the direction of System One Models! We Give A FAQWhere do the names “System One Models” and “Jev” come from?

For reasons we will get into in the future, we believe System One Models can be made more reliable than its alternatives. Skip to main contentSave this storySave this storyA set of parents and their children in Illinois and California filed a lawsuit last week in federal court in Chicago alleging that Meta illegally used their Facebook and Instagram photos to build NameTag, an unreleased face-recognition system for its smart glasses, and to train generative AI models including Emu and Muse Image.

At OpenAI, I helped build the methods that made language models useful at following instructions and talking with people. At the time, I thought maybe chat models would lead to AGI, but despite the hype it became obvious to me that there was something really big missing.

The wider picture

Jev achieves similar levels of intelligence on System One tasks compared to existing LLMs, while being two orders of magnitude faster and more efficient. Extraordinary claims require extraordinary evidence so see below for the receipts. 💅Frontiers, Old and NewExisting LLMsSystem One + JevOptimized withReinforcement Learning with Human Feedback (RLHF) / Reinforcement Learning with Verifiable Rewards (RLVR)Reinforcement Learning for Calibrated Decisions (RLCD)Optimizes forHuman preference: writeups and chat responses that human raters prefer. Calibrated decisions: answers with epistemically honest probabilities on System One tasks.

SpeedEnd-to-end response time is 3 to 329 seconds for frontier models. This can range from 40x-200x faster for the same levels of frontier intelligence for System One shaped queries. ConfidenceEven if prompted for a confidence estimate, models tend to be overconfident and inconsistent. The surrounding code constrains their freedom, making them easier to compose into reliable systems. Side-by-side demonstrationOur side-by-side demo shows a key difference between our models and LLMs: Jev outputs all probabilities in parallel instead of autoregressively generating by token.

What has been reported

Instead, we assume there is a correct compute graph (a “workflow” represented in code) and use the predictions of the largest, smartest, and most expensive external models as reference probabilities. We test how they compare to the average of the smartest models (in this case, Astra and Fable). We also compare to models with a generated prompt doing all the logic in their chain-of-thought, but this tends to do significantly worse than using the workflow itself.

1 as the reference answer, which biases answers towards OpenAI and Anthropic’s models. We likely underestimate the relative performance of our model and DeepSeek’s models. The LLMs use our System One LLM wrapper, which constrains LLMs to output structured decisions compatible with our API. Having a hallucinated tool call is inconvenient in an agent, but is an absolute deal-breaker if it’s part of a system with latency guarantees or it’s buried several layers deep in a dependency chain.

Existing models, no matter how smart, still hallucinate and have type errors.

What happens next

NuanceThe numbers for LLMs are from OpenRouter i. e. , there almost certainly is bias here: more complex queries might be routed to better models. That’s because this is against the non-reasoning modes of the models (except Astra which was set to the lowest reasoning setting). For the higher cardinality choices, we do a 2 stage-system of scoring independently then making an explicit choice, hence the occassional slowdown. The model class name draws on the distinction between fast, intuitive System 1 thinking and slow, deliberate System 2 reasoning.

“System 1 thinking” has also implied error-prone. There’s generally no need to buy a fresh phone just because new models have been released, as hardware updates are broadly iterative, adding small bits to an already accomplished package. Whether your priority is the longest battery life, the best camera, the biggest screen or simply the optimal balance of features and price, there’s more to choose from in the Apple ecosystem than you may expect.

The report has been compiled by The Daily Waves using information reported across typesafe.ai, theguardian.com, wired.com. Details are presented according to the information available at the time of publication and may change as authorities, organisers or other relevant parties provide updates.

Was this article useful?