Jev AI for Developers: A Practical Guide to TypeSafe’s Decision Model
Jev AI for Developers: A Practical Guide to TypeSafe’s Decision Model
Many AI integrations ask a language model to do something surprisingly small: classify a support ticket, select a tool, rank an item or decide which workflow should run next. The application needs a structured decision, yet the model is built to generate text.
Jev AI, developed by TypeSafe AI, approaches this problem through a different interface. Developers supply context and explicitly defined questions. The model returns typed decisions that software can consume, rather than paragraphs that need interpreting. TypeSafe describes Jev as its first public “System One” model, designed for automation inside applications.
For developers, the useful question is where this approach improves an existing system. This guide examines Jev’s decision primitives, API integration, pricing and production considerations, with practical examples of how to evaluate it.
What Is Jev AI?
Jev evaluates a supplied state against structured questions. That state might contain a customer message, document content or relevant application context. Each question defines the judgement your software needs.
Its interface exposes three primitives: Choice, Score and Noul. These can be combined in one request, with questions evaluated independently against the same state. TypeSafe recommends asking narrow, atomic questions and combining their results in application code.
Consider a support system. Instead of asking a model to “analyse this ticket and decide what to do”, a developer could ask:
- Which department should handle the request?
- Does the message indicate urgency?
- How severe is the reported disruption?
Your application then applies its own routing rules. A billing complaint might enter the payments queue, while an urgent technical incident triggers a separate escalation process.
This design makes the boundary between model judgement and business logic explicit. Developers can inspect and change the routing policy without rewriting an all-purpose prompt.
How Does Jev Work? System One Models and RLCD
TypeSafe says Jev combines a new architecture, parallel sampling and a training method called Reinforcement Learning for Calibrated Decisions, or RLCD. Its stated objective is to return decisions with probabilities that meaningfully represent uncertainty.
Calibration matters because a decision system needs more than a predicted label. It also needs a way to determine when that prediction deserves further checking.
For example, if predictions assigned approximately 80% probability are correct approximately 80% of the time across a representative dataset, those predictions are well calibrated at that level. Calibration is a property measured across examples; it cannot establish whether one individual answer is correct.
Developers should therefore evaluate probability quality alongside classification accuracy. A useful model might correctly identify most tickets but still express excessive certainty on ambiguous requests.
The practical consequence is straightforward: measure uncertainty on your own data before using it to control automation.
Jev vs LLMs: Which Should Developers Use?
The distinction is primarily about the task and output your application requires.
- Select one category from a defined list: evaluate Jev Choice.
- Rate content against an explicit rubric: evaluate Jev Score.
- Evaluate a narrowly defined proposition: evaluate Jev Noul.
- Generate explanations, articles or source code: use a generative language model.
- Perform complex planning and produce a detailed response: evaluate a reasoning-capable language model.
- Apply an exact, fully specified business rule: use conventional application code.
Jev’s decision interface gives up free-form text generation. That makes it relevant to classification and branching, while a generative model remains appropriate when the output itself must be prose or code.
A sensible architecture can use both. A decision stage evaluates the request, application code chooses the next operation, and a language model generates a response where necessary.
Before introducing either, check whether ordinary rules already solve the problem. If a ticket contains a trusted department identifier, routing it through AI adds cost and another possible failure point.
Understanding Choice, Score and Noul
Choice: Select a Defined Outcome
Choice selects an option from a developer-defined set. Its response includes the selected option, probabilities across the available options and a confidence value.
For ticket routing, options might include billing, technical, sales and other. Descriptions should clearly distinguish the categories. TypeSafe’s documentation recommends including an alternative such as other when the list might not cover every input.
Avoid overlapping definitions. If both “technical” and “integrations” include connection failures, the ambiguity begins in your schema.
Score: Evaluate Against a Rubric
Score evaluates the state against ordered criteria. It returns a score, probabilities over the rubric levels and confidence.
An incident rubric could distinguish a minor inconvenience, reduced functionality and a complete service interruption. Those descriptions are more actionable than asking for an unexplained severity score out of ten.
Noul: Evaluate a Proposition
Noul evaluates a proposition and returns a value between zero and one. Unlike Choice and Score, it does not include a separate confidence property.
A suitable question might ask whether a message explicitly requests cancellation. Decide what threshold should trigger your workflow through evaluation, rather than assuming any universal cut-off.
How to Use the Jev API
TypeSafe documents an HTTP endpoint at https://api.typesafe.ai/v1/systemone, authenticated with a bearer API key. Requests include state, model and questions. The quick start uses jev-latest as the model identifier.
To build a basic support-routing integration, obtain an API key and store it securely on your server. Send a POST request to the endpoint with an Authorization header containing your bearer token and a Content-Type header set to application/json.
For example, your request could use the following configuration:
- Model: jev-latest.
- State: “I was charged twice for my subscription.”
- Question identifier: department.
- Question type: choice.
- Instructions: “Which department should handle this request?”
- Criteria: billing for charges, invoices and subscription payments; technical for software faults and integration failures; sales for pre-purchase enquiries and product selection; and other for requests outside these categories.
This configuration defines both the information being evaluated and the possible outcomes. Your application remains responsible for deciding what happens after the classification.
The response’s answers object contains the result under your question identifier. For Choice, inspect choice, probabilities and confidence.
Keep credentials on the server. Wrap the integration in a small application service so request construction, timeouts, logging and fallback behaviour remain consistent.
Treat this example as a starting configuration, rather than a complete production integration. Add error handling and verify current API requirements in the official documentation.
Jev AI Pricing and Performance
TypeSafe’s launch pricing lists $0.042 per million input tokens, with output tokens free. Its launch announcement reports end-to-end response times of 70–500 milliseconds. These are published vendor figures, rather than performance guarantees for every application.
At that input rate, 100 million billable input tokens would cost $4.20. This calculation excludes any additional infrastructure, integration or operational costs.
TypeSafe also advertises substantial speed and cost improvements in its workflow evaluations. Developers should benchmark their own payloads rather than use a headline multiplier as a capacity-planning assumption.
Record median and tail latency, request failures, throughput under concurrency and total cost per completed workflow. Include retries and fallback calls: a cheap first decision may become expensive if it frequently requires another model.
For production applications, run measurements from the infrastructure region that will actually serve users.
Practical Jev Use Cases for Developers
The following are application patterns worth testing, rather than guarantees of suitability.
Support Ticket Classification
Evaluate department, urgency and disruption separately. Combine the results with account information and service-level rules in code.
Start by suggesting routes to support staff. Their corrections provide useful evidence before automatic assignment.
LLM Routing
Use a decision stage to distinguish requests that require text generation, specialist reasoning or a deterministic operation.
Assess routing by the quality and total cost of the final answer. A cheaper model selection is useful only if the resulting response still meets your requirements.
Document Triage
Classify incoming documents into known categories, then pass them to the appropriate processing pipeline.
Include an unknown category and test mixed documents. A file containing both an invoice and a complaint should not silently disappear into the wrong queue.
Agent Action Selection
Present a constrained set of actions appropriate to the current application state.
Keep permission checks in conventional code. A model selecting “delete record” should never bypass authorisation, confirmation or transaction controls.
Does Jev Really Have Zero Hallucinations?
TypeSafe connects its “zero hallucinations” claim to guaranteed schema matching: outputs stay within the defined structure. That is a narrower claim than guaranteeing that every decision is correct.
A valid billing label can still be the wrong classification.
Likewise, TypeSafe explains that its confidence value is derived from the probability distribution. Developers should not interpret a confidence value of 0.9 as an automatically established 90% correctness rate for their application.
Test ambiguous inputs, missing context and unfamiliar categories. Measure the errors that matter operationally, including confidently incorrect predictions.
How to Evaluate Jev Before Production
Build a labelled dataset containing common cases, edge cases and examples your current system handles poorly. Reserve a separate test set so repeated prompt adjustments do not overfit the evaluation.
Compare Jev with your existing implementation and a simple baseline. Measure per-category precision and recall, calibration, latency and cost.
Choose automation thresholds from observed results. Routing a recoverable support ticket can tolerate a different error rate from executing an account change.
Deploy initially in shadow mode, recording proposed decisions without changing live behaviour. Then automate a limited subset and monitor corrections, escalations and distribution changes.
Is Jev Worth Using?
Jev is worth evaluating when software needs frequent, narrowly defined judgements over contextual information. Its appeal is the possibility of placing those judgements behind an explicit interface that application code controls.
Begin with one measurable workflow. Define the options carefully, test against real examples and retain a reliable fallback.
The strongest adoption case is an improvement you can demonstrate: faster processing, lower total cost or better decisions at an acceptable error rate.











