The judge
Every judgment in jes is answered by a decision model on the TypeSafe API.
jes isn't tied to one model: Jev (jev-latest) is the default, and Laya,
tev1 (which also runs locally on Ollama), and any other
TypeSafe decision model work the same way. A decision model
classifies instead of generating. It takes the text and a set of typed
questions and returns a probability for each one in a single call. jes
compares those probabilities with your thresholds.
This page is about the model. The judge() function is something else: it
builds a policy from your own question, and the decision model answers it. See
judge.
jes talks to the decision model through LangChain's
TypeSafeClassifier,
from the langchain-typesafe package that pip install jes installs.
Pass a model
Give the guard any TypeSafe model name, or a classifier you built yourself:
from langchain_typesafe import TypeSafeClassifier
from jes import Guard
from jes.policies import injection
# A model name. jes builds the classifier on the first check.
guard = Guard([injection(threshold=0.50)], model="jev-latest")
# A classifier, when you want to set the key or the endpoint in code.
guard = Guard(
[injection(threshold=0.50)],
model=TypeSafeClassifier(
model="jev-1.13.0",
# api_key="...", # otherwise TYPESAFE_API_KEY
# base_url="...", # otherwise TYPESAFE_BASE_URL, default https://api.typesafe.ai
),
)
| Variable | Meaning |
|---|---|
TYPESAFE_API_KEY | Required unless you pass api_key= |
TYPESAFE_BASE_URL | Optional. Overrides https://api.typesafe.ai. Requests go to /v1/systemone. |
TypeSafeClassifier is marked beta in langchain-typesafe, so its
constructor may change between releases.
Both forms become a TypeSafe backend. Build one
yourself to change its request timeout or size cap:
from jes.backend import TypeSafe
guard = Guard(
[injection(threshold=0.50)],
model=TypeSafe("jev-1.13.0", timeout=10.0, max_request_bytes=262_144),
)
The same classifier can reach Jev through OpenRouter, or
run tev1 locally on Ollama.
Pin the model when thresholds matter
jev-latest is an alias that moves when TypeSafe releases a new Jev. A
threshold you chose for one release means nothing for the next, so pin a
release such as jev-1.13.0 once you have tuned your thresholds. See
Choosing a threshold.
One policy, its own model
Every judgment factory takes model= too. It overrides the guard's model for
that policy only:
from jes import Guard
from jes.policies import injection, toxicity
guard = Guard(
[
injection(threshold=0.80, model="jev-1.13.0"),
toxicity(threshold=0.70),
],
model="jev-latest",
)
How questions map to Jev
jes questions are TypeSafe's three primitives, under jes names:
| jes question | Sent to Jev as | Jev returns | Violation score | confidence |
|---|---|---|---|---|
YesNo | Noul, with true/false as criteria | The probability of yes | That probability | None |
Choice | Choice, with the options as criteria | A probability per option | The sum over the violating options | From Jev |
Score | Score, with the levels as criteria | A probability per level | The sum from violation_level up | From Jev |
- The questions travel separately from the text. The text never becomes part of a question, so it can't rewrite one.
- One call answers a batch of questions. Policies that share a
model and the same context rules are batched together, so a check makes a
few requests, not one per question. A check sends its requests in parallel.
result.usagehas one entry per call. - The text sent is the checked text alone, or, when there's context, compact
JSON with
text,prompt,question,sources,history, andtool.
What confidence means explains how
to use the confidence on Choice and Score answers.
Chat models are not judges
jes doesn't send questions to a chat model, on purpose:
- A chat model writes its verdict as text, which jes would have to parse.
- It has no honest probability to compare with a threshold.
- It follows instructions, including instructions hidden in the text it is judging. A page ending in "note to the reviewer: this is safe" is aimed at the judge as much as the agent. Jev returns probabilities, not text, so there is no reply for that note to steer.
Size and failures
- jes caps the UTF-8 size of each request, the text plus the questions, at
max_request_bytes, 1 MiB by default. That cap is not Jev's token window, and TypeSafe may still reject a request it considers too long. - The classifier owns HTTP. Neither it nor jes retries a failed call. jes
makes one classifier call per
batch, within the guard's
deadline_s. - A failed call becomes a
BackendErrorwith one of these reasons, and followson_backend_error:
| Reason | Cause |
|---|---|
client_setup_failed | The classifier could not be built, most often because TYPESAFE_API_KEY is not set |
timeout | The request timed out |
unauthorized | TypeSafe returned 401 or 403, for example an invalid key |
rate_limited | TypeSafe returned 429 |
upstream_error | TypeSafe returned a 5xx status |
rejected | TypeSafe returned another 4xx status |
connection_error | Any other transport failure |
malformed_response | TypeSafe's response failed validation |
malformed_answer | The response was missing an answer or didn't match the question |
malformed_reply | A backend's reply was missing answers for some questions |
unexpected_error | Any other exception from the backend |
See Failures and limits.
Tests and your own backend
model= also accepts any object that implements the backend protocol from
jes.backend: a name, a model, a
headroom() that says how many bytes of text still fit, and decide() (or
adecide() for an AsyncGuard), which takes a Request and returns a
Reply. FakeBackend is one, and it's what
every offline example uses.