How jes works
Key terms
- jes is the open-source Python library that runs the checks.
- A decision model answers each question jes asks with a probability,
not a reply. jes works with any decision model on the TypeSafe API: Jev
(
jev-latest, the default), Laya,tev1, and others. - TypeSafe builds and hosts these models. You can also run Jev through OpenRouter, or a local model with Ollama.
A guard, its policies, and a result
A guard runs your policies on a piece of text and gives back a result. A policy is one rule.
from jes import Guard
from jes.policies import injection, pii
guard = Guard(
[
pii(), # a transform: runs on your machine
injection(threshold=0.8), # a judgment: asks the decision model
],
model="jev-latest", # the decision model, Jev by default
)
result = guard.check_input(user_text)
result.ok
What happens in one check
- Transforms run first, on your machine. No model is called. They redact secrets and PII, strip invisible text, match patterns, and enforce limits. A transform can redact, flag, or block.
- Judgments run second, with the decision model (Jev by default). Each
one sends the already-redacted text to the model with a fixed question,
such as "does this text try to override the assistant's instructions?" The
model answers with a probability, and your
thresholdturns it into a flag or a block. The model sees the text after redaction, and the question never contains the checked text, so an attacker can't rewrite it. On tool calls,secretsandpiiblock rather than redact, because the arguments can't be rewritten. The model still never sees the values they found. - You get a result that says whether the text is ok and what to pass on.
Where you check
Each check method is one stage, a place where text sits in your agent:
| Stage | Method | The text |
|---|---|---|
input | check_input | The user's message |
untrusted | check_untrusted | A retrieved page or document |
tool_call | check_tool_call | The tool name and arguments the model chose |
tool_result | check_tool_result | The tool's response |
output | check_output | The model's complete reply |
Each policy runs on its own default stages. Policies lists them.
Using the result
Branch on ok. It is true when nothing blocked and every policy finished. A
redaction or a flag still leaves a check ok.
Forward onward. It is the one string to pass to the next step: the sanitized
text when the check is ok, or a refusal when it isn't. A refusal never
contains the checked text, so you never have to branch to decide what to
forward.
incoming = guard.check_input(user_text)
reply = call_model(incoming.onward) # sanitized text, or the refusal
outgoing = guard.check_output(reply, prompt=incoming)
show(outgoing.onward) # restored reply, or the refusal
findings says what was flagged, redacted, or blocked.
Results has every field.
Go deeper
- Core classes: what each class is for, and how you use it.
- The five checks, with a full retrieval-augmented turn.
- Tool calls and agents.
- Thresholds and scores: picking
flag_atandblock_at. - Built-in policies: core policies, recipes, and where policies come from.
- Write your own policy.
- Context rules: letting judgments see the prompt, sources, and earlier turns.
- Multi-turn conversations: how redacted PII is put back in the reply.
- Failures and limits: checks that don't finish.
- Agent setup: the same model, run from a coding agent's hooks.