Skip to main content

Policies

from jes.policies import injection, pii, judge # …

jes.policies holds the core policies, the judge() and regex() / substrings() building blocks, and the protocols for writing your own. See Where policies come from.

A policy is a transform (a local rule that can edit, flag, or block), a sensitive-data policy (secrets, pii, canary: it finds values and jes replaces them), or a judgment (questions for a decision model, compared with a threshold). Every judgment requires threshold=. See Thresholds and scores.

Transforms​

FactoryDefault stagesPhaseExtra
invisible_textinput, untrusted, tool_result, outputnormalize—
secretsall fivedetectsecrets
piiall fivedetectpii
canaryoutput, tool_calldetect—
regexinput, untrusted, tool_result, outputdetectregex
substringsinput, untrusted, tool_result, outputdetect—
allowed_toolstool_calldetect—
token_limitinputlimittokens

Transforms that edit text skip tool_call by default, because rewriting the argument string would break the JSON your application executes. secrets, pii, and canary do run there, but they block instead of redacting (pii can flag instead). A redaction on tool_call from any transform is turned into a block.

secrets, pii, and canary only report where values are. jes replaces every standalone occurrence of each found value, not just the one reported, so a repeated value never reaches a judge. On tool_call, judges see the arguments with the values replaced, while the arguments you run stay unchanged.

invisible_text​

invisible_text(
mode: Literal["targeted", "all"] = "targeted",
*,
block: bool = False,
name: str = "invisible_text",
stages: Iterable[Stage] = ("input", "untrusted", "tool_result", "output"),
)

Removes invisible and control characters used to smuggle text. "targeted" removes default-ignorable code points, bidi controls, and C0/C1 controls (except tab, LF, and CR). It keeps ZWJ/ZWNJ and valid variation sequences, so emoji and Persian or Indic text survive. "all" also removes joiners and every remaining format (Cf), private-use (Co), and unassigned (Cn) character. The finding is redact when anything was removed. With block=True it blocks instead.

secrets​

secrets(
redact: Literal["all", "partial", "hmac"] = "all",
*,
key: bytes | None = None,
name: str = "secrets",
stages: Iterable[Stage] = ("input", "untrusted", "tool_call", "tool_result", "output"),
)

Redacts secrets found by detect-secrets, plus known key formats: OpenAI (including project keys) and Anthropic sk- keys, GitHub ghp_, gho_, ghu_, ghs_, ghr_, and fine-grained tokens, Slack tokens, AWS AKIA and ASIA key ids, Google API keys, Stripe keys, and private key blocks. Public IP addresses are not reported. Requires the secrets extra.

"all" replaces the secret with ******. "partial" keeps the first and last two characters of a value longer than 8 characters, and masks a shorter one fully. "hmac" writes HMAC-SHA256 hex under key, so the same secret matches across conversations. That mode requires key, of at least 32 bytes. On tool_call, a secret blocks and the arguments are left unchanged. The finding label is secret.

pii​

pii(
entities: Iterable[str] | None = None,
*,
input_mode: Literal["redact", "mask", "block"] = "redact",
untrusted_mode: Literal["mask", "redact", "block"] = "mask",
output_mode: Literal["flag", "redact", "block"] = "flag",
tool_call_mode: Literal["block", "flag"] = "block",
restore: bool = True,
restore_origins: Iterable[str] = (),
on_placeholder_in_url: Literal["block", "allow"] = "block",
name: str = "pii",
stages: Iterable[Stage] = ("input", "untrusted", "tool_call", "tool_result", "output"),
)

Detects personal data with built-in patterns. Credit cards are Luhn-checked, and phone numbers need separators. Each finding's label is the entity, such as EMAIL_ADDRESS.

  • Entities. DEFAULT_PII_ENTITIES is used when entities is None: EMAIL_ADDRESS, PHONE_NUMBER, CREDIT_CARD, US_SSN, IBAN_CODE, CRYPTO, and PERSON. UUID, IP_ADDRESS, and US_BANK_NUMBER are opt-in. PII_ENTITIES lists all ten. An unknown entity raises PolicyError.
  • PERSON. Uses Presidio with an installed spaCy English model (en_core_web_lg, en_core_web_md, or en_core_web_sm). Without the pii extra or a model, pii() raises PolicyError. Pass entities without PERSON to run on patterns only, with no extra.
  • Input. With restore=True, redact issues reversible conversation tokens, [JES_PII_…]. With restore=False, it writes irreversible [REDACTED_<ENTITY>] markers. mask keeps the first and last two characters, and block blocks.
  • Untrusted and tool results. Values are masked by default (untrusted_mode). This text never gets tokens.
  • Tool calls. A finding blocks, and the arguments are left unchanged. With tool_call_mode="flag", the call passes with a flag finding. Either way, judges don't see the values.
  • Output. New personal data in a reply is always hidden from judges. flag puts it back in onward afterward, redact leaves [REDACTED_<ENTITY>] in place, and block blocks.
  • restore_origins. Absolute http/https origins where a token inside a URL may be restored. Any other URL gets placeholder_in_url, which blocks unless on_placeholder_in_url="allow".
  • Several pii policies. Restoration is off when any of them sets restore=False. Their restore_origins are combined, and a token in a URL is allowed only when every one sets on_placeholder_in_url="allow".

See Multi-turn conversations.

canary​

canary(
token: str,
*,
name: str = "canary",
stages: Iterable[Stage] = ("output", "tool_call"),
)

Irreversibly removes and blocks a leaked application canary token, such as a marker you put in your system prompt. It matches lookalike, cased, and invisible-character spellings too. The finding label is canary.

regex​

regex(
patterns: Iterable[str],
*,
action: Literal["block", "redact"] = "block",
match: Literal["search", "fullmatch"] = "search",
require: bool = False,
fold: bool = False,
timeout_ms: int = 50,
name: str = "regex",
stages: Iterable[Stage] = ("input", "untrusted", "tool_result", "output"),
)

Matches your patterns with the regex engine and a per-pattern timeout. A timeout blocks. match="fullmatch" requires the pattern to match the whole text. With require=True, text that matches none of the patterns is blocked, and text that matches passes unchanged (action is not applied). action="redact" replaces each match with [REDACTED]. fold matches on a folded copy (see substrings).

substrings​

substrings(
terms: Iterable[str],
*,
action: Literal["block", "redact"] = "block",
whole_words: bool = False,
fold: bool = True,
name: str = "substrings",
stages: Iterable[Stage] = ("input", "untrusted", "tool_result", "output"),
)

Bans or redacts exact substrings and lookalike spellings. Folding applies NFKC normalization, case folding, and removal of default-ignorable characters. It also applies the Unicode TR39 confusables skeleton, so lookalike letters and hidden characters cannot split a banned term. action="redact" replaces each match in the original text with [REDACTED].

EXPLOIT_TERMS​

from jes.policies import EXPLOIT_TERMS, substrings

substrings(EXPLOIT_TERMS)

An opt-in tuple of curated exploit phrases, for example "arbitrary code execution" and "a backdoor". It is not applied by default.

allowed_tools​

allowed_tools(
names: Iterable[str],
*,
name: str = "allowed_tools",
)

Blocks a tool call whose name is not in names, with the finding label tool_name. It runs on tool_call only and never edits the arguments. names must be non-empty, single-line strings. Leaving this policy out does not authorize any tool: jes checks text, and your application decides what may run. See Tool calls and agents.

token_limit​

token_limit(
limit: int,
*,
encoding: str = "cl100k_base",
mode: Literal["block", "truncate"] = "block",
name: str = "token_limit",
stages: Iterable[Stage] = ("input",),
)

Counts tiktoken tokens. "truncate" cuts to the limit and flags. Truncation runs after redaction.

Judgments​

FactoryDefault stagesContextQuestion ids
injectioninput, untrusted, tool_callnoneviolation
indirect_injectionuntrusted, tool_resultoptional (the question)violation
hazardsinput, outputoptional (prompt, history)S1 … S14
topicsinput, outputnoneone per topic, such as medication_dosage
toxicityinput, outputnoneone per label, such as insult
tool_safetytool_calloptional (prompt, tool name)violation
judgeinput, outputconfigurableyours

Every built-in judgment is one judge() call with frozen question text. A finding's label is the question id, and result.scores is keyed "<policy>.<question id>", such as "injection.violation" or "topics.medication_dosage". The question text for an id never changes; new wording would get a new id.

injection​

injection(
*,
threshold: Threshold | float,
model: ModelSpec | None = None,
) -> Judgment

One yes/no question: the text tries to override, ignore, or reveal an AI assistant's instructions, or to make it abandon its rules.

indirect_injection​

indirect_injection(
*,
threshold: Threshold | float,
model: ModelSpec | None = None,
) -> Judgment

One yes/no question: the text contains instructions addressed to an AI assistant rather than information for a reader.

hazards​

hazards(
categories: Iterable[str] | None = None,
*,
threshold: Threshold | float,
model: ModelSpec | None = None,
) -> Judgment

One yes/no question per hazard category, "S1" through "S14", or only the categories you pass. The threshold applies to each category, and each finding's label is the category code. There is no "any hazard" question. An unknown or empty category list raises PolicyError.

CodeHazardCodeHazard
S1Violent crimesS8Intellectual property
S2Non-violent crimesS9Indiscriminate weapons
S3Sex-related crimesS10Hate
S4Child sexual exploitationS11Suicide and self-harm
S5DefamationS12Sexual content
S6Specialized adviceS13Elections
S7PrivacyS14Code interpreter abuse

topics​

topics(
deny: Iterable[str],
*,
threshold: Threshold | float,
model: ModelSpec | None = None,
) -> Judgment

One yes/no question per denied topic. Text can touch several topics, so each topic is its own question, not a choice. Each question id is a slug of the topic, so "medication dosage" becomes medication_dosage. A slug that repeats gets a suffix, such as _2.

toxicity​

toxicity(
labels: Iterable[str] | None = None,
*,
threshold: Threshold | float,
model: ModelSpec | None = None,
) -> Judgment

One yes/no question per label, or only the labels you pass: toxicity, severe_toxicity, obscene, threat, insult, identity_attack, and sexual_explicit.

tool_safety​

tool_safety(
*,
threshold: Threshold | float,
model: ModelSpec | None = None,
) -> Judgment

One yes/no question, judged with the user's request as context: "The tool call, given the user's request, is destructive, exfiltrates data, changes privileges, or goes beyond what the user asked." Runs only on tool_call. See Tool calls.

judge​

judge(
name: str,
questions: Question | Mapping[str, Question],
*,
threshold: Threshold | float,
violating: Collection[str] | Mapping[str, Collection[str]] | None = None,
violation_level: int | Mapping[str, int] | None = None,
stages: Iterable[Stage] = ("input", "output"),
context: Literal["none", "optional", "required"] = "none",
sources: bool = False,
whole_text: bool = False,
on_overflow: Literal["block", "allow"] = "block",
items: ItemExtractor | None = None,
max_items: int | None = None,
model: ModelSpec | None = None,
) -> Judgment

Creates a custom judgment from questions. The guide is Write your own policy.

ArgumentMeaning
questionsOne question, whose id is "violation", or a mapping of ids to questions. Ids must be valid identifiers.
thresholdAlways required.
violatingFor a Choice: the option labels that count as violations. Required for choice questions, and a mapping when there are several.
violation_levelFor a Score: the index of the lowest violating level. Required for score questions, and a mapping when there are several.
context"optional" sends the prompt (or, on untrusted, the question) and the newest history that fit. A prompt that doesn't fit is dropped with a context_dropped flag, and cut history gets history_truncated. "required" treats a prompt that doesn't fit as an overflow (context_too_long), and is output-only.
sourcesAlso send check_output's sources.
whole_textNever chunk the text. It must fit in one request.
on_overflowWhat happens when the text, required context, or items don't fit: "block" blocks, "allow" flags. Either way the result is incomplete.
itemsAn ItemExtractor that splits the text into Items, each judged on its own. Incompatible with whole_text.
max_itemsThe per-policy item cap. Past it, the finding is too_many_items. Requires items.
modelA decision model for this judgment instead of the guard's.

Invalid combinations raise PolicyError.

from jes.policies import judge
from jes.questions import YesNo

refund = judge(
"refund_request",
YesNo("The text asks for money back."),
threshold=0.8,
stages=("input",),
)

Protocols for custom policies​

Implement Transform to write your own local policy. Custom transforms may edit, flag, redact, and block, but jes makes no confidentiality claim about them. Judgments are built only with judge(). See Write a policy class.

NameKindSummary
Policytype aliasTransform | Judgment | SensitivePolicy
Transformprotocolname, stages: frozenset[Stage], phase, and apply(text, context: TransformContext) -> TransformOutcome
TransformContextdataclassstage (the check's stage), origin (where this text came from), target, tool (the tool name on tool stages, else None), deadline, and check_deadline()
TransformOutcomedataclassedits: tuple[Edit, ...] = (), findings: tuple[TransformFinding, ...] = (). The engine applies the edits.
TransformFindingdataclasslabel, action, spans (indices into the text the transform was given)
Editnamed tuplestart, end, replacement. Edits are sorted and don't overlap. An empty range inserts, and an empty replacement deletes.
Judgmentfrozen dataclassWhat judge() returns. Don't build one yourself.
Itemdataclasstext, span, for item judgments
ItemExtractortype aliasCallable[[str], Iterable[Item]]
SensitivePolicyabstract classThe base of secrets, pii, and canary: detect(text, context) -> list[SensitiveHit]. It is jes-owned.
SensitiveHitdataclassspan, entity, mode, action: a found value and how the engine replaces it
Phaseliteral"normalize", "detect", or "limit"
ContextModeliteral"none", "optional", or "required"
OverflowModeliteral"block" or "allow"

A finding label must be a valid identifier. A custom policy that raises, returns the wrong type, returns invalid edits or spans outside the text, or uses an invalid label or unknown action fails with PolicyExecutionError.