Skip to main content

Write your own policy

The built-in policies cover common risks. For a rule that is specific to your app, write your own. There are three ways to do it:

You want toUseRuns
Ask a question about the textjudge()On the judge, a decision model such as Jev
Match a fixed pattern or list of wordsregex(), substrings()On your machine
Run your own code on the textA class that implements TransformOn your machine

The recipes in jes.recipes are built in these ways too. sentiment is a judge() question, and competitors is a substrings() list.

Ask your own question​

judge() turns a question into a policy. Give it a name, a question, and a threshold=. The threshold is always required, because jes has no measured default for a question you wrote.

from jes import Guard
from jes.policies import judge
from jes.questions import YesNo

refund = judge(
"refund_request",
YesNo("The text asks for money back."),
threshold=0.8,
stages=("input",),
)

guard = Guard([refund], model="jev-latest")
result = guard.check_input("Please refund my order.")

A question can take one of three shapes:

from jes.policies import judge
from jes.questions import Choice, Score, YesNo

# Yes or no. Blocks when P(yes) >= threshold.
refund = judge("refund", YesNo("The text asks for money back."), threshold=0.8)

# Pick one option. Blocks when the summed probability of the
# violating options >= threshold.
route = judge(
"route",
Choice(
"Which team should handle this?",
{"billing": "Money problems", "other": "Anything else"},
),
threshold=0.8,
violating=["billing"],
)

# A graded level. Blocks when the summed probability of level 2
# ("high") and above >= threshold.
severity = judge(
"severity",
Score("How severe is this?", ("low", "mid", "high")),
threshold=0.8,
violation_level=2,
)

Write the question about the text, not about the user. The question never contains the checked text, so attacker text never enters the instructions. See Questions and thresholds for every argument, and Thresholds for picking a number.

Shape the check​

judge() exposes the options the core judgments are built with but don't let you change.

Stages. Pick where the policy runs. The default is ("input", "output").

judge("refund", YesNo("The text asks for money back."), threshold=0.8, stages=("input",))

Context. Let the judge see the prompt with the reply. context="optional" sends the prompt, question, or history when the check has them. context="required" runs on output only and always sends the prompt you pass to check_output. Add sources=True to also send the sources you pass to check_output.

on_topic = judge(
"on_topic",
YesNo("The reply does not answer the user's question."),
threshold=0.8,
stages=("output",),
context="required",
)

Whole text. jes splits long text into chunks and asks about each one. For a question about the text as a whole, set whole_text=True. Text that doesn't fit then follows on_overflow, which blocks by default.

Items. Judge each part of the text on its own. Pass an items= function that returns Items, and cap them with max_items. More items than that follows on_overflow too, with the finding too_many_items.

import re

from jes import Span
from jes.policies import Item, judge
from jes.questions import YesNo

def urls(text):
for match in re.finditer(r"https?://\S+", text):
yield Item(match.group(), Span(match.start(), match.end()))

phishing = judge(
"phishing_url",
YesNo("The URL looks like phishing."),
threshold=0.8,
items=urls,
max_items=20,
)

The full argument list is in the judge reference.

Match a pattern​

When a pattern decides the answer, you don't need the judge. regex() and substrings() run on your machine, cost nothing, and need no threshold.

from jes import Guard
from jes.policies import regex, substrings

guard = Guard([
regex([r"(?i)\bdrop\s+table\b"], name="sql_drop"),
substrings(["Acme", "Globex"], action="redact", name="competitors"),
])
  • regex() blocks or redacts matches. It uses a regex engine with a timeout (timeout_ms=50 by default). match="fullmatch" with require=True checks that the whole text matches a format.
  • substrings() matches exact words and their lookalikes, so Acme written with Cyrillic letters still matches. whole_words=True skips matches inside longer words.

Give each one a name=. Policy names must be unique within a guard, so two regex() policies need different names.

Write a policy class​

For logic that a pattern can't express, write a class. A transform runs on your machine and returns the edits to make plus what it found. jes applies the edits. It needs these attributes and one method:

import re
from dataclasses import dataclass

from jes import Guard, Span
from jes.policies import Edit, TransformFinding, TransformOutcome

ORDER_ID = re.compile(r"ORD-\d{6}")

@dataclass(frozen=True)
class OrderIds:
name: str = "order_ids"
stages: frozenset[str] = frozenset({"input", "output"})
phase: str = "detect"

def apply(self, text, context):
matches = list(ORDER_ID.finditer(text))
if not matches:
return TransformOutcome()
edits = tuple(Edit(m.start(), m.end(), "[ORDER]") for m in matches)
finding = TransformFinding(
"order_id",
"redact",
tuple(Span(m.start(), m.end()) for m in matches),
)
return TransformOutcome(edits=edits, findings=(finding,))

guard = Guard([OrderIds()])
guard.check_input("Where is ORD-123456?").onward # "Where is [ORDER]?"

The rules jes checks:

  • Edits must be sorted and must not overlap. Spans index into the text the transform was given.
  • A finding's label must be a valid identifier, such as order_id.
  • phase is "normalize", "detect", or "limit". Transforms run by phase, then in list order.
  • context is a TransformContext that tells you the stage, where the text came from, and, on tool stages, the tool name.

A policy that breaks these rules fails the check with PolicyExecutionError. jes makes no confidentiality claim for a custom transform. What it does with the text is up to your code.

There is no judgment class to implement. Every judgment, built-in or yours, is built with judge().

Test it offline​

FakeBackend stands in for the judge, so you can test a judge() policy without a key. A single question is answered under "<policy>.violation":

from jes import Guard
from jes.policies import judge
from jes.questions import YesNo, YesNoAnswer
from jes.testing import FakeBackend

refund = judge(
"refund_request",
YesNo("The text asks for money back."),
threshold=0.8,
stages=("input",),
)
backend = FakeBackend(
answers={"refund_request.violation": YesNoAnswer(0.95)},
)
result = Guard([refund], model=backend).check_input("Please refund my order.")
assert not result.ok

The score comes from you, not from a model. Use it to test your branching. To choose a threshold, use real traffic.

Next​