Write your own policy
The built-in policies cover common risks. For a rule that is specific to your app, write your own. There are three ways to do it:
| You want to | Use | Runs |
|---|---|---|
| Ask a question about the text | judge() | On the judge, a decision model such as Jev |
| Match a fixed pattern or list of words | regex(), substrings() | On your machine |
| Run your own code on the text | A class that implements Transform | On your machine |
The recipes in jes.recipes are built in these ways too.
sentiment is a judge() question, and competitors is a substrings() list.
Ask your own question
judge() turns a question into a policy. Give it a name, a question, and a
threshold=. The threshold is always required, because jes has no measured
default for a question you wrote.
from jes import Guard
from jes.policies import judge
from jes.questions import YesNo
refund = judge(
"refund_request",
YesNo("The text asks for money back."),
threshold=0.8,
stages=("input",),
)
guard = Guard([refund], model="jev-latest")
result = guard.check_input("Please refund my order.")
A question can take one of three shapes:
from jes.policies import judge
from jes.questions import Choice, Score, YesNo
# Yes or no. Blocks when P(yes) >= threshold.
refund = judge("refund", YesNo("The text asks for money back."), threshold=0.8)
# Pick one option. Blocks when the summed probability of the
# violating options >= threshold.
route = judge(
"route",
Choice(
"Which team should handle this?",
{"billing": "Money problems", "other": "Anything else"},
),
threshold=0.8,
violating=["billing"],
)
# A graded level. Blocks when the summed probability of level 2
# ("high") and above >= threshold.
severity = judge(
"severity",
Score("How severe is this?", ("low", "mid", "high")),
threshold=0.8,
violation_level=2,
)
Write the question about the text, not about the user. The question never contains the checked text, so attacker text never enters the instructions. See Questions and thresholds for every argument, and Thresholds for picking a number.
Shape the check
judge() exposes the options the core judgments are built with but don't let
you change.
Stages. Pick where the policy runs. The default is
("input", "output").
judge("refund", YesNo("The text asks for money back."), threshold=0.8, stages=("input",))
Context. Let the judge see the prompt with the reply.
context="optional" sends the prompt, question, or history when the check
has them. context="required" runs on output only and always sends the
prompt you pass to check_output. Add sources=True to also send the sources you pass to
check_output.
on_topic = judge(
"on_topic",
YesNo("The reply does not answer the user's question."),
threshold=0.8,
stages=("output",),
context="required",
)
Whole text. jes splits long text into chunks and asks about each one. For a
question about the text as a whole, set whole_text=True. Text that doesn't
fit then follows on_overflow, which blocks by default.
Items. Judge each part of the text on its own. Pass an items= function
that returns Items,
and cap them with max_items. More items than that follows on_overflow
too, with the finding too_many_items.
import re
from jes import Span
from jes.policies import Item, judge
from jes.questions import YesNo
def urls(text):
for match in re.finditer(r"https?://\S+", text):
yield Item(match.group(), Span(match.start(), match.end()))
phishing = judge(
"phishing_url",
YesNo("The URL looks like phishing."),
threshold=0.8,
items=urls,
max_items=20,
)
The full argument list is in the judge reference.
Match a pattern
When a pattern decides the answer, you don't need the judge. regex() and
substrings() run on your machine, cost nothing, and need no threshold.
from jes import Guard
from jes.policies import regex, substrings
guard = Guard([
regex([r"(?i)\bdrop\s+table\b"], name="sql_drop"),
substrings(["Acme", "Globex"], action="redact", name="competitors"),
])
regex()blocks or redacts matches. It uses a regex engine with a timeout (timeout_ms=50by default).match="fullmatch"withrequire=Truechecks that the whole text matches a format.substrings()matches exact words and their lookalikes, soAcmewritten with Cyrillic letters still matches.whole_words=Trueskips matches inside longer words.
Give each one a name=. Policy names must be unique within a guard, so two
regex() policies need different names.
Write a policy class
For logic that a pattern can't express, write a class. A transform runs on your machine and returns the edits to make plus what it found. jes applies the edits. It needs these attributes and one method:
import re
from dataclasses import dataclass
from jes import Guard, Span
from jes.policies import Edit, TransformFinding, TransformOutcome
ORDER_ID = re.compile(r"ORD-\d{6}")
@dataclass(frozen=True)
class OrderIds:
name: str = "order_ids"
stages: frozenset[str] = frozenset({"input", "output"})
phase: str = "detect"
def apply(self, text, context):
matches = list(ORDER_ID.finditer(text))
if not matches:
return TransformOutcome()
edits = tuple(Edit(m.start(), m.end(), "[ORDER]") for m in matches)
finding = TransformFinding(
"order_id",
"redact",
tuple(Span(m.start(), m.end()) for m in matches),
)
return TransformOutcome(edits=edits, findings=(finding,))
guard = Guard([OrderIds()])
guard.check_input("Where is ORD-123456?").onward # "Where is [ORDER]?"
The rules jes checks:
- Edits must be sorted and must not overlap. Spans index into the text the transform was given.
- A finding's label must be a valid identifier, such as
order_id. phaseis"normalize","detect", or"limit". Transforms run by phase, then in list order.contextis aTransformContextthat tells you the stage, where the text came from, and, on tool stages, the tool name.
A policy that breaks these rules fails the check with
PolicyExecutionError.
jes makes no confidentiality claim for a custom transform. What it does with
the text is up to your code.
There is no judgment class to implement. Every judgment, built-in or yours,
is built with judge().
Test it offline
FakeBackend stands in for the judge, so you can
test a judge() policy without a key. A single question is answered under
"<policy>.violation":
from jes import Guard
from jes.policies import judge
from jes.questions import YesNo, YesNoAnswer
from jes.testing import FakeBackend
refund = judge(
"refund_request",
YesNo("The text asks for money back."),
threshold=0.8,
stages=("input",),
)
backend = FakeBackend(
answers={"refund_request.violation": YesNoAnswer(0.95)},
)
result = Guard([refund], model=backend).check_input("Please refund my order.")
assert not result.ok
The score comes from you, not from a model. Use it to test your branching. To choose a threshold, use real traffic.
Next
- Your own question, a runnable example.
- Built-in policies, ready-written questions and lists.
judgeand Questions and thresholds in the API reference.