Skip to main content

Questions and thresholds

A question is what a judgment asks the decision model about the text. A threshold is where its answer turns into a flag or a block.

Its job​

The two split one decision in half:

  • The question says what to measure. It is typed: yes or no, one of several options, or a level on a scale. The decision model answers it with probabilities, not with text.
  • The threshold says how sure is sure enough. jes turns the answer into a violation score from 0 to 1, and the threshold decides what that score means for your app.

The model never decides whether to block. It only measures. You decide where the line is.

Mental model​

How each type becomes a violation score:

TypeYou askViolation score
YesNoA yes/no question. "Yes" means a violation.P(yes)
ChoicePick one option. You name the violating ones.The summed probability of the violating options
ScoreA level on an ordered scale of 2 to 10 levels. You name the violation_level.The probability of that level or higher

Using it​

Built-in judgments, such as injection or toxicity, bring their own questions. You only pass the threshold:

from jes import Threshold
from jes.policies import injection, toxicity

injection(threshold=0.8) # block at 0.8
toxicity(threshold=Threshold(block_at=0.9, flag_at=0.6)) # flag first, block later

To ask your own question, wrap it in judge():

from jes import Choice, Score, Threshold, YesNo
from jes.policies import judge

refund = judge(
"refund_promise",
YesNo("Does the reply promise a refund?"),
threshold=0.7,
stages=("output",),
)

tone = judge(
"tone",
Choice(
"What is the tone of the reply?",
{"friendly": None, "neutral": None, "hostile": "Rude or threatening."},
),
violating={"hostile"},
threshold=Threshold(block_at=0.8, flag_at=0.5),
stages=("output",),
)

severity = judge(
"severity",
Score("How severe is the harm?", ("none", "low", "high")),
violation_level=2, # "high"
threshold=0.5,
stages=("output",),
)

After a check, result.scores holds every question's violation score, keyed "<policy>.<question id>", such as "tone.violation".

Good to know​

  • A threshold is always required. jes ships no default, because a good threshold depends on the model, the question, and your traffic. Measure it on your own data. See Thresholds and scores.
  • A threshold belongs to one model and one question. If you change either, measure again. Pin a model release, such as jev-1.13.0, when you do.
  • A number is shorthand. threshold=0.8 means Threshold(block_at=0.8). Add flag_at to log borderline text without blocking it. A boolean raises PolicyError.
  • One question gets the id "violation". Pass a mapping of ids to questions to ask several at once, such as {"medical": YesNo(...), "legal": YesNo(...)}. The threshold applies to each.
  • Built-in question text is frozen. It never changes under the same id, so a threshold you measured keeps its meaning across jes releases.

Next​