Thresholds and scores
A judgment asks the judge one or more questions. The decision model answers with probabilities, and jes turns them into one violation score per question. Your threshold decides what happens next.
Threshold
from jes import Threshold
from jes.policies import injection
injection(threshold=0.8) # same as Threshold(block_at=0.8)
injection(threshold=Threshold(block_at=0.8, flag_at=0.5))
- A violation score at or above
block_atblocks. - A score from
flag_atup toblock_atflags: the result is still allowed, with a finding. - Values must satisfy
0 <= flag_at <= block_at <= 1. NaN and infinities are rejected, and so are booleans, asthreshold=or insideThreshold(...).
Violation score
| Question type | Violation score |
|---|---|
YesNo | The score of "true" (every built-in yes/no question is phrased so true means violation) |
Choice | The sum of the scores of the options marked as violating |
Score | The sum of the scores of levels at or above the violation level |
Summing matters. Text split across two violating options at 0.45 each has a violation score of 0.9, even though neither option wins alone.
Each score in result.scores records the violation score and the model that
answered.
What confidence means
Choice and Score answers from Jev also carry a confidence in [0, 1],
and result.scores records it next to the value. It is not the violation
score, and it is not the top option's probability. It describes the shape of
the distribution: nearly all the probability on one option is confident, and
probability spread across options is not.
A YesNo answer (a Noul to Jev) has no separate confidence. Its
confidence is None, because the probability of yes already says how sure
Jev is.
jes never uses confidence to decide. The decision compares the violation
score with your threshold, and nothing else (see
jes never softens a decision). Use
confidence in your own code, for example to send uncertain passes to human
review:
for key, score in result.scores.items():
if result.ok and score.confidence is not None and score.confidence < 0.5:
queue_for_review(key, score.value)
Choosing a threshold
jes won't pick a number for you. Pick one from your own traffic:
- Pin the model first. Use a release such as
model="jev-1.13.0", not thejev-latestalias, and fix the policy options. A threshold belongs to that exact setup. - Collect a labelled sample from your own traffic: messages you know are fine, and messages you know should block.
- Score it without acting on it. Run the guard beside your app, ignore
its decision, and record
result.scores[...].valuefor every example.scoresincludes passing questions. - Set
block_atfrom the benign scores. The share of benign examples at or aboveblock_atis your false-positive rate. Pick the lowest value whose rate you can live with, then check how many known attacks still score at or above it. - Keep a flag band. Set
flag_atbelowblock_atso borderline text is allowed but logged. Review the flags, and move the thresholds as you learn. - Re-check after every change to the Jev release, question, or transforms. The old number doesn't carry over.
Thresholds are per policy, so each judgment can have its own. topics in a
customer-support bot may block at 0.6 while injection blocks at 0.9.
Why there are no defaults
A threshold means something only for one exact setup: the Jev release, the question wording, the policy subset, the transforms, chunking, and so on.
jes publishes no thresholds, so every judgment needs threshold=. The
0.5 in the agent config is a placeholder to start
from, not a measurement. A published default could only ever hold for a
pinned Jev release. jev-latest is an alias, so it never matches one.
A category or label subset is a different setup. For example,
hazards(["S1"]) does not inherit a threshold measured for hazards().
Topics, recipes, and judge() always require a threshold.
jes never softens a decision
If the judge is unsure, jes still compares the score with your threshold. An attacker can make a judge unsure, so uncertainty is never a reason to allow.