Skip to main content

Thresholds and scores

A judgment asks the judge one or more questions. The decision model answers with probabilities, and jes turns them into one violation score per question. Your threshold decides what happens next.

Threshold​

from jes import Threshold
from jes.policies import injection

injection(threshold=0.8) # same as Threshold(block_at=0.8)
injection(threshold=Threshold(block_at=0.8, flag_at=0.5))
  • A violation score at or above block_at blocks.
  • A score from flag_at up to block_at flags: the result is still allowed, with a finding.
  • Values must satisfy 0 <= flag_at <= block_at <= 1. NaN and infinities are rejected, and so are booleans, as threshold= or inside Threshold(...).

Violation score​

Question typeViolation score
YesNoThe score of "true" (every built-in yes/no question is phrased so true means violation)
ChoiceThe sum of the scores of the options marked as violating
ScoreThe sum of the scores of levels at or above the violation level

Summing matters. Text split across two violating options at 0.45 each has a violation score of 0.9, even though neither option wins alone.

Each score in result.scores records the violation score and the model that answered.

What confidence means​

Choice and Score answers from Jev also carry a confidence in [0, 1], and result.scores records it next to the value. It is not the violation score, and it is not the top option's probability. It describes the shape of the distribution: nearly all the probability on one option is confident, and probability spread across options is not.

A YesNo answer (a Noul to Jev) has no separate confidence. Its confidence is None, because the probability of yes already says how sure Jev is.

jes never uses confidence to decide. The decision compares the violation score with your threshold, and nothing else (see jes never softens a decision). Use confidence in your own code, for example to send uncertain passes to human review:

for key, score in result.scores.items():
if result.ok and score.confidence is not None and score.confidence < 0.5:
queue_for_review(key, score.value)

Choosing a threshold​

jes won't pick a number for you. Pick one from your own traffic:

  1. Pin the model first. Use a release such as model="jev-1.13.0", not the jev-latest alias, and fix the policy options. A threshold belongs to that exact setup.
  2. Collect a labelled sample from your own traffic: messages you know are fine, and messages you know should block.
  3. Score it without acting on it. Run the guard beside your app, ignore its decision, and record result.scores[...].value for every example. scores includes passing questions.
  4. Set block_at from the benign scores. The share of benign examples at or above block_at is your false-positive rate. Pick the lowest value whose rate you can live with, then check how many known attacks still score at or above it.
  5. Keep a flag band. Set flag_at below block_at so borderline text is allowed but logged. Review the flags, and move the thresholds as you learn.
  6. Re-check after every change to the Jev release, question, or transforms. The old number doesn't carry over.

Thresholds are per policy, so each judgment can have its own. topics in a customer-support bot may block at 0.6 while injection blocks at 0.9.

Why there are no defaults​

A threshold means something only for one exact setup: the Jev release, the question wording, the policy subset, the transforms, chunking, and so on.

jes publishes no thresholds, so every judgment needs threshold=. The 0.5 in the agent config is a placeholder to start from, not a measurement. A published default could only ever hold for a pinned Jev release. jev-latest is an alias, so it never matches one.

A category or label subset is a different setup. For example, hazards(["S1"]) does not inherit a threshold measured for hazards(). Topics, recipes, and judge() always require a threshold.

jes never softens a decision​

If the judge is unsure, jes still compares the score with your threshold. An attacker can make a judge unsure, so uncertainty is never a reason to allow.