Guard and AsyncGuard
from jes import Guard, AsyncGuard, Limits
Guard
Runs transforms in the calling thread, and sends a check's backend requests in parallel threads.
Guard(
policies: Iterable[Policy],
*,
model: ModelSpec | None = None,
limits: Limits | None = None,
deadline_s: float | None = 30.0,
on_backend_error: Literal["raise", "block", "allow"] = "raise",
fail_fast: bool = False,
)
| Argument | Meaning |
|---|---|
policies | Transforms, sensitive-data policies, and judgments. Transforms run first, by phase, then judgments. |
model | The decision model for judgments: a TypeSafe model id such as "jev-latest" or "jev-1.13.0", a TypeSafeClassifier, or your own backend such as FakeBackend. A judgment's own model= overrides it. See ModelSpec. |
limits | A Limits. None uses the defaults. |
deadline_s | Seconds one check may take. Requests still running then count as backend errors. None disables the deadline. |
on_backend_error | "raise", "block" (fail closed), or "allow" (fail open). With "block" and "allow" the result is incomplete, so ok is false either way. See Failures and limits. |
fail_fast | Skip the judgments once a transform has blocked. The result is then complete=False. |
A guard holds no per-conversation state. One guard can serve many threads.
Construction errors
The constructor raises PolicyError when:
- an entry in
policiesis not a transform, sensitive-data policy, or judgment; - two policies share a name, or a name is not a valid identifier;
- a transform has an unknown phase;
- a judgment has no model, from
model=on the guard or on the judgment; - (
Guardonly) a judgment's backend is async-only. UseAsyncGuard; - a judgment's questions leave no room for text: less than 64 bytes of headroom on one of its stages;
on_backend_erroris not"raise","block", or"allow", ordeadline_sis not a positive finite number orNone.
Limits(...) raises PolicyError itself when a field is not a positive
integer. judge() raises when a judgment with context="required" has a stage
other than output.
Closing a guard
Guard keeps a thread pool for its requests, of max_concurrency threads.
guard.close() stops it. Guard is a context manager, and AsyncGuard is an
async context manager. Both call close() on exit. AsyncGuard.close() does
nothing; it exists so both guards close the same way.
with Guard([injection(threshold=0.8)], model="jev-latest") as guard:
result = guard.check_input(text)
check_input
Guard.check_input(
text: str,
*,
redactions: Redactions | None = None,
history: Sequence[History] = (),
) -> Result
Checks user text at the input stage. Without redactions, a fresh store is
created and kept on the result.
check_untrusted
Guard.check_untrusted(
text: str,
*,
question: str | Result | None = None,
redactions: Redactions | None = None,
) -> Result
Checks retrieved or third-party text at the untrusted stage. question is
the user question the text was retrieved for.
check_output
Guard.check_output(
text: str,
*,
prompt: str | Result,
sources: Sequence[str | Result] = (),
history: Sequence[History] = (),
redactions: Redactions | None = None,
) -> Result
Checks one complete model reply at the output stage. When the result is
ok, result.onward holds the reply with authorized placeholders restored.
This is the only restoration entry point.
check_tool_call
Guard.check_tool_call(
name: str,
arguments: str | Mapping[str, object],
*,
prompt: str | Result,
redactions: Redactions | None = None,
) -> Result
Checks the tool name and arguments a model chose, at the tool_call stage.
namemust be one line of 1 to 256 characters. Otherwise the check raisesPolicyError.- A mapping in
argumentsmust hold only JSON values. Otherwise the check raisesPolicyError. jes serializes it canonically, with sorted keys and compact separators. - Every judgment receives the tool name. Context-aware judgments such as
tool_safetyalso seeprompt;injectiondoes not. - The arguments are never edited. When the call is allowed,
result.onwardis the original arguments. Judges see them with any secret or personal value replaced, so a found value never reaches a model. secretsandcanaryblock on this stage instead of redacting.piiblocks too, or flags withpii(tool_call_mode="flag").
check_tool_result
Guard.check_tool_result(
text: str,
*,
name: str,
prompt: str | Result | None = None,
redactions: Redactions | None = None,
) -> Result
Checks a tool's response at the tool_result stage. prompt is the user
request that led to the call. Transforms redact as they do on untrusted text.
Context rules
History = Result | Message, oldest first. It is exported fromjes.guard. Atool_resultresult, or aMessagewith role"tool", counts as a tool result, not as an assistant reply.- A
Resultwhoseredactionsis this check's store is already sanitized, so itssanitizedtext is reused. If that result is notok, the check adds a blockingcontext_not_okfinding. - A raw string, or a
Resultfrom another store, is sanitized again from itsoriginaltext. - The store comes from
redactions=, or else from aResultpassed aspromptorquestion, or else is new. If these name different stores, the check raisesRedactionError. - At most 1,024 context entries fit in one check. More block with
too_many_context_items. - Every result has
onward, the one string to forward. See Results.
AsyncGuard
Takes the same arguments as Guard. The five methods have the same signatures,
but they are async.
- A backend with
adecide, such asTypeSafeandFakeBackend, is awaited. One with onlydecideruns in a worker thread. UnlikeGuard,AsyncGuardaccepts an async-only backend. - Transforms and planning run in a worker thread, so they don't block the event loop.
- A check's requests run concurrently, at most
max_concurrencyat once per event loop. Requests still running atdeadline_sare cancelled, and so are all of them when the caller cancels the check.
guard = AsyncGuard([injection(threshold=0.8)], model="jev-latest")
incoming = await guard.check_input(text)
doc = await guard.check_untrusted(text, question=incoming)
call = await guard.check_tool_call("search", {"q": "notes"}, prompt=incoming)
result = await guard.check_tool_result(raw, name="search", prompt=incoming)
outgoing = await guard.check_output(reply, prompt=incoming)
Limits
Limits(
max_input_bytes: int = 1_048_576,
max_context_bytes: int = 2_097_152,
max_chunks: int = 32,
max_items: int = 100,
max_requests: int = 128,
max_concurrency: int = 8,
max_redactions: int = 1_000,
)
A frozen dataclass of per-guard bounds. Pass it as Guard(limits=Limits(...)).
Every field must be a positive integer, or Limits raises PolicyError.
| Field | Bounds |
|---|---|
max_input_bytes | The checked text, in UTF-8 bytes. Larger text blocks with input_too_long. |
max_context_bytes | Prompt, question, sources, and history together, after sanitizing. More blocks with context_too_long. |
max_chunks | Chunks per judgment when the text doesn't fit one request. More block with too_many_chunks. |
max_items | Items, such as URLs, extracted across all item judgments in a check. More block with too_many_items. |
max_requests | Backend requests per check. More block with too_many_requests. |
max_concurrency | Backend requests in flight across the whole guard. |
max_redactions | Sensitive values replaced in one check. More block with too_many_redactions. |
Crossing a limit blocks with complete=False. It does not raise. See
Failures and limits for every label.