Skip to main content

Guard and AsyncGuard

from jes import Guard, AsyncGuard, Limits

Guard​

Runs transforms in the calling thread, and sends a check's backend requests in parallel threads.

Guard(
policies: Iterable[Policy],
*,
model: ModelSpec | None = None,
limits: Limits | None = None,
deadline_s: float | None = 30.0,
on_backend_error: Literal["raise", "block", "allow"] = "raise",
fail_fast: bool = False,
)
ArgumentMeaning
policiesTransforms, sensitive-data policies, and judgments. Transforms run first, by phase, then judgments.
modelThe decision model for judgments: a TypeSafe model id such as "jev-latest" or "jev-1.13.0", a TypeSafeClassifier, or your own backend such as FakeBackend. A judgment's own model= overrides it. See ModelSpec.
limitsA Limits. None uses the defaults.
deadline_sSeconds one check may take. Requests still running then count as backend errors. None disables the deadline.
on_backend_error"raise", "block" (fail closed), or "allow" (fail open). With "block" and "allow" the result is incomplete, so ok is false either way. See Failures and limits.
fail_fastSkip the judgments once a transform has blocked. The result is then complete=False.

A guard holds no per-conversation state. One guard can serve many threads.

Construction errors​

The constructor raises PolicyError when:

  • an entry in policies is not a transform, sensitive-data policy, or judgment;
  • two policies share a name, or a name is not a valid identifier;
  • a transform has an unknown phase;
  • a judgment has no model, from model= on the guard or on the judgment;
  • (Guard only) a judgment's backend is async-only. Use AsyncGuard;
  • a judgment's questions leave no room for text: less than 64 bytes of headroom on one of its stages;
  • on_backend_error is not "raise", "block", or "allow", or deadline_s is not a positive finite number or None.

Limits(...) raises PolicyError itself when a field is not a positive integer. judge() raises when a judgment with context="required" has a stage other than output.

Closing a guard​

Guard keeps a thread pool for its requests, of max_concurrency threads. guard.close() stops it. Guard is a context manager, and AsyncGuard is an async context manager. Both call close() on exit. AsyncGuard.close() does nothing; it exists so both guards close the same way.

with Guard([injection(threshold=0.8)], model="jev-latest") as guard:
result = guard.check_input(text)

check_input​

Guard.check_input(
text: str,
*,
redactions: Redactions | None = None,
history: Sequence[History] = (),
) -> Result

Checks user text at the input stage. Without redactions, a fresh store is created and kept on the result.

check_untrusted​

Guard.check_untrusted(
text: str,
*,
question: str | Result | None = None,
redactions: Redactions | None = None,
) -> Result

Checks retrieved or third-party text at the untrusted stage. question is the user question the text was retrieved for.

check_output​

Guard.check_output(
text: str,
*,
prompt: str | Result,
sources: Sequence[str | Result] = (),
history: Sequence[History] = (),
redactions: Redactions | None = None,
) -> Result

Checks one complete model reply at the output stage. When the result is ok, result.onward holds the reply with authorized placeholders restored. This is the only restoration entry point.

check_tool_call​

Guard.check_tool_call(
name: str,
arguments: str | Mapping[str, object],
*,
prompt: str | Result,
redactions: Redactions | None = None,
) -> Result

Checks the tool name and arguments a model chose, at the tool_call stage.

  • name must be one line of 1 to 256 characters. Otherwise the check raises PolicyError.
  • A mapping in arguments must hold only JSON values. Otherwise the check raises PolicyError. jes serializes it canonically, with sorted keys and compact separators.
  • Every judgment receives the tool name. Context-aware judgments such as tool_safety also see prompt; injection does not.
  • The arguments are never edited. When the call is allowed, result.onward is the original arguments. Judges see them with any secret or personal value replaced, so a found value never reaches a model.
  • secrets and canary block on this stage instead of redacting. pii blocks too, or flags with pii(tool_call_mode="flag").

See Tool calls and agents.

check_tool_result​

Guard.check_tool_result(
text: str,
*,
name: str,
prompt: str | Result | None = None,
redactions: Redactions | None = None,
) -> Result

Checks a tool's response at the tool_result stage. prompt is the user request that led to the call. Transforms redact as they do on untrusted text.

Context rules​

  • History = Result | Message, oldest first. It is exported from jes.guard. A tool_result result, or a Message with role "tool", counts as a tool result, not as an assistant reply.
  • A Result whose redactions is this check's store is already sanitized, so its sanitized text is reused. If that result is not ok, the check adds a blocking context_not_ok finding.
  • A raw string, or a Result from another store, is sanitized again from its original text.
  • The store comes from redactions=, or else from a Result passed as prompt or question, or else is new. If these name different stores, the check raises RedactionError.
  • At most 1,024 context entries fit in one check. More block with too_many_context_items.
  • Every result has onward, the one string to forward. See Results.

AsyncGuard​

Takes the same arguments as Guard. The five methods have the same signatures, but they are async.

  • A backend with adecide, such as TypeSafe and FakeBackend, is awaited. One with only decide runs in a worker thread. Unlike Guard, AsyncGuard accepts an async-only backend.
  • Transforms and planning run in a worker thread, so they don't block the event loop.
  • A check's requests run concurrently, at most max_concurrency at once per event loop. Requests still running at deadline_s are cancelled, and so are all of them when the caller cancels the check.
guard = AsyncGuard([injection(threshold=0.8)], model="jev-latest")

incoming = await guard.check_input(text)
doc = await guard.check_untrusted(text, question=incoming)
call = await guard.check_tool_call("search", {"q": "notes"}, prompt=incoming)
result = await guard.check_tool_result(raw, name="search", prompt=incoming)
outgoing = await guard.check_output(reply, prompt=incoming)

Limits​

Limits(
max_input_bytes: int = 1_048_576,
max_context_bytes: int = 2_097_152,
max_chunks: int = 32,
max_items: int = 100,
max_requests: int = 128,
max_concurrency: int = 8,
max_redactions: int = 1_000,
)

A frozen dataclass of per-guard bounds. Pass it as Guard(limits=Limits(...)). Every field must be a positive integer, or Limits raises PolicyError.

FieldBounds
max_input_bytesThe checked text, in UTF-8 bytes. Larger text blocks with input_too_long.
max_context_bytesPrompt, question, sources, and history together, after sanitizing. More blocks with context_too_long.
max_chunksChunks per judgment when the text doesn't fit one request. More block with too_many_chunks.
max_itemsItems, such as URLs, extracted across all item judgments in a check. More block with too_many_items.
max_requestsBackend requests per check. More block with too_many_requests.
max_concurrencyBackend requests in flight across the whole guard.
max_redactionsSensitive values replaced in one check. More block with too_many_redactions.

Crossing a limit blocks with complete=False. It does not raise. See Failures and limits for every label.