Limits
Limits is the resource budget a guard gives each check: how much
text, context, and work one check may cause.
Its job
Some of the text jes checks comes from people you don't trust. Without bounds,
an attacker could send a megabyte of crafted text, or a document full of
URLs. That would make one check send thousands of model requests, or fill
memory. Limits puts a ceiling on every axis of that work.
Crossing a limit blocks the check, but doesn't raise. You get a normal
result with ok false, complete false, and a finding that names
the limit. Your code handles it like any other block.
Limits doesn't cap time. That is the guard's deadline_s, which defaults
to 30 seconds.
Mental model
One frozen object per guard, with seven caps. The defaults suit a typical chat app, so most code never sets them.
| Cap | Default | Protects against | When it's crossed |
|---|---|---|---|
max_input_bytes | 1 MiB | A huge text to check | Blocks with input_too_long |
max_context_bytes | 2 MiB | Huge prompt, sources, or history passed as context | Blocks with context_too_long |
max_chunks | 32 | A long text split into too many judgment requests | Blocks with too_many_chunks |
max_items | 100 | Too many extracted items, such as URLs, each judged on its own | Blocks with too_many_items |
max_requests | 128 | Too many backend requests in one check | Blocks with too_many_requests |
max_concurrency | 8 | Flooding the backend. Caps requests in flight across the whole guard | Never blocks. Extra requests wait their turn. |
max_redactions | 1,000 | A text stuffed with values to redact | Blocks with too_many_redactions |
too_many_chunks, too_many_items, and context_too_long on a judgment
follow that judgment's on_overflow. With "allow", they flag instead of
block, and the result is still incomplete.
Using it
from jes import Guard, Limits
from jes.policies import pii
guard = Guard(
[pii(["EMAIL_ADDRESS"])],
limits=Limits(max_input_bytes=64 * 1024, max_concurrency=4),
)
result = guard.check_input("x" * 100_000)
result.ok # False
result.findings # (Finding(policy='jes', label='input_too_long', action='block', ...),)
result.onward # "Blocked: input_too_long."
Set only the fields you want to change. The rest keep their defaults.
Good to know
- Lower the caps when your inputs are small and known, such as short chat messages. A tighter
max_input_bytesrejects abuse early, before any transform runs. - Raise them when long documents are expected and get blocked, for example with
too_many_chunks. Every extra chunk is another model request, so cost and latency grow with them. max_concurrencyis shared by the whole guard, not per check. It is how many threads aGuarduses, or how many requests anAsyncGuardruns at once.- Every field must be a positive integer.
Limits(max_requests=0)raisesPolicyErrorwhen you build it. - A conversation store has its own caps. These are
max_entriesandmax_bytesonRedactions, and crossing one blocks withredaction_store_full.
Next
- Limits reference: the exact signature.
- Failures and limits: every blocking label, deadlines, and backend errors.
- Handling failures: a runnable lesson.