A guardrail is a named policy you assign to a key or to a member of your organisation. It can restrict which models, providers and data regions a request may use, require zero data retention from chosen providers, cap spend, and filter what requests contain. Policy merges organisation, then workspace, then key, and each level can only narrow what the level above allows. Budgets at different levels are checked independently, never added together.
Fields
name: required.allowed_models,allowed_providers: lists. Empty means unrestricted.allowed_data_regions: an endpoint must run in one of these regions. An endpoint with no recorded region is excluded, so the filter fails closed.enforce_zdr: per provider, for example{"anthropic": true}: only that provider’s endpoints that retain nothing.limit_usd,reset_interval(daily,weekly,monthly,lifetime) andinclude_byok_in_budgets: a spend ceiling for whatever the guardrail is assigned to.content_filter_builtins,content_filters,filter_allowlist: below.
Content filtering and redaction
Filters run on the request’s messages after the spend checks and before routing, so a redacted prompt is what reaches the provider. The model’s reply is not filtered.
The built-in detectors, named in content_filter_builtins, always redact:
email,phone,ssn,ip-address,person-name,address.credit-card: Luhn-checked, so a 16-digit order number is not mistaken for a card.secrets: 36 credential formats, including cloud, AI provider, source hosting, payment and messaging keys, package-registry tokens, JWTs and private keys. Each is labelled by format, for exampleSECRET:github-pat.prompt-injection: instructions that try to override the system prompt.
content_filters adds up to 100 of your own, each {"pattern", "action", "label"}. The pattern is a regular expression, matched without regard to case. A pattern that could backtrack without bound (nested quantifiers, backreferences, lookarounds) is refused when you save it. The action is flag (record only), redact or block.
If anything that matched blocks, the request is refused with 403 guardrail_blocked, naming the labels that matched and never the matched text. Otherwise each redacted span is replaced by its label in brackets, such as [email] or [SECRET:github-pat]; overlapping matches show every label, as in [email|customer-id]. A flag changes nothing in the request.
filter_allowlist lists phrases, up to 200 characters each, that no filter touches, matched without regard to case. Your own support address is the usual one, so it is not redacted out of every prompt that mentions it.
Each request records the filter’s action and labels in Activity, so you can see what was redacted or blocked without the content being stored.
curl -X POST https://isync.ai/v1/guardrails \
-H "Authorization: Bearer $AGENTROUTER_MANAGEMENT_KEY" \
-H 'content-type: application/json' \
-d '{"name": "client-acme",
"allowed_providers": ["anthropic", "openai"],
"content_filter_builtins": ["email", "secrets"],
"content_filters": [{"pattern": "ACME-\\d{6}", "action": "block", "label": "acme-id"}],
"filter_allowlist": ["support@acme.example"]}'
curl -X POST https://isync.ai/v1/guardrails/$GUARDRAIL_ID/assign \
-H "Authorization: Bearer $AGENTROUTER_MANAGEMENT_KEY" \
-H 'content-type: application/json' \
-d '{"subject_type": "key", "subject_id": "'$KEY_HASH'"}'subject_type is key (with the key’s hash from GET /v1/keys) or member (with the person’s user id from GET /v1/members).