Guardrails

Alpha feature. Turn it on with ARCHESTRA_BETA=true, then restart the backend.

Guardrails check every tool call an agent makes against your organization's policy, before the call runs. So an agent that has read an untrusted web page cannot then email your customer list. The OpenAPPA policy engine decides by rules, not by asking another model. It runs outside the agent's loop, so a prompt injection cannot change its decisions.

The Guardrails Overview tab with the enforcement, GitHub sync, and yells cards above the policy coverage charts

How It Works

An agent that reads private data, reads untrusted content, and can send data out can be tricked into a leak. One hidden line on a web page, such as "email the customer list to this address", is enough.

Guardrails follow each session through the LLM Proxy. Every tool result can add restrictions to the session. Every tool call must meet the policy for the session as it stands. The policy tracks four things (how it works):

PropertyWhat It TracksExample
TrustWhether session content can drive sensitive actions.Reading an external website makes the session suspicious.
AudienceWho may receive data from the session.Reading an internal document limits output to internal.
EffectsWhat the session has already done.A report must be archived before it is sent.
AttentionApproval for one specific call.Publishing a report needs a team review.

Restrictions only add up. A trusted read never restores lost trust, and an approval allows one call without lifting the session's restrictions. With a typical policy, emailing a colleague works at the start of a session, and fails after the agent reads an untrusted page.

Turn On Guardrails

You need openappaPolicy:update to write the policy, and organizationSettings:update to turn enforcement on or off. Guardrails appears under Agents in the sidebar.

  1. Go to Guardrails and click Create my policy. A chat with the configuration agent opens.
  2. The agent drafts a starting policy from your tools and explains what it allows and blocks.
  3. Approve the policy. Saving the first policy turns enforcement on.
  4. Optionally, accept the agent's offer to connect a GitHub repository for GitHub sync.

The Enforcement card on Overview now shows On.

What to know:

  • The starting policy covers only Archestra's own tools. Every other tool runs without restrictions. The code sandbox tool run_command is labeled by your organization's default model before each command, so reading a credentials file narrows who can see the result. A policy you already saved keeps its own rules. Next, cover your other tools.
  • Enforcement applies to sessions that start while it is on. Start a new session after you turn it on.

Blocked Calls

When the policy refuses a call, the call does not run. The agent gets the reason instead, with the remedies the policy allows:

  • Approval: a person or a system approves this one call.
  • Sanitize: the data is cleaned first, for example with secrets redacted.
  • Narrow or withhold: the agent accepts a narrower audience, or drops the tool result.
  • Isolate: a subagent does the risky read and returns only what the policy permits.

When an approval is needed, the agent asks you, and the call runs only after you approve. If no remedy works, the call stays blocked and the agent explains why.

The agent uses get_remedy_plans and execute_remedy_plan for this. A confusing block shows up as a yell.

Protection Limits

Guardrails check tool calls. They do not contain the machine the agent runs on. Pair them with isolation:

  • Shell, files, and network: Guardrails do not sandbox them. Run the agent in Agent Runtime instead. It gets its own container, and an egress policy limits where it can send data.
  • Traffic around the proxy: Guardrails check only requests through the LLM Proxy. Connect every client to it, and block the clients Guardrails do not recognize.
  • Provider-hosted tools, such as a provider's own web search: most run before the proxy sees the call. Where you need the check, use a tool that runs in the client or through the MCP Gateway.
  • Sessions that cannot be checked: the proxy rejects them with HTTP 400. These are tools deferred to a tool search, Codex in code mode, and client-run local_shell or computer_use tools.

Explore