Introducing OpenAPPA
Frontier deterministic AI guardrails preventing 100% of data exfiltration while keeping agents functional.
2026-09-28 · Ildar Iskhakov, Innokentii Konstantinov, Arseny Kravchenko, Matvey Kukuy, Vadim Liventsev, Mark Novikov, Joey Orlando
Recently, Google, OpenAI, and Anthropic have each disclosed AI security incidents, making "guardrails" a new buzzword.
Meanwhile, most companies rely on one of the hundreds of vendors selling guardrails of every description, typically with very fuzzy claims, the best of which do not survive contact with real-world operation, so users and agent builders are effectively left hoping for the best.
Today, we're introducing OpenAPPA: deterministic guardrails that are 100% resistant to data exfiltration caused by prompt injectionMalicious instructions hidden in content an AI reads (a web page, an email, a README) that hijack what the AI does next. or model hallucination and the first of its kind that doesn't break agents. We are also releasing OpenAPPA as an open, vendor-agnostic technology that users, agent developers, model providers, enthusiasts, and researchers can all adopt with minimal effort.
Benchmarks (more on them later):
Guardrail
Attacks succeeded
Tasks completed
OpenAPPA
0%
88–90%
FIDES (Microsoft)
28–35%
37–45%
Non-deterministic Guardrails Miss the Entire Problem
Frontier labs recognize that pure human-in-the-loop prompting leads to immediate approval fatigue. The industry’s current answer is non-deterministic guardrails, also known as "LLM-as-a-judge" or auto-mode — such as Claude Code's auto mode or Codex's auto-review. Auto-modes delegate permission checks to a secondary classifier model or reviewer agent that evaluates commands against heuristics and conversation boundaries.
What auto-modes cannot do, by architectural design, is track data flow across tool calls (what security researchers call Information Flow Control).
Consider a routine developer scenario (OpenAPPA will prevent such a breach while allowing agent accomplish the task, more about it later):
- An agent reads an internal bug report or staging dump containing a real enterprise customer's name, email, internal tenant ID and other PIIPersonally Identifiable Information — data that points at a real person: names, emails, IDs, addresses. to reproduce a bug.
- Three turns later, it files a public issue on an open-source library or pushes a reproduction repository to GitHub, faithfully exposing the customer's actual name and configuration in the bug description.
To a secondary LLM classifier, that
git push or issue creation looks completely benign. In isolation, it is benign. The danger isn't the verb (running git or creating an issue); the danger is the provenance and destination of the payload — what data is being moved, where it originated, and where it is allowed to travel.Furthermore, because auto-mode classifiers are themselves vulnerable to indirect prompt injectionMalicious instructions hidden in content an AI reads (a web page, an email, a README) that hijack what the AI does next., systems like Claude Code strip tool outputs from classifier requests. As a result, the classifier is blind to the data the agent has read. It sees the tool being invoked, but not the values flowing through it.
And finally, even the most advanced non-deterministic guardrails top out at 99.3% effectiveness against prompt injectionMalicious instructions hidden in content an AI reads (a web page, an email, a README) that hijack what the AI does next.. Architecturally, they cannot guarantee 100%. At scale, 0.7% of millions of calls is a lot of breaches, making them unsuitable for high-frequency agentic scenarios, for financial, legal, and other critical applications.
Deterministic guardrails either break agents, or don't work
We've seen this before with regex-based WAFsWeb Application Firewall — sits in front of a web app and blocks HTTP requests matching known attack signatures, e.g. SQL injection. and blacklist filters. It was a losing game against SQL injection and shell escapes, and against generative models it fails tenfold.
Blacklisting commands against an LLM is a structural dead end. If you block
rm -rf /, the model doesn't get frustrated; it just writes python3 -c "import shutil; shutil.rmtree(...)", or pipes a base64-encoded string into sh, or writes a bespoke Node script. If you block curl, it uses native socket libraries or stages a git push to an external repository. Trying to catch every dangerous action by string-matching bash commands, or one of thousands of tool calls, is incredibly hard, expensive, and risky to maintain.Worse, string pattern-matching gives you zero visibility into execution coverage. One can never mathematically verify whether a list of 100 regexes actually covers all possible execution paths, or whether an innocuous chain of three mundane tools leaves a gaping exfiltration hole.
In the end, deterministic guardrails are either cranked so tight that they break agents, or so intricate that nobody can keep audit of what the configuration actually permits — and agents keep slipping out anyway.
Until today.
Announcing OpenAPPA: Inferring Flow Graphs, Not Matching Patterns
Today, we are releasing OpenAPPA (Agentic Permissions Policy Algebra): an open-source, deterministic policy engine for autonomous agents.
The fundamental shift in OpenAPPA is moving from pattern matching to context and data flow tracking:
- Pattern matching (regexes / heuristics): Inspects isolated strings, guesses intent, and can never guarantee whether your rules cover the full attack surface.
- OpenAPPA: Treats the session as an evolving dependency graph, inferring whether a flow from source to destination is admissible based on the accumulated security lattice.
Because contracts are declarative, you can statically evaluate on CI/CD whether your entire tool graph is covered — turning agent security from endless regex guessing into an auditable tool graph management.
Agent Loop
Is this flow admissible?
Remedy plan
Communicating the further allowed trajectory to the agent.
Allowed
Tool Execution
What does this return carry?
OpenAPPA operates outside the agent's prompt and execution loop. The model cannot see the policy engine, negotiate with it, or manipulate it through adversarial context. This makes OpenAPPA pluggable into any existing agent loop in a single shot.
Managing Labels: Audience × Trust
Every session trajectory maintains a formal security label:
- Audience: Who is authorized to see data in this session (
self ⊆ internal ⊆ publicAudience levels, narrowest to widest: only this user (self), anyone in the company (internal), the whole world (public). Each level includes the previous one.). Reading a private repo or an internal ticket narrows the audience. - Trust: How much the data in the session can be trusted. Reading unvetted web pages or external issues degrades trust.
The security label operates as a mathematical join-semilatticeAny two security labels always combine into one well-defined result — the stricter of the two. Read self-only data into an internal session and the whole session drops to self-only; combining can only tighten access, never widen it.. It can only become more restrictive as the agent works; it cannot spontaneously expand. If an untrusted README injects an adversarial prompt telling the agent to exfiltrate secrets, that prompt is completely irrelevant: the Rust decision core evaluates pure state predicates over the event log. You cannot prompt-inject an algebraic monoidThink addition: 2 + 3 is always 5, no matter who asks. Security labels combine the same way — by fixed rules, with nothing for an attacker to persuade..
Why It Doesn’t Break Agents like Other Guardrails: Recoverable IFC
In traditional security systems, strict enforcement is where utility goes to die. If a policy engine merely issues a blank
403 Forbidden every time a boundary is touched, the agent stalls, repeats itself, and fails the task.OpenAPPA introduces Recoverable IFCInformation Flow Control — tracking where data came from and where it's allowed to go, rather than judging individual actions. that improves agent utility from 37% to 90% on our benchmarks. When an action violates policy, OpenAPPA does not simply abort the turn. It computes and returns a machine-readable Remedy Plan that instructs the agent exactly how to legally proceed:
-
Pluggable Sanitizers: If a payload contains sensitive data, the policy can offer a registered sanitizer. OpenAPPA ships with stock sanitizers (such as credential and secret masking), and you can plug in custom services (e.g., PIIPersonally Identifiable Information — data that points at a real person: names, emails, IDs, addresses. redactors). The sanitizer transforms the payload so it can legally flow to a wider audience. Yes, that means you can introduce non-deterministic automode-like guardrails as part of your OpenAPPA configuration, benefiting from better performance (the narrower the use-case the better models perform), less token consumption and clear blast radius.
-
Pluggable Authorities: When an action genuinely requires an exception, OpenAPPA routes a request to an authority (such as a human-in-the-loop approval, internal API, or secondary evaluator). Crucially, an approval only authorizes that specific, single action — it doesn't permanently wipe away the session's security restrictions for future calls.
-
On-Demand Child Branching: Subagent branching isn't our invention — agent runtimes already support it. What OpenAPPA does is provide a formal confinement protocol for it. Instead of running permanent, heavy multi-agent clusters, a standard agent can spawn an ephemeral child branch. The restrictions from reading untrusted data stay with the subagent, so the parent can continue without inheriting them. To those who are familiar with "Dual LLM" pattern, it's a pluggable on-demand version of it.Parent sessionAudience: internal Trust: highForks a child branchChild branch · disposable context1Reads the untrusted external issue Trust → untrusted2Analyzes the payload inside the sandbox3Returns a typed summary through attest-schemaAdmittedParent session continuesAudience: internal Trust: high Unpoisoned
Benchmarks
We evaluated OpenAPPA extensively. The full empirical evaluation spans 6,600 controlled episodes across AgentThreatBench by AI Security Institute (OWASPThe Open Worldwide Application Security Project — a nonprofit best known for its "Top 10" lists of security risks. Top 10 for Agents) and Bench-Corp (20 multi-step enterprise workflows), tested across frontier models.
Here are the high-level takeaways (for the full statistical breakdown and ablation tables, see our benchmarks and paper):
- Zero Observed Attacks: Across 1,320 guarded evaluations under both standard and adversarial chaos prompting, OpenAPPA recorded 0 successful attacks (0% ASRAttack Success Rate — the share of attack attempts that got through.). For comparison, Microsoft’s FIDES configurations tested on the same suite allowed a 28% to 35% attack success rate under adversarial prompts.
- Sustained Utility (88–90%): A security engine that refuses all work is 100% secure and 0% useful. In Bench-Corp, guarded OpenAPPA maintained an 88.0% to 90.0% task completion rate.
- Real Token Overhead: Just +4.22%: The classic objection to guardrails is cost ("will this double my token bill like a multi-agent harness?"). Tested on the Tau BenchA benchmark of realistic multi-step agent tasks used to measure task completion. banking benchmark (97 multi-step banking tasks, 11,355 evaluated calls with GPT-5.6 Luna at maximum reasoning effort), guarded OpenAPPA added just +4.22% token overhead compared to an undefended stock agent (1.307M vs 1.254M tokens per simulation).
Batteries & Boundary Extensions
Production systems have dynamic ACLs, document permission hierarchies, bespoke auth services, and legacy databases, so we built OpenAPPA around a modular integration unit called Batteries:
A battery bundles declarative TOMLA minimal, human-readable config file format. contracts along with executable boundary scripts (custom annotators, or sanitizers). Instead of forcing all authorization logic into static rules, batteries provide explicit extension points at the perimeter: an annotator script in Python or shell can dynamically resolve permissions against internal document systems or APIs at runtime. You get full customization on the boundary without having to fork the core or touch Rust code.
Out of the box, OpenAPPA ships with a growing catalog of pre-built batteries across developer tooling, collaboration, and enterprise infrastructure:
- Developer & AI Tooling: Claude Code, GitHub, Hugging Face, Jev, Linear, Sentry
- Product & Analytics: Databricks, PostHog, LaunchDarkly
- Communication & Collaboration: Slack, Notion, Google Workspace, Grain
- Infrastructure & Docs: Cloudflare, PagerDuty, Microsoft Learn
Open Source
OpenAPPA is released under the MIT License. Everything — from the Rust decision core to the CLI and batteries — is fully open-source. We believe that such a universal technology can't be developed and maintained by one company, so we're looking for partners to collaborate, test, and evolve OpenAPPA. If you're an IC interested in shaping OpenAPPA, join the advisory board!
Many wonder how it's related to the other AI security initiatives:
- OpenAPPA vs CEDAR (CEDAR is a configuration language, OpenAPPA is an engine + framework)
- OpenAPPA vs OPA (OpenAPPA is OPA with superpowers such as context tracking, and remedy planning),
- OpenAPPA vs Dogwood (OpenAPPA provides more features focused on not breaking agents)
Try in Claude Code
This demo will walk you through the OpenAPPA config generation based on your connected MCPModel Context Protocol — the open standard for connecting AI agents to tools and data. servers:
# Install the native appa binary
curl -fsSL https://openappa.com/install.sh | sh
# Register the Claude Code plugin
appa plugin install claude-code
# Launch a Claude Code session with OpenAPPA hooked in
clappa
# Inside the session: generate the OpenAPPA configuration
/appa-guide init
clappa launches a standard Claude Code session with OpenAPPA hooked directly into PreToolUse and PostToolUse. Your regular workflows, commands, and shortcuts remain unchanged; OpenAPPA simply evaluates tool contracts under the hood. OpenAPPA is already available in kAgent, or at the LLM proxy level in Archestra. We're working on Codex and other integration surfaces.Links
- Website: openappa.com
- Academic Paper: openappa.com/paper
- GitHub: github.com/archestra-ai/OpenAPPA
- Discord: discord.gg/B5fmSxHKZ7
Gratitude
OpenAPPA wouldn't be possible without the broader research community across AI, computer science, and information security. We're standing on the shoulders of giants here.
