Hi Hacker News! Matvey, one of the authors, is here.
While building enterprise agents, we ran into a problem: the more tools you connect to the AI, the higher the chance it will run out of control and leak sensitive data.
Guardrails, in theory, should prevent this, but the situation is worrying:
- Non-deterministic guardrails (LLM as a judge, auto modes, etc.) are vulnerable to prompt injections, or they lack knowledge of the data, making them inefficient (~10% data leaks on our benchmarks).
- Existing deterministic guardrails (Cedar, OPA, FIDES, Dogwood) require massive case-specific IF-ELSE-like policies and break agents (~59% utility loss on our benchmarks).
We did something differently.
We’ve taken the best of existing deterministic guardrails and built a policy language that is data-specific, not use-case specific. It lets you scale agents without updating a policy.
On top of that, we’ve added multiple tricks (like a remedy plan or a DualLLM pattern) to help agents operate within those restrictions, raising utility from ~40% to ~90% and making it the first deterministic guardrail that doesn't break agents.
Finally, we’ve designed it to be pluggable into any agent loop with pre- and post-tool-call hooks.
aussi, si vous êtes à Montréal et vous aimerez apprendrez plus sur OpenAPPA, viens à l'evenement CNCF demain soir, le 29 - je ferai un p'tit discours sur l'integration kagent d'OpenAPPA.
Hi! One of the OpenAPPA authors here. Ask me anything!
My favorite part of APPA is “batteries”: you can run arbitrary programs as part of an authorization decision. For example, a battery could call the GitHub API to check whether a repository is public or private, then use that result to decide whether its contents can be posted to Slack.
a few days ago I started an agent on gpt-5.6-terra to work on a project, and one of the website pages had a sentence to create GH issues. Agent read it and that was enough to derail and go creating issues with my context
I’ve had the chance to play with OpenAppa for a bit and if there’s one thing that I love with this project: it’s simple to get started with and easy to tweak. imo agentic security shouldn’t have to be painful to setup.
Give it a shot and hopefully ya’ll will find this project useful. It's also open source :)
While building enterprise agents, we ran into a problem: the more tools you connect to the AI, the higher the chance it will run out of control and leak sensitive data.
Guardrails, in theory, should prevent this, but the situation is worrying: - Non-deterministic guardrails (LLM as a judge, auto modes, etc.) are vulnerable to prompt injections, or they lack knowledge of the data, making them inefficient (~10% data leaks on our benchmarks). - Existing deterministic guardrails (Cedar, OPA, FIDES, Dogwood) require massive case-specific IF-ELSE-like policies and break agents (~59% utility loss on our benchmarks).
We did something differently.
We’ve taken the best of existing deterministic guardrails and built a policy language that is data-specific, not use-case specific. It lets you scale agents without updating a policy.
On top of that, we’ve added multiple tricks (like a remedy plan or a DualLLM pattern) to help agents operate within those restrictions, raising utility from ~40% to ~90% and making it the first deterministic guardrail that doesn't break agents.
Finally, we’ve designed it to be pluggable into any agent loop with pre- and post-tool-call hooks.
We invite you to check out our benchmarks: https://www.openappa.com/evaluation
Play with it in Claude Code: https://www.openappa.com/claude-code
Try plugging it into your agent: https://www.openappa.com/add-to-agent
Or check the academic paper: https://arxiv.org/abs/2607.24625
We'd love to hear any feedback!
Finally some determinism in our high-temperature sampling world!
aussi, si vous êtes à Montréal et vous aimerez apprendrez plus sur OpenAPPA, viens à l'evenement CNCF demain soir, le 29 - je ferai un p'tit discours sur l'integration kagent d'OpenAPPA.
https://www.meetup.com/kubernetes-montreal/events/316689391/
I still remember the times when ai/ml security was about perturbing pixel gradients to misclassify a panda
Hi! One of the OpenAPPA authors here. Ask me anything!
My favorite part of APPA is “batteries”: you can run arbitrary programs as part of an authorization decision. For example, a battery could call the GitHub API to check whether a repository is public or private, then use that result to decide whether its contents can be posted to Slack.
a few days ago I started an agent on gpt-5.6-terra to work on a project, and one of the website pages had a sentence to create GH issues. Agent read it and that was enough to derail and go creating issues with my context
Guardrails with builtin remediation instead of simply blocking my agent is a mind blowing long awaited experience! Sooo good. Can't recommend more!
Quick disclaimer, I work at Archestra.
I’ve had the chance to play with OpenAppa for a bit and if there’s one thing that I love with this project: it’s simple to get started with and easy to tweak. imo agentic security shouldn’t have to be painful to setup.
Give it a shot and hopefully ya’ll will find this project useful. It's also open source :)