Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Washington Taxed Its Millionaires. Now the Rich Want It Repealed (nytimes.com)
    —discuss
  2. Microsoft tells nonprofits their deleted M365 data isn't coming back (theregister.com)
    —discuss
  3. Humans Are Reading Copilot Prompts – and They're Horrified (404media.co)
    —discuss
  4. Modders Bring Nvidia's DLSS 5 Neural Rendering to AMD Radeon GPUs (tomshardware.com)
    —discuss
  5. Visual browser of 1,750 YC AI startup homepages (acolorbright.com)
    —discuss
  6. CrankGPT – hand powered offline physical AI box (squeezlabs.github.io)
    —discuss
  7. Show HN: RIFC-light, offline, pseudocode based flow chart generator (github.com/ramathesungod)
    —discuss
  8. I Put 63 Buying Questions to ChatGPT and G2 Was Cited Zero Times (cuescout.com)
    —discuss
  9. Cloudflare: Birthday Week 2026 (cloudflare.com)
    —discuss
  10. How to Tell If Someone Is Cheating on Lichess (chesscheatdetector.com)
    —discuss
  11. Show HN: Jeva.cpp – a llama.cpp fork with JEV-compatible API for all LLMs (github.com/pragmatwice)
    —discuss
  12. Where Trust in Automated Review Comes From (jakelundberg.dev)
    —discuss
  13. Aspiring Firmware Engineer (chrisgammell.com)
    —discuss
  14. Log internet outages and show whether they're on your side or the ISP's (github.com/boazcohenj)
    —discuss
  15. Expert Asterisks (nedbatchelder.com)
    —discuss
  16. Show HN: Decide – Jev decisions in the shell, scripts, and agent skills (github.com/vsekhar)
    —discuss
  17. Analyzing the ClickFix social engineering technique (2025) (microsoft.com)
    —discuss
  18. Galileo's first civil authenticated position fix under spoofing conditions (esa.int)
    —discuss
  19. Dot Plot (Statistics) (wikipedia.org)
    —discuss
  20. Show HN: X402 Inspector (Find misconfigurations in your x402 API) (trango-compute.com)
    —discuss
  21. UK study finds high-mileage electric cars more durable than petrol (electrek.co)
    —discuss
  22. AMD Takes the Lid off of Zen 6 as EPYC 9006 (servethehome.com)
    —discuss
  23. 2026 Small World in Motion Competition (nikonsmallworld.com)
    —discuss
  24. Accountable Algorithms (pennlawreview.com)
    —discuss
  25. Show HN: Senzii – open-source staff scheduling with a native MCP interface (github.com/senzii-app)
    —discuss
  26. The Most Valuable Engineer Isn't Shipping Features – Rails World 2026 Keynote [video] (youtube.com)
    1comments
  27. My time away from tech (2026) (jezhou.com)
    2comments
  28. Tell HN: New Chat Control vote tomorrow
    —discuss
  29. Dstv Installation Glenferness Tech Installations (techinstallations.co.za)
    1comments
  30. Show HN: IngotDB – SQL-based memory for LLM agents (github.com/tjbroodryk)
    1comments

Show HN: OpenAPPA – open-source deterministic guardrails that don't break agents

21 pointsby 1h agoopenappa.com
11 comments
Hi Hacker News! Matvey, one of the authors, is here.

While building enterprise agents, we ran into a problem: the more tools you connect to the AI, the higher the chance it will run out of control and leak sensitive data.

Guardrails, in theory, should prevent this, but the situation is worrying: - Non-deterministic guardrails (LLM as a judge, auto modes, etc.) are vulnerable to prompt injections, or they lack knowledge of the data, making them inefficient (~10% data leaks on our benchmarks). - Existing deterministic guardrails (Cedar, OPA, FIDES, Dogwood) require massive case-specific IF-ELSE-like policies and break agents (~59% utility loss on our benchmarks).

We did something differently.

We’ve taken the best of existing deterministic guardrails and built a policy language that is data-specific, not use-case specific. It lets you scale agents without updating a policy.

On top of that, we’ve added multiple tricks (like a remedy plan or a DualLLM pattern) to help agents operate within those restrictions, raising utility from ~40% to ~90% and making it the first deterministic guardrail that doesn't break agents.

Finally, we’ve designed it to be pluggable into any agent loop with pre- and post-tool-call hooks.

We invite you to check out our benchmarks: https://www.openappa.com/evaluation

Play with it in Claude Code: https://www.openappa.com/claude-code

Try plugging it into your agent: https://www.openappa.com/add-to-agent

Or check the academic paper: https://arxiv.org/abs/2607.24625

We'd love to hear any feedback!

1h agoHN ↗

Finally some determinism in our high-temperature sampling world!

1h agoHN ↗

I still remember the times when ai/ml security was about perturbing pixel gradients to misclassify a panda

1h agoHN ↗

Hi! One of the OpenAPPA authors here. Ask me anything!

My favorite part of APPA is “batteries”: you can run arbitrary programs as part of an authorization decision. For example, a battery could call the GitHub API to check whether a repository is public or private, then use that result to decide whether its contents can be posted to Slack.

58m agoHN ↗

a few days ago I started an agent on gpt-5.6-terra to work on a project, and one of the website pages had a sentence to create GH issues. Agent read it and that was enough to derail and go creating issues with my context

53m agoHN ↗

Guardrails with builtin remediation instead of simply blocking my agent is a mind blowing long awaited experience! Sooo good. Can't recommend more!

49m agoHN ↗

Quick disclaimer, I work at Archestra.

I’ve had the chance to play with OpenAppa for a bit and if there’s one thing that I love with this project: it’s simple to get started with and easy to tweak. imo agentic security shouldn’t have to be painful to setup.

Give it a shot and hopefully ya’ll will find this project useful. It's also open source :)

37m agoHN ↗

i suspect we’ll see more of this: flexible agents but deterministic boundaries. Congrats on launch!

15m agoHN ↗

Really interesting direction. What resonated with me is that you're treating agent security as an information-flow problem rather than a prompt-classification problem. It was not so obvious to me.

A key question I agree isn't just "is this tool call allowed?", but "given everything the agent has read so far, is this information now allowed to flow to this destination?" That feels like a much more fundamental abstraction.

The part I'm particularly curious about is how this will work with policy authoring at scale. What would be the main adoption challenge?

3m agoHN ↗

great question, very practical.

we have a layered answer here: 1) we ship over a dozen "batteries" now (and plan to grow the number) - they contain base annotations for popular services and helper scripts where relevant; 2) we also ship a skill helping you write your own policies for custom services or adopt the default ones based on your specific needs. The criteria "what's acceptable for each particular scenario" varies, there is no "one size fits all" solution; 3) finally, there is a designed placeholder to cover the rest via wildcard AI annotator if needed. The difference between that and regular "auto mode" in coding agents is that APPA's annotator emits local label (e.g. "does this call require a trusted env?"), not wide allow/block, while decision making stays within label algebra.

6m agoHN ↗

the paper is good! thorough. I like it.