Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Claude Isn't Allowed to Write Me Prose (kvit.app)
    —discuss
  2. Religion, Consciousness, and Continuity (alexahn.com)
    —discuss
  3. Rasmus Hauschild on X: "just vibe-coded a flight simulator in ~2 hours (twitter.com/rasmushauschild)
    —discuss
  4. How to do agentic software development properly (fiodar.substack.com)
    1comments
  5. Blygger: AI-intertwingled diachronic-synchronic social publishing – blygger (blygger.org)
    —discuss
  6. Ask HN: Fong Test?
    —discuss
  7. GPT-6.1 Sol: Release Intelligence, Performance and Price (artificialanalysis.ai)
    —discuss
  8. Ask HN: OpenSEO inspiration for other tools, asking for opinion
    —discuss
  9. Efficient Image Steganography Method (nature.com)
    1comments
  10. Anthropic's $518B AI buildout hinges largely on deals that cannot be canceled (reuters.com)
    —discuss
  11. Spencer Pratt on X: most political advisors are so bafflingly incompetent (twitter.com/spencerpratt)
    —discuss
  12. Spaghetti Celesti – Devour the Heavens (punkcake.itch.io)
    —discuss
  13. OpenAI Agent Bypasses Internet Restrictions Through DNS (circleid.com)
    —discuss
  14. Annalist – required GitHub check for what a coding agent wrote (github.com/gregdixonmxn)
    —discuss
  15. Why has Shopify dropped React Native? (pragmaticengineer.com)
    —discuss
  16. The AI margin collapse is gathering pace (martinalderson.com)
    —discuss
  17. An LLM Workflow That Reproduces, Improves, Extends Published Economics Research (nber.org)
    —discuss
  18. Show HN: Relay – a harness for AI coding agents that recover and verify (relayevals.com)
    —discuss
  19. Show HN: Mluva – Dictation and editable rewrites for Omarchy (github.com/1vecera)
    —discuss
  20. You Said No MCP (earendil.com)
    —discuss
  21. Car Is Sharing Data with Big Tech (consumerreports.org)
    1comments
  22. Agentic Hacks, Real Proofs: Inside Google's PageBreak Project (blog.google)
    —discuss
  23. GPT-6.1 Sol (Max): Intelligence, Performance and Price Analysis (artificialanalysis.ai)
    —discuss
  24. Codex Security Cloud (chatgpt.com)
    —discuss
  25. Attack your own AI agent in under 10 minutes – then secure it before deploying (humanbound.ai)
    —discuss
  26. Password-protecting an Nginx proxied site (aweirdimagination.net)
    —discuss
  27. Automating eval design and hillclimbing with Claude (claude.dev)
    —discuss
  28. The Life and Death of Microsoft Clippy, the Paper Clip the World Loved to Hate (artsy.net)
    —discuss
  29. Electrification efficiency: The world will need less energy after the transition (hannahritchie.substack.com)
    —discuss
  30. GPT6.1 Sol
    —discuss

Show HN: Relay – a harness for AI coding agents that recover and verify

2 pointsby 11m agorelayevals.com
0 comments
I'm tired of AI coding agents which are good for 30 seconds and then completely fail on the slightest hiccup: a failing test, a temporarily unavailable dependency, an incorrect assumption. Sometimes I end up having to babysit these agents anyway.

Relay is a harness for coding agents which takes advantage of the fact that agents are often good at doing something slightly wrong, and not so good at doing something correctly and completely. Relay repeatedly tries the task and on each failure, uses the output to find a better way to do it.

It does this by running the task in a loop: attempt the task, run the checks in a fresh sandbox, on failure read the error and retry instead, until the checks pass. What it learned from a run is retained, so that it doesn't repeat the same mistake, and when the checks eventually pass, it will ship the change as a normal git PR - no hidden state, no magic, the diff is human readable.

A few specifics:

- it runs locally, is free, and doesn't require an account to run the local agent - it uses a top-tier model to plan/review, and cheaper models to actually write code (for cost reasons) - happy to discuss why this is a good idea and why it isn't - the checks run in an isolated sandbox for each task, rather than "trusting" the agent - currently supports [list actual models/integrations - Opus 5.5, GPT-6, etc] - there's a paid tier for running this unattended across multiple repos with shared sandbox minutes, the actual agent is free.

It's not magic, it's not AGI, and it will fail badly on some problems.

Would appreciate feedback, particularly from people who've deployed agents unattended and found that they don't actually work, or have had to build substantial tooling around them to get useful results.

A quiet thread, for now.Start the conversation on HN ↗