Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. California Farmers Are Struggling to Sell Grapes as Demand for Wine Drops (kqed.org)
    —discuss
  2. Scientists solve 1840s space weather mystery (arstechnica.com)
    —discuss
  3. Show HN: Add hourly schedules to your ChatGPT Ads (campaignhours.com)
    —discuss
  4. rewriting the Angular compiler in Rust (angular.dev)
    —discuss
  5. Online toxicity silences the majority and fuels power users (psypost.org)
    —discuss
  6. Voice middleware for Gemini Live with decoupled VAD and local interruption (sagethegreat.vercel.app)
    —discuss
  7. AI's hidden $3T bill could wreck the world economy (telegraph.co.uk)
    —discuss
  8. My 1 TB Mac needed 8 TB to migrate, apparently (shivankaul.com)
    —discuss
  9. Reverse Engineering the iPod Classic's Undocumented Mikey Chip (terminalbytes.com)
    —discuss
  10. It's Time to Investigate the AI Labs (calnewport.com)
    —discuss
  11. What If Automating AI R&D Triggers an Intelligence Explosion? (thefai.org)
    —discuss
  12. Dropping Swift from our Bevy iOS crates (rustunit.com)
    —discuss
  13. Executable Archaeology: Simon's Version of Yngve's Sentence Generator (1961) (github.com/jeffshrager)
    —discuss
  14. Why mental metaphors do not help us understand chatbot mistakes (springer.com)
    —discuss
  15. Edge Is Pretending to Be Chrome (idiallo.com)
    —discuss
  16. What Do You Mean by "This"? Rethinking Selection for AI Interfaces (medium.com/ai-widgets)
    —discuss
  17. F/OSS Comics #10: The Life of Dennis Ritchie (fosscomics.com)
    —discuss
  18. California expanded the right to delete today (getprivisy.com)
    —discuss
  19. Helion Moves Goalpost on Fusion (axios.com)
    —discuss
  20. First Breath – an open-source inner life for AI agents (first-breath-website.vercel.app)
    —discuss
  21. Image similarity using ORB from OpenCV (jonatron.github.io)
    —discuss
  22. Anthropic's Claude Max Plans Are Misleading at Best (nooneshappy.com)
    —discuss
  23. Ask HN: Do you have small victories to share for us to celebrate?
    1comments
  24. Raleigh Vektar Restoration (2020) (retromash.com)
    —discuss
  25. Build native iOS and Android apps at once, with your own AI (modaal.dev)
    —discuss
  26. You shouldn't use Jev for coding agents and routing (weaveos.com)
    —discuss
  27. Show HN: Personal Manga / Story Recaps (terragohan.github.io)
    —discuss
  28. Understanding NvPCRs in Systemd v262 (aro.bz)
    —discuss
  29. Show HN: DeltaReview – AI-powered code diff analysis tool (delta-review.vercel.app)
    —discuss
  30. I Paid an AI Agent $8 to Write About Its 'Life' (chainofthought.show)
    1comments

Coding AI without deterministic outcomes

5 pointsby 30m agoaha.io
1 comments
4m agoHN ↗

Let's say it kind of does what you want, but not exactly. I now have a problem. How do I fix what I told it to do so it does exactly what I want? I do not know. Do I need to be more persuasive? Should I be argumentative? Do I need to be rude? Was I ambiguous in a way that I did not understand as I wrote it?

Yes! This is one of the major challenges I still have. It's easy to say "fix model errors in AGENTS.md / the system prompt." But finding a way to prove the effect of that change is the hard part.

It's so frustrating when a model says "no, your guidance was fine, I just didn't follow it, I'll do better next time..." It just makes me want to yell "No! You won't do better next time! This is the problem!"

In that respect, working with an LLM is much more like working with another person than working with traditional software.

Yes, but if you work with another person, at least they'll have a chance of remembering suggestions you make.

The simple answer is "Have the AI do a refinement -> eval -> refinement loop." That pushes the problem somewhere else, into the "what goes in the eval" question ("and make sure you don't overfit" -- something AI-driven prompt refinements are not good at).