Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. How Environment Variables Work (infisical.com)
    —discuss
  2. Aho-Corasick Algorithm (compiler.club)
    —discuss
  3. Mosquitoes Are a Choice (worksinprogress.co)
    —discuss
  4. Four Years of Coding with AI (av.codes)
    —discuss
  5. Show HN: Worklaude – learn by speaking; video scenes branch on TypeSafe JEV (worklaude.com)
    —discuss
  6. Grok is obsessed with our docs. Just not the parts we built for it (handsontable.com)
    1comments
  7. Pipe absorbs sound going through it [video] (youtube.com)
    —discuss
  8. The Weird Museum Map for Your Italian Trip (medium.com/counterarts)
    —discuss
  9. Netanyahu claims ability to hack any iPhone and plant false evidence (twitter.com/dllambo)
    1comments
  10. Single-Writer Bottleneck: Why Write Serialization Limits RealTime Graph Workload (memgraph.com)
    —discuss
  11. The Cambrian Explosion in Software (jeffammons.com)
    —discuss
  12. Changing the difference between Codex plans to be Plus=1x Pro 100=5x Pro 200=10x (twitter.com/thsottiaux)
    —discuss
  13. Show HN: An MCP to control other Windows computers (github.com/boxerbk)
    —discuss
  14. Car's data privacy problems are worse than you think (theverge.com)
    —discuss
  15. Show HN: MineSweeper, but guessing is not needed, or forgiven (billpg.com)
    —discuss
  16. Mark Zuckerberg – Colossus (colossus.com)
    —discuss
  17. Is Spotify Down? Streaming Service Reports Issues (cnet.com)
    —discuss
  18. OpenAI blocked its agent's web access. Then it tunneled out through DNS (thenewstack.io)
    —discuss
  19. Show HN: Sensored – streaming-first PII/PHI redaction (github.com/atomicpages)
    —discuss
  20. Florida invokes extinction fears in legal bid to halt OpenAI development (arstechnica.com)
    1comments
  21. Side Hustle Tax Calc – Paste your 1099 income, see what you owe in 30 seconds (botlabcollective.com)
    —discuss
  22. There's a new way to break RSA that's faster than anything we've seen before (arstechnica.com)
    —discuss
  23. Gates Foundation 5 Year Goal: Help 3B People to Use AI in Their Language (gatesfoundation.org)
    —discuss
  24. Star.vision Unveils 'Space.IDC' Computing Constellation with Partners (china-in-space.com)
    —discuss
  25. Reading Progress Bars Are Bad UX (maxschmitt.me)
    —discuss
  26. Strutt's Powered Wheelchair Is Not a Wheelchair (For Legal Reasons) (core77.com)
    —discuss
  27. DIY X-RAY generator made of eBay parts (highvoltageforum.net)
    —discuss
  28. No errors, no warnings, no gods, no masters – HTML Purity is a Fetish (shkspr.mobi)
    —discuss
  29. Oracle Extends Fusion Agentic Applications with Introduction of Fusion Claw (oracle.com)
    —discuss
  30. Apple's New CEO Moves to Overhaul Company to Run Faster and Leaner (bloomberg.com)
    2comments

Every model (incl. Jev) we tested inflates security finding severity

7 pointsby 27m agocasco.com
1 comments
27m agoHN ↗

Eleanor from my team and I collab'ed on this benchmark. A few things worth stating up front because they shape how to read this. We constantly see LLMs saying "THIS IS A CRITICAL SECURITY FINDING" and we wanted to put it to the test. Most LLMs naturally gravitate towards using CVSS for severity scoring.

We wanted a task where model judgment could be checked against verified results from security engineers (ground truth) and root cause "why" security findings are always inflated. TL;DR LLMs still make too many severity judgements without the proper context, so naturally they bias towards the "worst case" scenario. Happy to answer any questions on this topic