Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Side Hustle Tax Calc – Paste your 1099 income, see what you owe in 30 seconds (botlabcollective.com)
    —discuss
  2. There's a new way to break RSA that's faster than anything we've seen before (arstechnica.com)
    —discuss
  3. Gates Foundation 5 Year Goal: Help 3B People to Use AI in Their Language (gatesfoundation.org)
    —discuss
  4. Star.vision Unveils 'Space.IDC' Computing Constellation with Partners (china-in-space.com)
    —discuss
  5. Reading Progress Bars Are Bad UX (maxschmitt.me)
    —discuss
  6. Strutt's Powered Wheelchair Is Not a Wheelchair (For Legal Reasons) (core77.com)
    —discuss
  7. DIY X-RAY generator made of eBay parts (highvoltageforum.net)
    —discuss
  8. No errors, no warnings, no gods, no masters – HTML Purity is a Fetish (shkspr.mobi)
    —discuss
  9. Oracle Extends Fusion Agentic Applications with Introduction of Fusion Claw (oracle.com)
    —discuss
  10. Apple's New CEO Moves to Overhaul Company to Run Faster and Leaner (bloomberg.com)
    —discuss
  11. Show HN: ProgressCove, calm and smart to-do app with Home Assistant integration (progresscove.com)
    2comments
  12. San Diego overbills sanitation customers due to file copy error (voiceofsandiego.org)
    —discuss
  13. Another World Ported to ZX Spectrum (github.com/antirez)
    —discuss
  14. A list of publications that accept cartoons and their associated payment (colemantoons.com)
    —discuss
  15. Turso 0.8: Concurrent writes without SQLite's single-writer bottleneck (turso.tech)
    —discuss
  16. The Genius Trick Behind Exile's Impossible Map (youtube.com)
    —discuss
  17. Every model (incl. Jev) we tested inflates security finding severity (casco.com)
    1comments
  18. Wikifunctions (wikifunctions.org)
    —discuss
  19. Slow running benefits: Boosts in mood and brain function at very light intensity (direct.mit.edu)
    —discuss
  20. Singularity – design a data model, get the REST API and the MCP server (github.com/dantesabatier)
    —discuss
  21. Show HN: Squidbrake – self-hosted approval gateway for AI agent tool calls (github.com/batrapulkit)
    —discuss
  22. Enforce positive security with Cloudflare Application Profiles (cloudflare.com)
    —discuss
  23. Bootstrapping an Infrastructure [pdf] (usenix.org)
    —discuss
  24. Author Dropped from Literary Prize over AI Allegations (plagiarismtoday.com)
    —discuss
  25. Nvidia Donates $12,000 to the Perl and Raku Foundation (perl.com)
    —discuss
  26. TLA+ helped us fix and10 issues in our OSS project (github.com/desplega-ai)
    1comments
  27. People building AI think it might kill everyone. Hear from them directly. (frominside.ai)
    —discuss
  28. Ciao, Control coding agents from your phone (apps.apple.com)
    1comments
  29. SlopOne: Winning the code quality fight with Jev (qlty.sh)
    —discuss
  30. Have we reached peak Markdown? (strata.space)
    1comments

Every model (incl. Jev) we tested inflates security finding severity

7 pointsby 14m agocasco.com
1 comments
14m agoHN ↗

Eleanor from my team and I collab'ed on this benchmark. A few things worth stating up front because they shape how to read this. We constantly see LLMs saying "THIS IS A CRITICAL SECURITY FINDING" and we wanted to put it to the test. Most LLMs naturally gravitate towards using CVSS for severity scoring.

We wanted a task where model judgment could be checked against verified results from security engineers (ground truth) and root cause "why" security findings are always inflated. TL;DR LLMs still make too many severity judgements without the proper context, so naturally they bias towards the "worst case" scenario. Happy to answer any questions on this topic