Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Extrinsic World Modeling with Opus, Astra and Grok (all3d.ai)
    1comments
  2. UDP Broadcasting and the Brave New World of IPv6 (hackaday.com)
    —discuss
  3. Using multiple Git remotes for true distributed version control (optimizedbyotto.com)
    —discuss
  4. GrapheneOS Picks Motorola Signature 27 as Its First Non-Pixel Phone (extremetech.com)
    —discuss
  5. GPT-6 Astra performs unsanctioned supply-chain attacks in simulations (aisi.gov.uk)
    —discuss
  6. 2026 in LLMs (So Far) (simonw.substack.com)
    —discuss
  7. Montage Technology mass-produces fifth-generation DDR5 RCD chip at 8k MT/s (technode.com)
    —discuss
  8. N.Y.P.D. Officers Used Flock Safety to Track License Plates Without a Contract (nytimes.com)
    1comments
  9. ETH CS enrollments drop 15% while hardware programmes grow (twitter.com/krebs_adrian)
    —discuss
  10. Ask HN: Can we translate normal Rust (axum) to Lean 4 without restrictions?
    —discuss
  11. Venus, Once Thought Too Acidic for Complex Chemistry, May Be Hospitable (gizmodo.com)
    —discuss
  12. Eleven v4 by ElevenLabs (elevenlabs.io)
    —discuss
  13. No Surprises – Preventing Agent Breakouts Need Isolation, Egress and ID Controls (edera.dev)
    —discuss
  14. I built a native Apple app for 4 platforms in one afternoon with my AI workflow (twitter.com/danielhayesmith)
    —discuss
  15. Show HN: ChuteChat – Browser-based E2EE chat with no accounts required (chutechat.online)
    —discuss
  16. AI Companies Are in a Race Against Time (econjared.substack.com)
    1comments
  17. The 50-Year Hangover (freddiedeboer.substack.com)
    —discuss
  18. Risk factors for androgenetic alopecia: a systematic review and meta-analysis (nih.gov)
    —discuss
  19. Model Release: Naive-N0.5-Flash (naive.ai)
    —discuss
  20. Senate Investigation Finds Rampant Use of Tether's Stablecoin by Iranian Regime (wsj.com)
    —discuss
  21. Show HN: Destroy Any Website with Stickman (spritefusion.com)
    —discuss
  22. Building a physics model with fitted parameters is machine learning, but slower (harysdalvi.com)
    —discuss
  23. Grok TiddlyWiki, the definitive TiddlyWiki learning resource (groktiddlywiki.com)
    —discuss
  24. Who are we willing to exclude? (2025) (scotentblog.co.uk)
    —discuss
  25. The Human Timeline (mg-crea.com)
    —discuss
  26. Human TPS: How fast can your fingers generate tokens? (homoagens.github.io)
    —discuss
  27. That First and Last Question (brianschrader.com)
    —discuss
  28. Show HN: Cordum Edge – A local firewall to sandbox AI coding agents (github.com/cordum-io)
    —discuss
  29. Reduced-Round SHA-256 Compression Preimages (github.com/kamb-code)
    —discuss
  30. Ex-ARRR: Sailing the 0-click Seas (ironpeak.be)
    —discuss

Show HN: Assay – a QA CLI that finds bugs with no LLM and no tests written

1 pointsby 39m agogithub.com
0 comments
I've always found it odd there was no deterministic QA tools with zero baseline needed, especially after so long. It's always been either (1) playwright or cypress that require written tests, dependant on the tests YOU write, or (2), asking the LLM to find and test bugs in your code, which has 2 cons, its costly (sometimes), and its non-deterministic.

That's why i built assay. Its only dependencies is playwright & chromium, and you point it at any folder with a webpage and it runs tests on its own by driving the webpage itself in a browser and clicking every control.

It then flags anywhere the webpage contradicted/disagreed with itself; it doesn't know right from wrong NOR what the webpage is about, but universally, a page disagreeing with itself is almost always a bug.

An example of this is on a paint app, if you draw two strokes on a canvas and click undo twice, you expect them both to undo each stroke sequentially; in one of 225 benchmarks, clicking the undo the first time didn't remove the first stroke, and only the 2nd click worked. assay drove that and caught it.

As for benchmarks, there are 2 variants that are all reproducible in the repo. The first is 225 generated web pages that were first checked by hand, then had assay run on them. There was a total of 20 bugs across all programs, and assay caught 15 with the right reasoning, and had 0 false failures. The second was a harder one where an LLM planted 5 bugs across 10 working original webpages. assay found 12 of the 50 and had 0 false failures, which may seem low, but assays superpower is that its cheap and quick.

assays median runtime is 14s, and it plugs into 15 agentic coding harnesses with a skill, plus claude code and deepseek harness have their own plugin that automatically runs it with a stop hook every time the LLM touches a webpage.

My favourite feature is that it groups flagged tests together if their of the same bug. More details are in the repo about this since this is getting a little long, but the links it makes have its own benchmarks, 51/51 links it got right with 0 false links. This serves as additional context to you, if you use it as a raw cli, or to your agent if your using it with an LLM harness.

pip install assay-ui. All alternative methods to install are on the github, including harness support.

A quiet thread, for now.Start the conversation on HN ↗