I've always found it odd there was no deterministic QA tools with zero baseline needed, especially after so long. It's always been either (1) playwright or cypress that require written tests, dependant on the tests YOU write, or (2), asking the LLM to find and test bugs in your code, which has 2 cons, its costly (sometimes), and its non-deterministic.
That's why i built assay. Its only dependencies is playwright & chromium, and you point it at any folder with a webpage and it runs tests on its own by driving the webpage itself in a browser and clicking every control.
It then flags anywhere the webpage contradicted/disagreed with itself; it doesn't know right from wrong NOR what the webpage is about, but universally, a page disagreeing with itself is almost always a bug.
An example of this is on a paint app, if you draw two strokes on a canvas and click undo twice, you expect them both to undo each stroke sequentially; in one of 225 benchmarks, clicking the undo the first time didn't remove the first stroke, and only the 2nd click worked. assay drove that and caught it.
As for benchmarks, there are 2 variants that are all reproducible in the repo. The first is 225 generated web pages that were first checked by hand, then had assay run on them. There was a total of 20 bugs across all programs, and assay caught 15 with the right reasoning, and had 0 false failures. The second was a harder one where an LLM planted 5 bugs across 10 working original webpages. assay found 12 of the 50 and had 0 false failures, which may seem low, but assays superpower is that its cheap and quick.
assays median runtime is 14s, and it plugs into 15 agentic coding harnesses with a skill, plus claude code and deepseek harness have their own plugin that automatically runs it with a stop hook every time the LLM touches a webpage.
My favourite feature is that it groups flagged tests together if their of the same bug. More details are in the repo about this since this is getting a little long, but the links it makes have its own benchmarks, 51/51 links it got right with 0 false links. This serves as additional context to you, if you use it as a raw cli, or to your agent if your using it with an LLM harness.
pip install assay-ui. All alternative methods to install are on the github, including harness support.
That's why i built assay. Its only dependencies is playwright & chromium, and you point it at any folder with a webpage and it runs tests on its own by driving the webpage itself in a browser and clicking every control.
It then flags anywhere the webpage contradicted/disagreed with itself; it doesn't know right from wrong NOR what the webpage is about, but universally, a page disagreeing with itself is almost always a bug.
An example of this is on a paint app, if you draw two strokes on a canvas and click undo twice, you expect them both to undo each stroke sequentially; in one of 225 benchmarks, clicking the undo the first time didn't remove the first stroke, and only the 2nd click worked. assay drove that and caught it.
As for benchmarks, there are 2 variants that are all reproducible in the repo. The first is 225 generated web pages that were first checked by hand, then had assay run on them. There was a total of 20 bugs across all programs, and assay caught 15 with the right reasoning, and had 0 false failures. The second was a harder one where an LLM planted 5 bugs across 10 working original webpages. assay found 12 of the 50 and had 0 false failures, which may seem low, but assays superpower is that its cheap and quick.
assays median runtime is 14s, and it plugs into 15 agentic coding harnesses with a skill, plus claude code and deepseek harness have their own plugin that automatically runs it with a stop hook every time the LLM touches a webpage.
My favourite feature is that it groups flagged tests together if their of the same bug. More details are in the repo about this since this is getting a little long, but the links it makes have its own benchmarks, 51/51 links it got right with 0 false links. This serves as additional context to you, if you use it as a raw cli, or to your agent if your using it with an LLM harness.
pip install assay-ui. All alternative methods to install are on the github, including harness support.