Call for a style guide meeting every time you see feedback like this, and write down what everybody agrees on. You’ll never have to do it again after 3-4 of those. Problem solved.
Still not solved? Guess it was really about the commas and not the value delivered anyway, so do whatever you feel like.
There's been a lot of talk about the purpose of code review recently. It makes sense in the face of AI. Heres a link that was submitted a little while ago: https://mathstodon.xyz/@mjd/115096720350507897
And in response I wrote a non-exhaustive checklist of things that a code review can look for:
- Does it functionally achieve what it sets out to (as per tacker issue or PR description)?
- Does it have extraneous code? Leftover debug prints, private API keys etc...
- Does it have any obvious defects? Memory leaks, un-handled edge cases, security flaws, obsolete API calls, etc...
- Could it be more understandable? Add/remove abstractions, better variable/method names, more/less functional etc...
- Is the style consistent with the codebase and/or style guidelines?
- Are there obvious performance improvements? Hashset instead of list, lazy evaluations, etc...
- Is it sufficiently well tested?
I think LLMs are okay at most of these, and worst at the first.
The thing is, leadership relies on the people actually building the software to provide concrete, accurate feedback about cost. Without that they have no chance to do a decent cost/benefit analysis.
But AI has engendered a collapse in developers’ ability to actually do that. Those of us who are stuck on the vibecoding bandwagon have lost the comprehensive understanding of the systems under our care that we need to understand and explain the quality and maintenance implications of a change.
Worse, if you happen to lose your mind and suggest the initial development cost is anything more than ~zero, your friendly neighborhood Claude keener will publicly shame you for not having sufficient faith in the Glorious Agentic Future. Product leadership will then have no choice but to side with them, not necessarily because they agree, but because they, too, are aware that we’re still in the phase of the hype cycle where openly questioning said hype is a career-limiting move.
“Net new” is one it seems to be particularly bad at.
I don’t think I have ever even once seen an LLM solve a problem related to overengineering by simply removing the overengineering. They always choose to add more epicycles and further compound the complexity.
I actually think this might be a false positive. I think it got tripped up by the higher than average use of jargon (which I don't mind here because the article itself flowed well and raised good points).
Gpt-zero scores "human", and I've always found it to be a better judge
code reviews are just a gateway that can be whatever you want it to be, and is kind of legacy human coder thing now. At its basics it was a point to catch problems that humans were likely to make / would more likely make if they knew there wasn't a review. Now you can target it for AI mistakes. You can build your code review skills (AI skill) to be incredibly thorough. The points made in the article don't really seem like things you need to do at the "legacy" gateway of code review. Things are different now. Code is cheap. Validation, Product Coherence, Governance need to be done early and throughout.
In my experience, automated code review is more pointless than ever.
We have all the linters, tests, and AI writing code for us. I don’t need the left hand to tell the right hand it did a good job. I’m very certain my code runs when I push the PR.
What I need now is architectural, long-horizon and business perspective.
I believe code review is important (the article articulates some of the reasons) but I have to admit, having different models review a PR before passing to a human has proven valuable.
A defense of human code review I wish I saw more often, especially in light of the concerns people have about cognitive/comprehension debt: comprehension redundancy. At the end, if taken seriously, at least two people understand how the feature works (even if that number is, on average, trending closer to between one and zero). Ideally at least one of the two also comes away with a better understanding of the wider system and how the feature fits into or stands out from that landscape.
I think this applies to the writing of code as well
Unfortunately this often represents the only feedback given by the people in those „higher“ positions.
„The indent is wrong here“
„Comments should end with a period“
Because this kind of feedback is and was always easy.
If you get feedback like that, it’s time for your team to get automated linting/formatting
Call for a style guide meeting every time you see feedback like this, and write down what everybody agrees on. You’ll never have to do it again after 3-4 of those. Problem solved.
Still not solved? Guess it was really about the commas and not the value delivered anyway, so do whatever you feel like.
There's been a lot of talk about the purpose of code review recently. It makes sense in the face of AI. Heres a link that was submitted a little while ago: https://mathstodon.xyz/@mjd/115096720350507897
And in response I wrote a non-exhaustive checklist of things that a code review can look for:
- Does it functionally achieve what it sets out to (as per tacker issue or PR description)?
- Does it have extraneous code? Leftover debug prints, private API keys etc...
- Does it have any obvious defects? Memory leaks, un-handled edge cases, security flaws, obsolete API calls, etc...
- Could it be more understandable? Add/remove abstractions, better variable/method names, more/less functional etc...
- Is the style consistent with the codebase and/or style guidelines?
- Are there obvious performance improvements? Hashset instead of list, lazy evaluations, etc...
- Is it sufficiently well tested?
I think LLMs are okay at most of these, and worst at the first.
- Do we want this? Cost/Benefit etc
- Is the change architecturally right?
Particularly the latter LLMs seem still pretty useless at.
The former feels more like a product leadership problem.
Although I do think that LLMs have made it much easier to justify writing low-value code which can make this more common now.
I work on an open source project, so to-be-reviewed work can come in without any involvement by anyone :)
The thing is, leadership relies on the people actually building the software to provide concrete, accurate feedback about cost. Without that they have no chance to do a decent cost/benefit analysis.
But AI has engendered a collapse in developers’ ability to actually do that. Those of us who are stuck on the vibecoding bandwagon have lost the comprehensive understanding of the systems under our care that we need to understand and explain the quality and maintenance implications of a change.
Worse, if you happen to lose your mind and suggest the initial development cost is anything more than ~zero, your friendly neighborhood Claude keener will publicly shame you for not having sufficient faith in the Glorious Agentic Future. Product leadership will then have no choice but to side with them, not necessarily because they agree, but because they, too, are aware that we’re still in the phase of the hype cycle where openly questioning said hype is a career-limiting move.
Missing my biggest issues as you ask the agents to do larger tasks with less up front planning.
Is there already a pattern or code on in in the existing codebase that handles this functionality,
Do we really need net new code to achieve this functionality?
Can existing code be extended or abstracted to more cleanly implement this feature or functionality.
“Net new” is one it seems to be particularly bad at.
I don’t think I have ever even once seen an LLM solve a problem related to overengineering by simply removing the overengineering. They always choose to add more epicycles and further compound the complexity.
Code review also transfers knowledge to the reviewers!
I couldn’t agree more. Code review is integral to engineering, to sharing system understanding, to building sustainable systems.
Something is missing in the new ai bot review paradigm we’ve all sleepwalked into.
I’ve been building Archme.io for this reason. PR reviews for the age of AI
Pangram check on the article: 94% of this text is AI
it's really obvious too
Pangram check on this comment: 142% of this text is AI
I actually think this might be a false positive. I think it got tripped up by the higher than average use of jargon (which I don't mind here because the article itself flowed well and raised good points).
Gpt-zero scores "human", and I've always found it to be a better judge
code reviews are just a gateway that can be whatever you want it to be, and is kind of legacy human coder thing now. At its basics it was a point to catch problems that humans were likely to make / would more likely make if they knew there wasn't a review. Now you can target it for AI mistakes. You can build your code review skills (AI skill) to be incredibly thorough. The points made in the article don't really seem like things you need to do at the "legacy" gateway of code review. Things are different now. Code is cheap. Validation, Product Coherence, Governance need to be done early and throughout.
In my experience, automated code review is more pointless than ever.
We have all the linters, tests, and AI writing code for us. I don’t need the left hand to tell the right hand it did a good job. I’m very certain my code runs when I push the PR.
What I need now is architectural, long-horizon and business perspective.
What kind of code review tools did you try?
That's exactly what these tools are now good at. They have a huge gap when fixing these issues properly but they can spot these issues no problem
To get the context that isn't in the code, maybe it would be better to ask for a review of the prompt?
Thank you, well put! Bots reviewing code written by bots is a self licking ice cream cone.
I believe code review is important (the article articulates some of the reasons) but I have to admit, having different models review a PR before passing to a human has proven valuable.
A defense of human code review I wish I saw more often, especially in light of the concerns people have about cognitive/comprehension debt: comprehension redundancy. At the end, if taken seriously, at least two people understand how the feature works (even if that number is, on average, trending closer to between one and zero). Ideally at least one of the two also comes away with a better understanding of the wider system and how the feature fits into or stands out from that landscape.