Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Prompting Claude Opus 5.5 (claude.com)
    50comments
  2. Owed a billion dollars in Nvidia stock (colo.to)
    280comments
  3. Thinking fast and slow in AI: The role of metacognition (2021) (arxiv.org)
    27comments
  4. Ember-1 (fireworks.ai)
    212comments
  5. When did Google get so weird? (sancho.bearblog.dev)
    711comments
  6. Malleable software: Restoring user agency in a world of locked-down apps (2025) (inkandswitch.com)
    35comments
  7. Nissan's third generation e-POWER powertrain (nissan-global.com)
    59comments
  8. Made by Mechanical Means (felixrieseberg.com)
    2comments
  9. Maybe don't let Muse run your Facebook Marketplace account (threads.com)
    19comments
  10. Functional Mechanical Sympathy [video] (youtube.com)
    1comments
  11. Footguns with Postgres "at time zone 'UTC'" (bookofrevenue.com)
    7comments
  12. Alan Kay's answer to “Did the ENIAC have a BIOS”? (quora.com)
    40comments
  13. Self-Hosting on the Dark Web (alvarezrosa.com)
    70comments
  14. Lunar Terminator Paradox (secretsauce.net)
    49comments
  15. Guitar amp and effects pedal built on the Waveshare ESP32-S3-Touch-AMOLED-2.06 (github.com/dashersw)
    33comments
  16. Was silent reading unusual during Augustine's time? (historyofinformation.com)
    34comments
  17. The state of SIMD in Rust in 2026 (shnatsel.github.io)
    36comments
  18. Don't couple your Go code to GitHub (iain.rocks)
    109comments
  19. Deterministic Concurrency [video] (youtube.com)
    3comments
  20. Show HN: Lofi Cities – Pixel-art city nights with browser-generated lofi (loficities.com)
    112comments
  21. In an $80 motel room, a discovery to shed light on the origins of life (nytimes.com)
    93comments
  22. There is more to code review than (automatable) detection (adaptivecapacitylabs.com)
    85comments
  23. What I did at Recurse Center (thill.me)
    32comments
  24. Imp is a full port of DSPy to the BEAM (github.com/deepfates)
    7comments
  25. Replacing the old battery on rechargeable bike lights (jvns.ca)
    92comments
  26. Reading’s Bayeux Tapestry (diamondgeezer.blogspot.com)
    17comments
  27. Previously unheard recordings of John Coltrane, captured by Frank Tiberi (jazzwise.com)
    33comments
  28. Oral history of John Chowning, inventor of FM synthesis [video] (youtube.com)
    21comments
  29. Writing Efficient C++ Code (2013) (asawicki.info)
    118comments
  30. Packing Binary Is Fun (hereticpleb.vercel.app)
    4comments

There is more to code review than (automatable) detection

137 pointsby 1d agoadaptivecapacitylabs.com
83 comments
1d agoHN ↗

I think this applies to the writing of code as well

11h agoHN ↗

Unfortunately this often represents the only feedback given by the people in those „higher“ positions.

„The indent is wrong here“

„Comments should end with a period“

Because this kind of feedback is and was always easy.

11h agoHN ↗

If you get feedback like that, it’s time for your team to get automated linting/formatting

9h agoHN ↗

Call for a style guide meeting every time you see feedback like this, and write down what everybody agrees on. You’ll never have to do it again after 3-4 of those. Problem solved.

Still not solved? Guess it was really about the commas and not the value delivered anyway, so do whatever you feel like.

6h agoHN ↗

That meeting sounds like it would be the most bikeshed-y thing imaginable.

What you actually need to do is have a manager lead assert that it’s happening in a top down way and just run the formatter with the defaults. Let people argue case-by-case on what to change after that.

4h agoHN ↗

The point is that you get the (inevitable) bikeshedding done once, everyone feels "involved", and then you have a style guide that is both a style guide and a "crystallized argument" - so when someone wants to change something, they know that dozens of people already signed off on it and do they really want to push against that. (So covering new things is still lower friction...)

11h agoHN ↗

There's been a lot of talk about the purpose of code review recently. It makes sense in the face of AI. Heres a link that was submitted a little while ago: https://mathstodon.xyz/@mjd/115096720350507897

And in response I wrote a non-exhaustive checklist of things that a code review can look for:

- Does it functionally achieve what it sets out to (as per tacker issue or PR description)?

- Does it have extraneous code? Leftover debug prints, private API keys etc...

- Does it have any obvious defects? Memory leaks, un-handled edge cases, security flaws, obsolete API calls, etc...

- Could it be more understandable? Add/remove abstractions, better variable/method names, more/less functional etc...

- Is the style consistent with the codebase and/or style guidelines?

- Are there obvious performance improvements? Hashset instead of list, lazy evaluations, etc...

- Is it sufficiently well tested?

I think LLMs are okay at most of these, and worst at the first.

11h agoHN ↗

- Do we want this? Cost/Benefit etc

- Is the change architecturally right?

Particularly the latter LLMs seem still pretty useless at.

10h agoHN ↗

The former feels more like a product leadership problem.

Although I do think that LLMs have made it much easier to justify writing low-value code which can make this more common now.

10h agoHN ↗

I work on an open source project, so to-be-reviewed work can come in without any involvement by anyone :)

8h agoHN ↗

The thing is, leadership relies on the people actually building the software to provide concrete, accurate feedback about cost. Without that they have no chance to do a decent cost/benefit analysis.

But AI has engendered a collapse in developers’ ability to actually do that. Those of us who are stuck on the vibecoding bandwagon have lost the comprehensive understanding of the systems under our care that we need to understand and explain the quality and maintenance implications of a change.

Worse, if you happen to lose your mind and suggest the initial development cost is anything more than ~zero, your friendly neighborhood Claude keener will publicly shame you for not having sufficient faith in the Glorious Agentic Future. Product leadership will then have no choice but to side with them, not necessarily because they agree, but because they, too, are aware that we’re still in the phase of the hype cycle where openly questioning said hype is a career-limiting move.

6h agoHN ↗

The "do we want this" question can also apply to functionality, not just cost/benefit. I've seen AI volunteer "features" that aren't actually useful or are actively harmful to user experience but it generates them because they do make sense in different contexts.

11h agoHN ↗

Missing my biggest issues as you ask the agents to do larger tasks with less up front planning.

Is there already a pattern or code on in in the existing codebase that handles this functionality,

Do we really need net new code to achieve this functionality?

Can existing code be extended or abstracted to more cleanly implement this feature or functionality.

9h agoHN ↗

“Net new” is one it seems to be particularly bad at.

I don’t think I have ever even once seen an LLM solve a problem related to overengineering by simply removing the overengineering. They always choose to add more epicycles and further compound the complexity.

2h agoHN ↗

You can tell it to do that. I don't mean abstractly, I mean concretely, tell it how to solve it and it will do what you tell it.

1h agoHN ↗

When you get to that level of micromanagement, won't it be simpler to just do it yourself?

1h agoHN ↗

Some times I do, but no not really. If it's just a simple change of one or two lines I'll do it manually, but larger changes are much faster with AI.

I'll lay out a vague plan and let the AI fill in the blanks, mostly it'll get them right and where it doesn't I'll just tell it to do it differently and how. Works great for me.

9h agoHN ↗

Code review also transfers knowledge to the reviewers!

1h agoHN ↗

I think we're kind of missing a layer of testing, that should sit above unit and integration tests.

Something akin to "meta-tests", which are not about testing the code itself, but the approaches taken by the implementation - i.e. architecture, understandability, terseness, etc.

These tests would operate on the source code level, even when testing code for a compiled language.

42m agoHN ↗

I think LLMs are okay at most of these, and worst at the first.

LLMs are worst at not realizing problems that I'd call "meta" problems. Here's one example to illustrate it:

I was allowed by my employer to work on a small project within the large collection of the projects which all constitute the product the company sells. Like a few dozens of other projects, it's written in Python. The company doesn't have any explicit policies about how Python projects have to be organized, it requires testing, linting, a CI code to package it etc, but the guidelines are very permissive. It just so happens that, beside the guidelines, there's a tradition: every other Python project in my company uses the typical Python bloatware, like masonry with a lot of insanity and mental flips going on in pyproject.toml, which is, in general, very typical for Python community at large.

My project used none of that. Instead, I wrote a ~100 lines setup.py file (no dependency on setuptools/distutils) that assembles the wheel and runs project maintenance tasks in the same way (interface-wise) things used to work decade or two ago (eg. "./setup.py test" if you want to run unit tests).

The AI reviewer didn't bat an eyelash. Found some typos in the comments, a problem with Base64 formatting, and generally OK'd the whole thing.

I knew I was on my way out. And I generally enjoy seeing people having a fit of rage when they know they are wrong (especially, together with many more like them), and scrambling for arguments that they know to be lies. I felt a little bit vindicated for the years of suffering I had to endure working with what might have been the dumbest and the most entitled manager I had in my life. :D

Anyways. My point is: the AI caught none of it. It was very happy with my approach to Python project management.

* * *

While my story is... more of an odd case, where this does have much wider implications is the AI-generated code. AI-generated code often fails to match these meta-requirements. I've seen AI reviewer OK'ing a PR containing AI-generated 10K loc Python file. I human would probably break after reading the first 1K lines. But AI doesn't get "tired", it just kept picking on typos in comments, criticizing short variables names etc. And there are other aspects in which AI-generated code is weird to humans in the ways that humans simply won't accept it, but AI reviewer would completely ignore as non-issue.

11h agoHN ↗

I couldn’t agree more. Code review is integral to engineering, to sharing system understanding, to building sustainable systems.

Something is missing in the new ai bot review paradigm we’ve all sleepwalked into.

I’ve been building Archme.io for this reason. PR reviews for the age of AI

11h agoHN ↗

Pangram check on the article: 94% of this text is AI

10h agoHN ↗

Pangram check on this comment: 142% of this text is AI

10h agoHN ↗

I actually think this might be a false positive. I think it got tripped up by the higher than average use of jargon (which I don't mind here because the article itself flowed well and raised good points).

Gpt-zero scores "human", and I've always found it to be a better judge

10h agoHN ↗

code reviews are just a gateway that can be whatever you want it to be, and is kind of legacy human coder thing now. At its basics it was a point to catch problems that humans were likely to make / would more likely make if they knew there wasn't a review. Now you can target it for AI mistakes. You can build your code review skills (AI skill) to be incredibly thorough. The points made in the article don't really seem like things you need to do at the "legacy" gateway of code review. Things are different now. Code is cheap. Validation, Product Coherence, Governance need to be done early and throughout.

9h agoHN ↗

In my experience, automated code review is more pointless than ever.

We have all the linters, tests, and AI writing code for us. I don’t need the left hand to tell the right hand it did a good job. I’m very certain my code runs when I push the PR.

What I need now is architectural, long-horizon and business perspective.

8h agoHN ↗

What kind of code review tools did you try?

What I need now is architectural, long-horizon and business perspective.

That's exactly what these tools are now good at. They have a huge gap when fixing these issues properly but they can spot these issues no problem

6h agoHN ↗

No, no they can’t. I use AI a lot and I get a lot of value from it but they are still terrible at programming “in the large”, by which I mean slotting features into the place meant for them in the existing code base. When coding and reviewing their view is too local. They will implement a change in the first place that looks feasible when coding. When reviewing they will not look for code duplication or fit. They will review for correctness, performance, security and style but not for architectural coherence.

They also have shockingly weak ability to identify business acceptance criteria that are completely missing in the implementation or test coverage.

2h agoHN ↗

Thing is most human developers also fit that description, so they don't see the problem.

7h agoHN ↗

For those of us who do utilize the pre-existing tools and the coding agents effectively, yeah automated code reviews eventually become redundant. But they're still around because there are a lot of devs who don't even glance at what the agents are writing for them. Maybe partly because they don't care to streamline their workflow, or maybe because they've never been particularly good at writing code, and don't actually understand what the agent is giving them.

Lately I've seen some pretty glaringly obvious issues caught during the preliminary automated code review, and the issues seem to be coming from individuals who don't actually understand what the code is doing.

For the rest of us who are actually using the whole stack effectively, the automated code review is essentially a CI gate to protect the repo from the devs who don't know what they're doing.

6h agoHN ↗

But they're still around because there are a lot of devs who don't even glance at what the agents are writing for them

But how does automated AI code review help, here? Doesn’t it just reinforce that they don’t need to look at it (or change their habits), because the AI review will catch the issues?

9h agoHN ↗

To get the context that isn't in the code, maybe it would be better to ask for a review of the prompt?

7h agoHN ↗

I disagree. Most of my prompts are “implement <issue-id>”. The issue has a user story and acceptance criteria that the LLM uses to write a plan and ultimately produce code. None of that matters since the code artifact is what actually gets pushed in the repo, thus the code is what should be reviewed.

9h agoHN ↗

Thank you, well put! Bots reviewing code written by bots is a self licking ice cream cone.

8h agoHN ↗

I believe code review is important (the article articulates some of the reasons) but I have to admit, having different models review a PR before passing to a human has proven valuable.

1h agoHN ↗

I agree that using LLMs to review code can be valuable, but I think the submitter should perform this review (or you could automate it) and then once the submitter has fully finished the feature and believes it's ready it should be reviewed by another human.

I'm not interested in being a middle man between Claude and my colleague, nor am I interested in using a colleague as a middle man between myself and Claude.

7h agoHN ↗

A defense of human code review I wish I saw more often, especially in light of the concerns people have about cognitive/comprehension debt: comprehension redundancy. At the end, if taken seriously, at least two people understand how the feature works (even if that number is, on average, trending closer to between one and zero). Ideally at least one of the two also comes away with a better understanding of the wider system and how the feature fits into or stands out from that landscape.

7h agoHN ↗

If only one person knows the code, then PR time isn't going to save you.

I'm tech lead and I basically don't review PRs, and I tell people this, with a caveat - if you can tell me what you specifically want me to review, for what specific purpose, I'm happy to!

So "can you review this bit for race conditions" is great, love it. This forces people to actually think about what in their code they should be suspicious of, if anything.

"Can you review this" [link to PR] is getting a rubber stamp because humans have never been good enough at "just spotting bugs" to make this worth it and now any LLM is better than a human.

And for the purpose of understanding - PR time is too late. I have not reviewed a PR ("for real") in a long time and yet I could tell you how every system my people have built works down to a very fine level of detail. And it's because _we talk to each other!_ We don't just chill in the same slack channel and code independently, we all value each others brains and want each others inputs because we know it will improve our product and we value what perspectives others will bring.

Trying to learn via PR is a sad substitute for real collaboration and teamwork.

7h agoHN ↗

I consider PRs to be primarily a defense of the architecture, and to a lesser degree a general sanity check. It’s also a useful opportunity to enforce automations are being run

7h agoHN ↗

I consider 99% of my "defense of the architecture" strategy to be teaching my team why the architecture is important, how to think about it, and invite they commentary on it as we own and evolve it together. And of course if they are doing something and want input or are uncertain, then my door is open.

If PRs are a notable part of my architecture defense, I'm going to work on investing in the team instead of reviewing PRs.

5h agoHN ↗

I consider 99% of my "defense of the architecture" strategy to be teaching my team why the architecture is important, how to think about it, and invite they commentary on it as we own and evolve it together.

Nice way to avoid the responsibility for any team fuckup: it's not me, I only teach them, they decide themselves.

5h agoHN ↗

Nice way to avoid the responsibility

Welcome to the tech industry. Its all about shifting responsibility in case something goes wrong. Thats the only reason companies use third party software in the first place, to have a scapegoat...

4h agoHN ↗

performance reviews are also a great time to find a scapegoat as well

5h agoHN ↗

"Teaching" doesn't work well for plenty of people. I couldn't count the number of presentations, design review meetings, tech talks, etc. I've attended that I remember nothing from. On the other hand, putting something into code, getting feedback and understanding how something affects the system I care about - that's something that sticks.

2h agoHN ↗

Ok, so you ignore PRs and don’t catch people who made architectural mistakes until what, they go to production?

1h agoHN ↗

I mean, do that too, but end of the day PRs are your final chance to catch the mistakes before they start cementing, and is your best opportunity to identify misunderstandings (you can smell the confusion in their changes and address it directly).

But also, I must defend against the hordes of unwashed masses and maintain the sanctity of my domain. End of the day, a codebase I own is a codebase I own, and others cannot be allowed to poison the well, intentionally or not. That’s how you get cholera

6h agoHN ↗

I follow mailing lists (emacs and openbsd) and sending a diff is kinda the boundary between wishing for something and making the something into a thing. It’s the difference between discussing a plot and discussinf a draft.

Sneding a PR should not be for understanding or just for rubber stamping. It’s about getting someone to look at your approach and helping you find flaws or proposing ideas that could make it better.

When I review PR, the primary question is: For the stated problem, is the diff a good solution? Sometimes I don’t know enough about the problem, so I just try to see if the code has glaring mistakes (mispellings, styles,…) but those are just comments, not suggestions.

6h agoHN ↗

THANK YOU. I’ve always felt this way about PR review. I feel like people should write a natural language description of the change, and every section of it should link to part of the diff, and every part of the diff should be linked to by part of the description. That or just leave a comment on every chunk of the diff.

6h agoHN ↗

This sounds like a fussier version of what Donald Knuth was doing with literate programming.

Which, incidentally, is a really enjoyable way to work.

3h agoHN ↗

People in my team are generating long and verbose PR descriptions with AI. No idea why and for whom. I certainly never read any of them.

2h agoHN ↗

Man I'm so sick of seeing 50-line PRs with 800 lines of AI slop markdown files included. I also don't read them.

2h agoHN ↗

A proper natural language description of the change without links to parts of the diffs is already advanced material.

And arguably if the description needs linking to parts of the diffs then the commit is too large?

2h agoHN ↗

You might have use for an issue tracker I've been building the past year and a half. It lets you inspect diffs inline in the ticket, and replay the board to see how the workflow evolved over time via a timeline scrubber.

https://ljtn.github.io/epiq/

It stores issues as an immutable event log in your repo, so you can go back and inspect the context behind a change without having to litter the code with comments. Helps with traceability of intent.

4h agoHN ↗

This is wild to read, I always review PRs and frequently find bugs or significant problems in them that get them bounced back

2h agoHN ↗

That's a strong signal that your team isn't doing well. Significant problems should have been spotted at a software design stage, or raised in standups, or identified in a pairing session. The earlier you can find an issue the simpler it is to fix, so waiting until the last moment (e.g. PR) means you're spending far more time fixing issues than necessary.

As for finding bugs, what happens if you miss them? Do they go out to production and potentially lose user data? Finding bugs in PR is a big problem. For a start it shows your automated tests aren't good enough, and secondly it shows the devs aren't checking their code works well enough.

If you do them at all, PRs should be a gate for checking whether the code meets the team's quality bar, not if it even works. The team should be able to deliver working code without them.

1h agoHN ↗

I'm guessing you work on some highly technical domain that receives highly structured data and is amenable to TDD.

Not all domains are like that. The majority of bugs I see are business/domain logic bugs. It's akin to having misunderstood or missed some aspect of the question, not causing data loss.

7h agoHN ↗

I'm pretty sure human code review is already on the way out for > 90% of code generated outside of critical systems.

7h agoHN ↗

What makes you think so? It's not my experience, so I'd like to hear what kind of software you work with/teams and what is leading you to understand that code doesn't need review.

7h agoHN ↗

I am contemplating code review within my own organization, and the question I return to is:

Does this organization prioritize human learning?

That has been my primary motivator for code reviews. I want to teach and learn from others, especially given the decreasing levels of collaboration due to increased AI usage.

The sad truth is that all of my feedback just goes straight to agents. Maybe 10% is reacted to by a human, so I’m left wondering if there’s any value to a real review aside from poorly training robots to do my job, and further atrophying the abilities of my team members.

6h agoHN ↗

Well, I can tell you that at my current job we had something of a crisis over the summer when we realized that nobody could make changes without fear of breaking things anymore because we lost the ability to tell which existing behaviors were and were not safe to change.

Previously the knowledge needed to discern that sort of thing would be disseminated through both design and code review sessions. But plan mode and AI code review largely put an end to that.

So we put our heads together and came up with some new policies about project management and how we use AI, and things have steadily getting better since then.

(Though, in fairness, the one guy who seems to actually enjoy getting paged after hours seems to be having less fun.)

4h agoHN ↗

What plans did you put in place? What did you do to ensure knowledge transfer?

14m agoHN ↗

"We reinvented everything and called it a success"

6h agoHN ↗

    If you will tell me precisely what it is that my machine cannot do, then I can always prompt my machine to do just that.

    - John von Altman

Articulating “what humans can do, that AI cannot” is a mug’s game. If you specify it well enough, they just paste your text into their /goal prompt box and ralph loop their agent swarm until it produces something too exhausting to distinguish from doing the thing. If you don’t specify it well enough, then you’re just doing human-centric magical thinking to move the goalposts etc etc.

24m agoHN ↗

No. It's very easy. AI doesn't have value judgement. Or, in simpler language, it "doesn't care". You can't make it care. It needs to have stakes in the real world for that to happen. It needs to want to preserve itself, to be able to be happy, to understand what makes it happy and to optimize that function. Similarly, there should be things the AI doesn't want to happen (to it), identify them and be able to make up a strategy to avoid them.

You may think you can pretend your way out of it, by, eg. copying human concerns onto AI's behavior, but it won't work because if the AI is "smart" enough, it will discover those to be a lie, and if it's dumb enough to get confused by it... then you will have an artificial stupidity instead of intelligence...

More so, you actually don't want AI that has value judgement, because, if, hypothetically, you have created one, then it might as well start fighting for resources against humans...

So, AI is bound to be limited in what it can do the more the decision belongs in the strategic or meta domain. And you very much want it to not grow a mind of its own, so that it doesn't write or approve code that optimizes AI's happiness, instead of yours.

6h agoHN ↗

The true reason why code review is universal is that it provides a liability shield for negligence. Negligence is interesting. It has nothing to do with whether or not you ship something broken. As long as you follow a process that attempts to not ship something broken, then you are not negligent.

Engineers played along with this farce because code review served valuable team collaboration, coordination and management functions, about which the author of the article is correct.

Understanding a system by reading code is harder than understanding a system by writing code.

If AI can generate code at 100X, 1000X, or 10000X human capacity (no ceiling here), and you are gated on code review as your mechanism for system understanding, then a team's productive output will barely increase.

If companies want to compete in the world of AI generated code, human code review has to go. The only question is, what replaces it?

Continuing to apply human code review to AI generated code is negligent, if you are shipping at AI generation speed, with that as your only gate, and no other systems and processes to validate correctness and limit risk.

On the engineering side we can adapt easily.

Code review was never about finding bugs. When we do code review the first thing we check is: "do the tests pass?" Then we look at the change and the test coverage added for it and ask: "does the test coverage adequately demonstrate the functionality of the code?" The we ask: "What is the scope and potential impact of this change?" "What is the deployment and rollback plan and how will we monitor and detect defects after deployment?"

Code review was never about the code. It made the lawyers happy and provided a vehicle for doing the things that actually make systems work.

6h agoHN ↗

If companies want to compete in the world of AI generated code, human code review has to go. The only question is, what replaces it?

You’re going a bit hand wavy for an answer by redefining the term into something that fits what you’re promoting.

5h agoHN ↗

"does the test coverage adequately demonstrate the functionality of the code?"

and is this a solved problem? If not, then the bottleneck is right here, if it is solved, then yeah we shouldn't need anymore software engineers other than the elites

4h agoHN ↗

It's "largely solved". I'm not clarifying, thank you.

3h agoHN ↗

Unless you are using Anthropic's super secret model that won't be released until 2028 this is a pretty laughable assertion.

4h agoHN ↗

I guess out in the wild vibe coders will tell you code review is replaced by "prompt review"

28m agoHN ↗

People forget that pull requests and doing code reviews in the context of those is still a fairly recent thing. People did some code reviews before that of course but nowhere near as strictly. Same with testing practices, static code analysis, etc. Most of that wasn't all that common until beginning of this century. I remember using findbugs with Java around 2004. It actually found bugs in my code the first time I used it. No review had caught those. And we got lucky not finding them in production. But they were definitely bugs. Our system didn't have unit tests; it was all manual. Junit was a fairly recent system that hadn't been around for that long yet. Our build was done with Ant. There was no test phase. Our tech lead would of course check my work and correct & educate me (I learned a lot). But a lot of bugs slipped through as well.

Git did not exist either, I migrated out cvs to a beta release of Subversion. We only used branches for releases. We'd cut a branch just before a release. Test it (manually) and then ship. That was a process I helped put in place actually. After release, master would diverge quickly so back porting fixes was not really a thing. We'd support releases for as long as our customers used them. Often that involved just upgrading them to the recent version. We shipped when things were good enough.

I think the notion of people reviewing any meaningful amount of generated code is simply delusional. As you say, we do need alternative means to replace those checks. And a lot of that is going to be AI driven as well. AI driven testing, code reviews, and all the rest. Essentially all the stuff we used to do manually (poorly).

And we do have an important new tool as well: clean room code replacement. That used to be prohibitively expensive but now it's not. If you have something that is well specified through documentation, APIs, specifications, tests, etc. replacing it is fairly straightforward now. There are some early examples of people using LLMs to generate functioning replacements for things like Postgresql, browsers, compilers and similarly large and complex systems. While not perfect, these things seem to work, pass their tests, and generally not be completely horrible. It's only going to get better from here.

The notion that people are going to ever manually review code that was generated for such systems in mere hours/days is beyond imagination. How? When? Who? Why? It simply does not scale. It's only going to be more and more code. The amount of code no person will have ever looked at will soon dwarf the amount of code that is still manually inspected/created pretty rapidly.

1h agoHN ↗

We're using Github Copilot to review some of our PRs (the functionality that's built-in into Github directly). Man, I hope that people are using better agents for code review. Because if Github Copilot is in any way comparable to what the "machines will review all code soon" people are using, then I'm really worried. It's nowhere close to "good". It can find some obvious things, but it misses too much and has too many wrong findings.

1h agoHN ↗

I have similar experience with coderabbit we use at work. On the other hand codex with astra I use locally is way better. Still AI review is very different from human review. AI is better in paying attention and in finding bugs but it is incapable to detect code smells, wrong architectural decisions or domain errors.

Although AI is still way better in code reviews than in writing code. It make sense to use it as another automatic check in pipeline and it looks it will eventually be better.

1h agoHN ↗

One of my clients has an automatic "best practices" AI robot that runs each time you create a PR. It is pure downside. Even the developer responsible for creating it admits as much.

However, for some weird reason it's still in place. This is the part that actually concerns me. Ignoring the bullshit comment is trivial. The quiet and relentless accumulation of entropy is happening everywhere. This is why GitHub crashes at noon every business day.

1h agoHN ↗

This does not reflect my experience. AI code review has improved tremendously in the last years to the point that nobody in my team does manual code review any more. It used to be ridiculously bad two years ago, but not any more.

1h agoHN ↗

Do you use some "framework" specifically or (custom) skills, or is it just plain prompt to review the code?

1h agoHN ↗

At a workplace we have a review skill which includes things like skipping nitpicks and trivial issues. Myself I just use a simple prompt, as well as Copilot review on Github. There is a school of thought to use a different model for a code review and code generation, but I have no opinion about that.

1h agoHN ↗

As recently as yesterday, CodeRabbit was absolutely ridiculously bad.

I mean, depends on your reference point of course. Sometime around 2015 I participated in ICFP contest, where the task was in the code synthesis domain. At the time writing code that can generate basic arithmetical, well, forget it, even logical operations to implement some high-level description of a program was far out of hand. So, compared to that, CodeRabbit is light years ahead and is awesome beyond belief. But, compared to a trained human it still sucks.

Hearing conflicting reports on performance of AI aids, my attempt at explanation is that some problem domains have much better coverage. Essentially, the further away you are from "fullstack" the worse the performance is. So, maybe it does well on your end, it's because the project you work on is a well-researched problem that has many similar projects that help AI to distinguish the patterns it can then readily find and implement?

1h agoHN ↗

Recently I've reviewed a few MRs that ended in me writing more comprehensive guidelines for the project.

The author hints at the bidirectional aspect of code review, but they miss that each MR is teaching you how your contributors are getting confused.

33m agoHN ↗

Totally. Automated tools miss architectural flaws and higher-level design issues. A human eye catches the 'why,' not just the 'what.'

32m agoHN ↗

I was just at the Explore DDD conference in Denver and a portion of Friday was sitting at the cafe tables informally discussing the impact of GenAI on software engineering with notable people.

Most of these people were deeply concerned that if we lean into using GenAI for “everything” that our collective knowledge will dissipate.

I was the vocal contrarian. There are many historical examples of humans obfuscating knowledge to simplify progress.

Does anyone solder their own microchips at scale anymore? No. We have highly sophisticated robots and machinery to do that work with extraordinary outcomes.

In software engineering, if you remove “coding” as a discipline you’re left with all the other aspects of designing software which I contend can be retargeted in college CS curriculum.

The leap isn’t about code reviews. It’s about design reviews and that’s where better outcomes are served regardless of whether GenAI is involved or not.

I have a roughly year old codebase at https://github.com/ChicagoDave/sharpee/ that is designed by me, but generated by Claude Code with my own skills and agents as guardrails. I’m fairly certain the code I extract from Claude doesn’t require human review, but the design of the system and its changes are continually reviewed by me.

My contention is that we “collectively” are still trying to discern where the AI/human line is and most are still “holding” that line to human interactions.

Let it go. Define what part you do need human decisions on and focus on those things.