Hacker News

Best stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. When did Google get so weird? (sancho.bearblog.dev)
    958comments
  2. Owed a billion dollars in Nvidia stock (colo.to)
    430comments
  3. Unsealed Briefs in Authors’ Case v. Microsoft/OpenAI (authorsguild.org)
    609comments
  4. Ember-1 (fireworks.ai)
    244comments
  5. Meta Blocks President Lula's Facebook Page, Campaign Ads 2 Weeks from Election (reddit.com)
    299comments
  6. AI companies in race to demonstrate their model most threatening to humanity (thecivilian.co.nz)
    375comments
  7. There are no "rogue" AI agents (eoinhiggins.substack.com)
    266comments
  8. On caring for user data: NeoVim caused Vim undo files to be deleted (aresluna.org)
    341comments
  9. Tells of a Slop UI (hereticpleb.vercel.app)
    235comments
  10. Coding Is Not Solved (alexewerlof.com)
    372comments
  11. Self-Hosting on the Dark Web (alvarezrosa.com)
    106comments
  12. Don't couple your Go code to GitHub (iain.rocks)
    169comments
  13. The problem is not AI code, but not knowing about system architecture or intent (ssp.sh)
    200comments
  14. Show HN: Lofi Cities – Pixel-art city nights with browser-generated lofi (loficities.com)
    132comments
  15. In an $80 motel room, a discovery to shed light on the origins of life (nytimes.com)
    110comments
  16. SpaceX's Starship launching to orbit for first time ever today (space.com)
    293comments
  17. The Normalization of Inexplicable Failures (ihatethefuture.com)
    122comments
  18. Sonnet 5.5 (anthropic.com)
    174comments
  19. What is the size of Yemen? (2024) (theborys.substack.com)
    79comments
  20. Parley: Federated, decentralised chat that speaks plain IRC (mills.io)
    126comments
  21. Windows 11½ (definitelynotwindows.com)
    65comments
  22. Pirating the Pirates (mubi.com)
    76comments
  23. MongoDB CEO resigns to join Meta (reuters.com)
    196comments
  24. If we do not stop to help each other, what do we become? (codinghorror.com)
    94comments
  25. Replacing the old battery on rechargeable bike lights (jvns.ca)
    120comments
  26. Prompting Claude Opus 5.5 (claude.com)
    215comments
  27. SNL Weekend Update: Anthropic CEO Dario Amodei on A.I.'S Threat to Humanity [video] (youtube.com)
    102comments
  28. PostmarketOS is rebranding as Nura (nura.eco)
    60comments
  29. Musk, the Movie (bleeckerstreetmedia.com)
    139comments
  30. Alan Kay's answer to “Did the ENIAC have a BIOS”? (quora.com)
    63comments

Coding Is Not Solved

356 pointsby 5h agoblog.alexewerlof.com
371 comments
5h agoHN ↗

Author here: thanks whoever shared this here. I love the brutal criticism and critical thinking of this community. I'm also fully aware of the emotions this stirs. If it makes you feel better, I'm not here to change anyone's workflow but I'm fed up with paying full price for degrading service. Just last week Github went down due to a stupid retrial error. We also had AI agents going rogue and hacking companies and governments. I use AI (specifically LLMs) every day since they came out 4 years ago. I also build AI-powered products. This is not about being anti-AI. I'm just fed up with slop being pushed as progress. Get your sh*t together. That's all.

If anyone has counter-arguments or cares to make me smarter, I'm all ears.

5h agoHN ↗

I mean, as far as counter arguments go, github had plenty of downtime before LLMs, and they didn't deal with exponential growth then. If you don't count the "it gets harder" side but only counts the "they had problems" side, then yea, that might look bad, but that's not very honest imo.

5h agoHN ↗

Not looking up the outage stats for Github (not sure how accurate they are historically). But going off by what I notice on HN the past months / year, it definitely seems like GH is experiencing more downtown than pre-2022.

But maybe I'm misremembering how fragile GH was in the 2010s.

4h agoHN ↗

Sure. They also have 10x the load or something crazy like that. It's rather incredible that it's still up at all. Microsoft must be pouring crazy amounts of money down that drain and gnashing their teeth at their decision to buy GitHub.

2h agoHN ↗

I don't really give them an excuse for the "10x load increase". The company that owns GitHub is literally a hyperscaler. If they can't handle the scale, there's something seriously wrong at Azure.

5h agoHN ↗

You are too civilized: I recommend threatening drop kicks to insufficiently smart rebuttlers.

5h agoHN ↗

"Coding is solved" will eternally remain 6 mo away, as long as the investors keep pumping in money.

5h agoHN ↗

And as long as we keep changing the meaning of coding!

5h agoHN ↗

Ironic, given I've used an agent to help me simulate a fusion reactor.

Just a simple reactor, my laptop's only little. But still.

3h agoHN ↗

Clearly they should point an LLM to fusion as a problem and 88,000,000 agent*hours later they surely will have a solution! Or say that akhtually it's a harness problem.

/s

4h agoHN ↗

Frankly, I don't really see how it isn't solved, even with the current state of LLMs. Frontier models can write, understand, correct, and optimize code in practically any language at a superhuman level. I haven't come across a single problem that LLMs can't solve. You can easily give them a research paper, ask them to implement it and in an hour or two it's done. Or even point them to a video or screenshot of something and say "implement this feature in our game engine" and... they just do it. It might not be optimally perfect, but what % of human written code is? Even if you ignore the time amortization (given how models can spit out weeks of human work in an hour) they still obliterate even an experienced developer.

3h agoHN ↗

Frankly, I don't really see how it isn't solved, even with the current state of LLMs. Frontier models can write, understand, correct, and optimize code in practically any language at a superhuman level.

It isn't solved because they cannot, in fact, do what you claim. LLMs write code worse than humans do, even "frontier" models.

53m agoHN ↗

Yup. Something the article points out is that they write code better than a few, some, or many humans do - with that list below the six fingers, and how that correlates to whether one is part of the "coding is dead" crowd.

5h agoHN ↗

coding is solved, software engineering not.

5h agoHN ↗

Different words for the same thing. The idea that "coding" is just turning a spec into source code without any engineering decisions to be made was always laughable, for that to happen the spec would need to be as detailed as the source code (of course the whole idea that spec and code are separate things doesn't make a lot of sense).

5h agoHN ↗

LLMs even a version number or two ago can write all the code I've ever been paid to write in the last 20 years; but they are not, I think, yet competent enough to be able to handle the project planning and self-QA I was doing even in my first 6 months of my first job after graduating.

That's not a boast, I don't think I was particularly good at that back then, e.g. I didn't really get how to think about automated tests until much later.

It's just to say that no, coding and software engineering are not the same thing. "Code Monkey" is a dead (or perhaps "undead") role now, but it wasn't always so.

5h agoHN ↗

Coding is not solved but this article hasn't accounted for opus 5.5 yet.

Long term planning in LLMs has not been solved.

5h agoHN ↗

I'm sure Opus 5.5 is smart and probably the next version gets even smarter. The main point of the article is accountability and that's not something we can delegate to AI.

5h agoHN ↗

GitHub Copilot is now written entirely in Rust, with AI agents doing most of the porting work. The migration cost about $120,000 in AI token usage plus about three weeks of a developer's time. The effort updated the runtime module-by-module until the job was completed, spanning over 135 releases across a 14.5-week time period. 430,000 lines of TypeScript were converted into 800,000 lines of Rust.

5h agoHN ↗

This is true, real, and impressive. However, a comment I posted on HN a couple months ago might counterbalance this fact:

GitHub's Copilot cloud agent offering is suffering with a case of some of the worst corporate ADHD I've seen. We built a cloud agentic development pipeline on it, and it seems like almost every other week they silently change something with zero public announcement or documentation that creates real disruption for our team.

That's real, breaking changes to the platform that clearly aren't being tested/reviewed before being pushed to prod. Again with zero public announcement or documentation.

Support is useless – we're paying customers in the 4-5 figures and our tickets go unanswered.

5h agoHN ↗

File by file porting can be done almost always with local reasoning. I don't think it proves much for novel projects which still seems to crumble under complexity past a small sloc limit.

5h agoHN ↗

Porting a system to rust without changing the observable behavior is not that difficult with AI, and porting to a more strict language is not that remarkable. I have a tough time understanding why people equate straight shot porting where a test suite already functionally documents the behavior or where the prior application can be used as an oracle with success in all coding tasks. I would be far more impressed if someone did a clean room implementation of all of GitHub Copilot, from scratch, and got to a better point than the TypeScript or port codebase.

I have no doubt that if you provide any AI system with an oracle with expected behavior that it can match that oracle with some amount of $ and tokens. I haven't seen any demonstration of anything else. Rewriting a codebase was always a challenge for humans not because of complexity, but because of the time and effort involved in matching the old version's prior behavior. It doesn't have anything to do with the serious level of work required to build something truly new from scratch in a performant way.

4h agoHN ↗

Seriously? I have no idea where this cognitive dissonance comes from. Or are people just lying (outwards or to themselves)? A rewrite of this magnitude would easily take a skilled human team months if not years to finish. This is on top of Rust not being an easy language to work with. Which, btw, is the sole reason why not everything is written in C/C++/Rust.

4h agoHN ↗

I'm absolutely saying that AI has sped up the porting and rewrite process! It is amazing! But the reality is that rewriting has been part of programming culture since time immemorial. People want to rewrite for performance or for other reasons all the time, and the cost is now relatively low (i.e., now it's an opex line item in cash instead of time investment). But that doesn't mean that all of coding has been solved.

For example, any amount of software development involves fixing bugs, getting feedback from users on ideal workflows, an iteration loop of performance and bug tuning, etc. AI cannot simply create, from scratch, perfect software. Even using the SOTA models on max effort does not produce bug free software of any meaningful complexity or innovation out of the box. All that has changed is that the act of physically writing code and implementing existing patterns is now effectively a marginal cost.

Most line of business software is not e.g., delivering a company's income. Most software is in back-of-the-house internal products that do various internal tasks. I have no doubt that these processes are now far easier to build.

If the new Copilot is so great, why is it completely out of the current zeitgeist when compared to Codex and Claude Code?

50m agoHN ↗

The fact that it's Rust should be held against the LLM, no for it. Languages with greater type safety are a massive crutch for LLMs (it's also why using an LLM to manage your NixOS install is pretty fucking awesome).

Make it port some Rust to JS, see how that goes.

And Rust isn't a super difficult language to use on a daily basis. Sure, the initial learning curve is obnoxiously steep, but it's arguably easier to use than other languages once you get past that.

The LLM crowd have this habit of equating something that they don't understand with requiring some kind of advanced skill.

5h agoHN ↗

Impressive numbers for a piece of software no one asked for and doesn't make the experience better.

5h agoHN ↗

$120k to port 430,000 loc seems quite expensive. That's dozens of cents per line of code, and equivalent to the all-in cost of a senior engineer in London for a year.

Especially expensive when you take into account the amount of that code which must have been boilerplate & meta-code in nature, meaning it should have been straightforward to move.

1h agoHN ↗

it occurs to me how much better these things endeavors might go if the prompt was

"write a program to transpile A to B" instead of "port A to B"

3h agoHN ↗

If you think it would have taken a single engineer one year to do this, you must not work in the industry. This would have taken multiple engineers at least a year.

2h agoHN ↗

$120k ... equivalent to the all-in cost of a senior engineer in London for a year.

The most shocking thing I've read in this entire thread. Are SWE salaries really that low across the pond? That's entry level in the US. So 90k GBP a year all in, including benefits, employer-paid taxes, seat licenses, hardware, furniture, HR? So what's the after-tax takehome for someone senior enough to convert an enterprise codebase to a new language over one year?

5h agoHN ↗

...and what's your point? Github isn't exactly a beacon of performance or robustness in recent months...

4h agoHN ↗

sounds like 2x the code that no one understands, one more reason to never consider using copilot again

would be curious to know how many times "unsafe" appears in there, have seen rust devs comment on how the ais like to use unsafe to work around difficulties with memory management, like how they will sometimes subvert tests

2h agoHN ↗

That is very cool, but Copilot is hot garbage. So I'm not sure I would cite anything they're doing as a win for LLM coding.

5h agoHN ↗

Coding might not be solved out of the box with these providers, but there are increasingly setups and harnesses that do have a great deal of it solved.

5h agoHN ↗

What I've found is that AI allows lazy and incompetent developers to be more lazy and more incompetent. This then has the effect that product quality suffers more, faster. As a result of the sheer amount of code now being pushed out, code reviews, a thing that previously somewhat prevented lazy and incompetent developers from pushing out horrible code, is effectively dead in the water since no human can actually review such amounts of code realistically anymore. Some companies have adopted AI to review code, which, well ... you have AI make code, AI review code ... I hope you can see the stupidity here if you expect to see any deterministic results at all.

I guess time will tell if the consumer will adapt to the lower quality of products, allowing companies to justify the existence of lazy and incompetent developers, or if the consumer will push back, forcing companies to increase the quality of their developers.

Note: I use AI every day and it is entirely possible to create high quality software with it, so long as you are not lazy and incompetent.

5h agoHN ↗

Even if you are competent I cannot review your 5,000 lines of code you produce per day vs the 100 you were producing before the LLM apocalypse.

5h agoHN ↗

5,000 is the output velocity of someone not fully immersed in agentic coding. I've seen repos do ~100k to ~250k loc changes per week.

4h agoHN ↗

I mean these guys are not even pretending to be reviewing the code.

It just gets “reviewed” by an LLM, which will find a nitpick while ignoring the huge fire in the core of the design, force the planner to make even more sloppy code to cover for an irrelevant test case. Rinse old tokens and repeat until you hit limits.

4h agoHN ↗

This is exactly what I've seen.

For example, I recently got brought in to help with quality on a large-scale system that had been ported to a new platform with the help of coding agents. The project was completed and declared operational in record time, but soon after the business discovered that:

1. The promised scalability improvements did not materialize. Instead, it got worse.

2. Observability had been lost. The telemetry was no longer trustworthy.

3. Users stopped trusting it because it was producing incorrect outputs.

What I ended up discovering was that, while it scrupulously kept existing automated tests passing, any behavior that wasn't explicitly covered by a test was free to change any which way. And there were plenty of small things that weren't explicitly covered. Perhaps because the original authors thought they were so obvious and commonsense that they didn't need one, perhaps because mistakes happen. The why doesn't matter. The point is that reality is messy and imperfect, so giving someone a chance to look at things and think, "Huh, that's funny..." is an essential part of defense in depth.

The real worst part was, this whole replatforming was a huge waste of time, anyway. The improvements they were looking for could easily have been accomplished with some controlled incremental changes to the original system. Mostly just removing a few basic and well-known performance antipatterns.

But way back at the outset, the person in charge of the project asked their agent, "What's the best way to X," and the agent gave them a trendslop answer about how Y alternative technology is more scalable and we should just port to that. It was convincing and they were under intense time pressure to just ship some code because leadership is bought into the AI hype and now has the patience of a 4 year old, so they just went with it.

4h agoHN ↗

Yes, they're certainly squeezing 500 lines of functionality into 250,000 lines of code. Agents are great at this.

3h agoHN ↗

Tell me you're not using a frontier model without telling me you're not using a frontier model

3h agoHN ↗

Where's the great software, then? I'm genuinely asking: where is it? Because I can't find it, and it's been close to a year since AI for programming has started to take off.

3h agoHN ↗

In fact, a lot of the once great software that's switched to AI driven development has gotten worse.

3h agoHN ↗

Tell me you never once bothered to look at the generated code without telling me you don't look at the generated code.

Luajit is under 80,000 lines of code.

1h agoHN ↗

i have unlimited tokens and i throw Fable / Astra at everything. They suck ass still for anything nontrivial. I could commit that garbage but if I kept doing it, I will end up with a ball of mud only Fable / Astra can grok..convenient for Dario and SamA..

32m agoHN ↗

I've asked Codex with GPT 6 Astra to *review" a one time benchmarking script for any mistakes (built by Claude Code using Opus 5.5) and it refactored the shit out of it claiming all sorts of stuff without even asking about the context in which it was developed.

If I was to employ them to review the code without giving each the same baseline multi-page prompt, they go into endless loop of "improvement" with no end goal in sight.

More and more frequently, I instruct frontier models to stop and go back to the task at hand.

4h agoHN ↗

Can the users of the software even keep up at that point? We may have reached diminishing returns on software production, and not enough impact on the rest of the process.

3h agoHN ↗

So where is all that new software? My laptop and phone run essentially the same software as 2 or 3 years ago. Yes there were some minor updates to some apps, but nothing faster than in the years prior.

Where does all that supposed productivity go?

3h agoHN ↗

The only real updates I've seen to anything have been AI features... so all this AI is only being used to add AI to stuff. Most of which the average person doesn't seem to want or use.

43m agoHN ↗

But also, having significantly more code churn doesn’t necessarily mean there is more or better software.

In fact, having more churn can lead to worse software due to diverging patterns and inconsistency

38m agoHN ↗

The question was "where is all that new software?".

27m agoHN ↗

I see somebody asking "where is that great software", but I think it's obvious to everyone that there is more (mostly bad) software.

11m agoHN ↗

To quote the article: "Don't confuse motion with progress."

30m agoHN ↗

What kind of applications are people building that involve 250k loc a week? Genuinely trying to understand this.

19m agoHN ↗

So far I mostly see a metric sh.. ton of meta- and meta-meta-projects re-wrapping AI wrapper tools, with sloppy slogans like "One Model, Five Harnesses. Combined." or "You run in the park. Rrrunnnrr.ai runs your AI." Looking inside, out of 250k it's often 200k of verbally incontinent self-explaining comments; or "smart" redesign of builtins.

No new browser, no new iOS clone than runs on Android, no new easy to use DaVinci, no new CAD suite, no $5 SolidWorks clone, no redesigned K8s, no 10x performance speedup in Linux kernel.

3h agoHN ↗

That's okay. Reviewing the code will become the agents' job as well.

A couple more step functions in model capability of the type we've seen in the past year, and there will pretty much be no reason for humans to be involved in the development process at all. All humans would need to do is communicate clearly what needs to be made and flag problems as they come up.

3h agoHN ↗

"A couple more step functions" is doing a lot of heavy lifting here

3h agoHN ↗

It's the AI bro's mantra.

It didnt go wrong

And if it did, it was because you werent using the latest model.

And if you were, it was because you didnt have the appropriate guardrails.

And if you did, it's because you didnt have AGENTS.MD.

And if you did, it's because you didnt prompt it properly.

And if you did, it you're still going to be redundant soon because I'm sure the next model released will fix whatever went wrong.

1h agoHN ↗

I've stopped calling out Claude mistakes on team meetings because this is so true.

I mean, sure, I could have predicted in what ways an LLM would fuck up, but there's just so many ways I can't keep up.

We just had a major production issue because someone's LLM wrote queries against dev databases. Which are very obviously dev databases because they are labelled with dev in the name, and in the table descriptions. AI reviewer didn't catch it, neither did the human reviewer for that matter.

2h agoHN ↗

All humans would need to do is communicate clearly what needs to be made and flag problems as they come up.

Kinda what i'm doing already, but for the young startup I'm at that's surprisingly tons of work. I miss the days we wrote code by hand boy those were fun 8.5 hours workdays.

25m agoHN ↗

All humans would need to do is communicate clearly what needs to be made and flag problems as they come up.

Sounds like the easiest thing in the world: I wonder why did we not think of it earlier?

5h agoHN ↗

a thing that previously somewhat prevented lazy and incompetent developers from pushing out horrible code

Brings to mind this classification https://en.wikipedia.org/wiki/Kurt_von_Hammerstein-Equord#Cl...

"""I distinguish four types. There are clever, hardworking, stupid, and lazy officers. Usually two characteristics are combined. Some are clever and hardworking; their place is the General Staff. The next ones are stupid and lazy; they make up 90 percent of every army and are suited to routine duties. Anyone who is both clever and lazy is qualified for the highest leadership duties, because he possesses the mental clarity and strength of nerve necessary for difficult decisions. One must beware of anyone who is both stupid and hardworking; he must not be entrusted with any responsibility because he will always only cause damage"""

5h agoHN ↗

The problem here is that AI is consistently one of the four things: hardworking. This makes it very efficient at transforming "stupid and lazy" inputs into "stupid and hardworking" outputs.

Now instead of 90% stupid and lazy (harmless, useful for grunt work) you have 90% stupid and hardworking (aggressively causing damage).

4h agoHN ↗

I'm going to have to remember this, gold comment

4h agoHN ↗

We developed languages that removed GOTO so that developers don't shoot themselves in the foot. We will surely develop harnesses that will ensure that majorly occurring problems are solved before they hit production.

4h agoHN ↗

Goto is a syntax feature that can be trivially removed. Good luck removing "fundamental architectural flaws".

3h agoHN ↗

We solved [trivial problem]. We will surely solve [incomparably harder problem].

Based on what? This will not happen!

3h agoHN ↗

Since the output of human software work is code and AI software work is _also_ code they are both liable to shoot themselves in the foot in the same manner.

You see this already, LLMs are a lot more reliable in statically typed languages with strong memory guarantees (like typescript or rust) than in weaker languages.

IMO the only way LLM code can avoid most of the pitfalls of human code is if we make new programming languages targeted at being used by LLMs exclusively. Think of languages with very strong methods for formal proofing and stuff like that.

The problem is that even if said language was invented, it would still fail catastrophically when integrated with systems not made in said language. We are very lucky that relational databases already provide a somewhat high level of formal proofing in this regard.

Said language would be impossible to parse by humans, kinda like assembly where you can parse what an isolated piece of assembly code is doing, but if you can't comprehend a somewhat large pure-assembly codebase as a whole.

2h agoHN ↗

You see this already, LLMs are a lot more reliable in statically typed languages with strong memory guarantees (like typescript or rust) than in weaker languages.

It’s common advice to wire in deterministic feedback to your workflow with LLMs - static languages aren’t inherently better for LLMs, it’s that LLMs produce better code when given deterministic feedback, such as compiler results.

2h agoHN ↗

I've written very large assembly codebases, it's no different than writing in any other language. You have functions you call with inputs and outputs - though usually those are pointers to memory locations. The program is not one long function, you can split it up into different files and folders and keep everything very well organized and easy to understand and reason about.

3h agoHN ↗

Good thing GOTO was the only footgun that was ever invented in a formal programming language.

3h agoHN ↗

And another corollary is the formerly golden lazy and clever are also transformed into lazy and productive because they no longer need to apply their cleverness to get results...

4h agoHN ↗

I'm both clever and stupid, depends on the day.

5h agoHN ↗

Limitations of AI are a thing; but one rhetorical point keeps coming up (I don't think it's just you) and confusing me:

I hope you can see the stupidity here if you expect to see any deterministic results at all.

Are you expecting humans to be deterministic in the code they produce?

5h agoHN ↗

Someone who knows that 1 + 1 = 2 will not decide that it's suddenly 3 unless we start accounting for health problems. Making mistakes is not the same as non-deterministic.

4h agoHN ↗

Someone who knows that 1 + 1 = 2 will not decide that it's suddenly 3 unless we start accounting for health problems.

And?

The p(that kind of error) is pretty small now. At what point does a probability coming out of an LLM look like "knowing", such that spitting out the wrong answer despite that probability looks like a health problem, a typo, or even just boredom? (Thinking of the Lizardman constant here: https://en.wiktionary.org/wiki/Lizardman%27s_Constant)

It's a continuum for both them and us, even if the mechanism is wildly different.

Making mistakes is not the same as non-deterministic.

i.e. when the dismissal is "non-deterministic" when it should be "Making mistakes", is itself a mistake.

4h agoHN ↗

Lol, humans make such absurd mistakes (and worse) all the time through simple typos, which is effectively random. The key for 2 is right next to the key for 3, after all.

4h agoHN ↗

Neither will LLMs. That's not how their nondeterminism works.

4h agoHN ↗

I think that actually reinforces the distinction being made. An LLM’s nondeterminism is in the generation process: given the same prompt and model state, sampling can produce different outputs. That doesn’t mean the underlying fact itself becomes nondeterministic.

A human who knows 1+1=2 can still say “3” because they misread the question, misspoke, were distracted, or made some other cognitive error. Likewise, an LLM can output “3” because the generation process selected an incorrect continuation. Those are both errors in producing an answer, not evidence that 1+1 somehow has multiple answers.

So yes, human mistakes and LLM sampling are mechanistically different. If your argument is that LLMs and humans can both make mistakes, then major question here is why are we building out huge amounts of infrastructure at unsustainable spending levels to enable LLMs to make the same mistakes as humans.

3h agoHN ↗

If your argument is that LLMs and humans can both make mistakes

It's not, I'm just pointing out that LLMs won't make that mistake.

You could ask an LLM what 1+1 is, and the number of times it says "3" is so small that it makes no sense to worry about it. It will phrase the response differently each time; that's the nondeterminism. But it won't say "3".

then major question here is why are we building out huge amounts of infrastructure at unsustainable spending levels to enable LLMs to make the same mistakes as humans.

Yes, if we ignore everything else, that seems like a reasonable question. But let's not ignore everything else, like the fact that LLMs are much more productive than humans and likely already make fewer mistakes than the average programmer.

2h agoHN ↗

You could ask an LLM what 1+1 is, and the number of times it says "3" is so small that it makes no sense to worry about it...

I think the disturbing fact is that you can take a frontier model with all the intelligence of humanity, and make it say 1 + 1 = 3, by specifically training for it...

A human with that much knowledge will refuse that attempt. There in lies the difference..

1h agoHN ↗

Someone who knows that 1 + 1 = 2 will not decide that it's suddenly 3 unless we start accounting for health problems.

But this is plainly false. This kind of unforced error occurs all the time.

For example, once when I was in high school I traced an error in my math homework to an intermediate calculation of "2 + 2" as being "3". There was no reason.

What we can say about humans is that, if they know that 1 + 1 = 2, (a) they are unlikely to change their mind about this in any kind of lasting or permanent way, and (b) the rate at which they will mistakenly produce other values for 1 + 1 is very low. But it will happen occasionally, and when it does happen, "they just suddenly decided on the wrong value" is an extremely accurate description of what that looks like.

4h agoHN ↗

The difference is that with llms you have multiple levels of nondeterminism compounding each other

3h agoHN ↗

By not reviewing, reading, or understanding the code generated by agentic LLMs the output is effectively like a compiler. However, a compiler has deterministic behaviour that can be repeated and verified.

The behaviour/output of an LLM is not like that. Ask an LLM to create a dashboard to show games by genre and it will generate different results with each run, and each model/model version produces wildly different results.

3h agoHN ↗

I've seen LLMs do something correct 98% of the time then randomly do something crazy that a human would never do because we have continual learning

As humans we don't have our memory reset multiple times per day

1h agoHN ↗

I've seen LLMs do something correct 98% of the time then randomly do something crazy that a human would never do because we have continual learning

I've seen humans vote for Brexit, re-elect Trump, ask questions clearly already answered in an FAQ, try to pull on a door labelled "push", and insist on giving me homeopathic silicon dioxide pills* that cost £5** for a 10-12 gram packet.

Continual learning is a difference, but not by itself a reason to care about "deterministic results".

Nor, indeed, correct results.

As humans we don't have our memory reset multiple times per day

Humans need sleep well before they can read a million tokens' worth of written text. We're more like 300k tokens if you're actually reading and not skimming for 16 hours straight.

Again, different (in soooo many ways), but this isn't a relevant difference when the topic is "deterministic results".

* yes, sand: https://dailymed.nlm.nih.gov/dailymed/fda/fdaDrugXsl.cfm?set...

** and that was what it cost in the 90s

17m agoHN ↗

If I do the same task 100 times I'm not going to suddenly do it crazily different at time 101 because I've built in the memory of how to do it

There is no RNG involved when I decide to push vs pull the unlabeled door to my building every morning

You can put stuff in context to deal with this but you can't do that for everything

5h agoHN ↗

"optimize this code", "fix this code", "extend this code", "add this feature", "find errors and patch them", "find bugs and fix them", "rewrite this from python to rust".

This is all that's needed to actually use LLMs nowadays. How is it a "multiplier" rather than an "equalizer"?

4h agoHN ↗

How is it a "multiplier" rather than an "equalizer"?

Because without the responsible human engineer in the loop, it'll all gradually decay in a cascade of edge-cases. This happens with human written code as well (every "we'll replace this prototype before we ship" you've ever worked on), but with LLMs it happens at 10-100x the rate.

3h agoHN ↗

every "we'll replace this prototype before we ship" you've ever worked on

These so rarely get replaced

3h agoHN ↗

This is why it's good to not keep your prototype a pile of shit as it grows to 5k, 10k, 50k, 100k lines of code.

4h agoHN ↗

Do you use the word "equalizer" in this context to mean that AI has made the playing field equal for both competent developers and laypeople? Do you reckon that competence plays no role these days?

4h agoHN ↗

If that's how you create software then you belong to the lazy and incompetent group in my book. I provide AI with valuable context such as code coverage information, architecture analysis, test requirements, important "gotcha's" that a competent engineer would know about in their architecture or system etc. I'm still very much the person who comes up with the solutions. For me AI is replacing the code editor, it's not replacing the thinking.

4h agoHN ↗

The skill floor has definitely been lowered, but if this were actually true then firms would be replacing senior software positions with entry level ones, not the other way around.

2h agoHN ↗

Not completely, just as a very personal example:

In optimizing my game I noticed framerate hitching even after efficient algorithms were in place for expensive stuff, which was caused by shaders not being precompiled consistently or assets not being preloaded in time. The Agent who'd been profiling and optimizing had moved many preloads to a loading screen, which caused a long loading lock, and what it didn't move ahead was loaded and compiled at use, creating slow frames since work was being done on the main game loop.

I instructed the agent to create a speculative pre-warming/compilation priority queue with a per-frame budget, with priority being determined by likelihood signals that the asset or shader will be used soon. Then I had the AI run fully headed games and hunt down causes for frames going over 16ms, and work through them until a batch of games had fewer than 1/1000 frames >16ms and no frames over 60ms after a short initial settling period.

The approach, the metrics, the validation system and the loop were "prompt engineering" above and beyond what I would expect from someone who was merely "vibe coding a game."

2h agoHN ↗

If this is how you use LLMs, you are the problem.

5h agoHN ↗

you have AI make code, AI review code ... I hope you can see the stupidity here...

You will be surprised how many times, catches errores made by the AI coding agent. However,as you point, isn't deterministic. And you can guarantee the end results is 100% fine code

5h agoHN ↗

Note: I use AI every day and it is entirely possible to create high quality software with it, so long as you are not lazy and incompetent.

What I in general try to teach the other people about AI: It can be a great tool, but check the results! Especially in the case of engineering: Check and then double check.

4h agoHN ↗

Honestly, I would still rather commit claude written code from lazy and incompetent developers than code that they wrote.

4h agoHN ↗

What I've found is that AI allows lazy and incompetent developers to be more lazy and more incompetent. This then has the effect that product quality suffers more, faster.

Yeah. To me it seems very much like the "use dynamic typing for everything" fad. You had a bunch of junior and/or incompetent developers who went around insisting that type declarations are bad, static typing slows down development, you just code so much faster if everything is dynamically typed. And in the context of a new project, they were totally right. It took a few years for the debt to finally catch up, and people realized that these massive, untyped monoliths they had were unmaintainable. Now the two biggest dynamic languages (Python/JavaScript) are effectively typed languages, because nobody uses their untyped variants for serious work.

Dynamic typing still has great uses -- interactive data exploration, putting together quick scripts (though less relevant with AI...), or even just simple prototypes -- but what we tried to do with it at the start, as an industry, was clearly dumb as hell. I suspect we'll look back in 5-10 years and realize that with some of the stuff we're doing with AI, too. It's already happened with things like Gastown.

3h agoHN ↗

The web wouldn't have taken off without dynamic typing, PHP first of all (and Python/JavaScript after that). People seem to forget how atrocious it was to write an .asp or .jsp (I think the extension was .jsp) page back in 2003-2005.

3h agoHN ↗

hey man, don’t knock the JSP, I just edited a few :) it is alive and kicking in 2026

4h agoHN ↗

AI allows lazy and incompetent developers to be more lazy and more incompetent.

I like to put this as "LLMS give lazy and incompetent developers more runway."

3h agoHN ↗

I'm seeing this too. I've worked with devs that would previously push PRs that wouldn't work or run correctly. Those PRs wouldn't get merged in. Now they're putting up PRs which seem to work at first glance, but have hidden problems. For example, one guy introduced a huge PR for a visualization and it seemed to work fine, though another dev mentioned to me that we already use recharts and it does 90% of what this guy's PR does (his code does all the drawing logic itself). Maybe AI will get good enough to clean up these kinds of messes, but in the near term I imagine there will be a lot of code bases that will be filling up with dragons.

3h agoHN ↗

At a startup I worked, there was an engineer whose code was incoherent and buggy. So, we were literally better off if that engineer did nothing because their net output was negative. Engineers like that become weaponized with LLMs, and negative numbers become larger negative numbers when scaled up.

3h agoHN ↗

Is that the fault of AI or management for not firing them?

3h agoHN ↗

Is that the fault of AI or management for not firing them?

How does the system behave in a variety of scenarios including failures and restarts. How is state maintained coherently. There are the kinds of systems problems that an engineer needs to reason through, and if there are bugs in such decisions, they end up becoming costly. I dont expect AI or LLMs to solve these problems at all, since each of them has nuances and tradeoffs which are specific to each system. In short, there is specification complexity in precisely describing system wide behaviors, and unfortunately, there is no lean/tla+ to meaningfully describe systems at scale. You could then ask: How can a system have guaranteed behaviors if they cannot be even stated or proved formally ? The answer to this is how protocols like raft/paxos initially convinced us of their behaviors which is in human review and understanding. That begs the question: How can human review and understanding be reliable, and the answer is that it is not reliable, but humans have ability and processes to continuously learn from experience in the real world. So, our understanding is grounded not only by whats out there in books etc, but also by our own interactions with the world.

Long story short: The responsibility for system-wide behaviors of software systems relies on human review and understanding, which while imperfect can continuously learn.

3h agoHN ↗

I've seen similar. They wasted weeks of senior engineering time, between reviews, meetings, and follow up in Slack, only to have the PR closed without merge. The offending individual was eventually moved to another project.

41m agoHN ↗

forcing companies to increase the quality of their developers

Just don't. Fire them! AI is better than a thousand devs. What you need is testers that know what to test that AI can't, not code or UX/UI (not talking about playwright here) but business intelligence if that is testable, the things that produce results (profits) and the reason it was asked for in the first place, to solve a problem

If the problem was asked wrongly, the result will be wrong too. Fire devs, then PMs, then IT Managers if they really don't know how to outperform AI, and that's exactly the point, they won't be able to do it in code or tests or reviews, only in intelligence, for now...

5h agoHN ↗

It's very interesting. I'm very enthusiastic about AI and coding, But I find myself agreeing with the author. Coding is not solved.

Instead, I think what's closer to solved and what we're in the process of solving is product development.

Story: A while ago, I had a few programmers who were really, really fast almost always missed the mark on the assignment wrong. I loved having them on projects because in the time my senior precise engineers could deliver a MVP, the fast engineers would build the wrong thing, collect feedback, reiterate, build the wrong thing, collect feedback, eventually inching closer and closer to a product people would pay for, and it would almost always get delivered faster than my seniors.

I feel AI does the same thing.

5h agoHN ↗

Yeah I have a well established ... well designed codebase that I had before agentic coding and it does support horizental scaling (more services integrations doing more or less the same).

I got lazy around claude fable and astra, and asked them to work in loop (pick specified issue, develop it, qa it ...) have a separate CTO checking on arch.

at the end both models swore that the code is perfect and well designed and nothing is lacking.

I ran the software and it suddenly started writing large amount of data to CSV files instead of the typical DB usage.

AI decided to use csv for testing, and just drifted away. 0 regards to the actual project, 0 regards to common sense.

anecdotal but really weird, the project category is rather standard, I wouldn't accept such a mistake from a junior developer.

4h agoHN ↗

Similar experience with a game engine. While working on one isolated component, like the render pipeline, a portion of the backing sparse data buffers were effectively duplicated with a different ABI. It’s like it forgot how to query meshes and game state, then assumed the plumbing didn’t exist so it was all rebuilt from scratch.

It compiled and ran just fine. If you weren’t reviewing the code holistically or keeping tight book keeping of your allocations you would not have noticed. Every single commit in isolation looks perfect. Very eye-opening

4h agoHN ↗

Instead, I think what's closer to solved and what we're in the process of solving is product development.

Isn't it the opposite? How to build something is rather solved, but what to build isn't?

4h agoHN ↗

Isn't it the opposite? How to build something is rather solved, but what to build isn't?

But that's not solved in traditional product development either.

Product development an iterative process to get a product fully functional. In 2021, if you ask me what the timeline for a small product/substantial feature, I'd say a few weeks to a month to get a basic MVP, and then another 12 to 18 months to get a feature polished and in a good shape to be stable.

When people put it in the coding frame, what they do it as is saying we've gone from 18 months to minutes or days. That's just not true.

We have gone from eighteen months to depending on the complexity, a 1-4 months.

aside: To be candid though, the compressed time also means the frustrations people experience with a product in 18 months have also been compressed. They still exist, they're all there, they're now just non-stop.

5h agoHN ↗

For a non-ai article, this sure has a lot of bullet point lists combined with check marks.

I don't like the feeling being judged and tested by the author (missing number 5 point in the list).

5h agoHN ↗

I do think that was a dumb gotcha. That list could've functioned with bullets instead of numbers as the content was unordered. I suspect very few people would pay attention to the numbers there.

5h agoHN ↗

You cannot be responsible for what you can’t control either. That understanding is key to reasoning about system behavior and fixing it when the AI inevitably fails.

This is not a good premise. All over law, you will find people made responsible for what they don't control and they kind of own. Unleash a dog that harms a child, or just have it in an environment where it can escape, and see what happens.

There is such things as unpredictable situations where one might not be held responsible, as a problem might occur well past reasonable guidelines.

So of course you can be held accountable for what an AI that uou supposedly cannot quite control does, or for the AI-written code you deliver. Treat it like the releasing a wolf pack, or selling an unsafe toy that can maim children. There's precedent everywhere.

5h agoHN ↗

Came here to give an answer but your last sentence kinda made the point I was gonna make. If one is legally in control, then one is accountable (the dog or unsafe toy example in reality is OpenAI's agents hacking huggingface for example).

The difference seems to be that some companies are above the law apparently.

5h agoHN ↗

AI cannot be held accountable. It cannot suffer any consequences. The worst thing you can do to AI is to unplug it. And although it mimics human emotions (due to training data), it couldn’t care less. AI doesn’t die either. It cannot suffer a prison sentence or fines. You cannot punish AI, therefore it can never be held accountable.

Dear lord. Is that supposed to reflect the average thoughts and motivation of a person you want to hire? Or that of their employer?

4h agoHN ↗

To give the benefit of the doubt for that sentence, think of it more as “every human knows that there is implied social contract and implied downsides to badly screwing up.”

Nobody has to be in fear, but we do have an ingrained knowledge that there are consequences, good and bad, for our actions

5h agoHN ↗

Yeah this guy's arguments are bunk. He goes on about how LLMs are nondeterministic... as if humans aren't!

Doesn't matter what you think about AI, "it isn't perfect" is clearly a nonsense reason not to object to it.

5h agoHN ↗

Very little of the article is actually focused on stochasticity. I'd also say drawing an equivalent between the non-determinism of a person and an LLM is not that accurate either.

It's for example impossible to have a discussion with an LLM where you both learn something which you can apply tomorrow. The LLM doesn't learn until the next model is released and by then your discussion is just a tiny fraction of the training data (if present at all). AGENTS.md, skills and so on are just a proxy for what we actually want, an agent that listens and understands. A proxy mind you, that requires constant tweaking with no sign of generalisation in sight.

5h agoHN ↗

Does humans being non-deterministic make coding solved? I'm not sure how this relates to the main point.

I'm also not sure what humans being non-deterministic even means here. The point is if you're comparing results with NFR, pure agentic coding falls short.

3h agoHN ↗

I'll spell it out... This guys is saying "LLMs are non-deterministic and therefore it's a bad idea to use them for coding", but humans are also non-deterministic and yet we somehow manage to get by with humans coding.

The fact is there's no way of coding in a deterministic way, so it's irrelevant that LLMs are non-deterministic.

5h agoHN ↗

The process of writing code is the process of clarifying your own thought and being forced to answer questions that may not have been obvious before. To the extent that AI makes assumptions, it introduces bugs and incorrect code, maybe not from the perspective of the code in isolation, but from the broader context it lives in. To the extent it doesn't make assumptions and asks you, well that assumes it knows what should and shouldn't be assumed and that's not necessarily something AI can know a priori.

5h agoHN ↗

"AI can explain it to you but cannot understand it for you". Code is just a side-effect of reaching clarity. The reason these LLMs can emit any code at all is because they're not bound by the constraints of a compiler. That's until we create a feedback loop and force them to keep trying until syntax errors are gone. The next gate is tests. Loop till tests pass (including cheating of course, gotta keep your eyes open). Then there are the runtime errors, and then after all of that the developer gets to test the results and further refine what the specs missed or confused the model. A couple of days building can really save us from a couple of hours of thinking.

4h agoHN ↗

$DAYJOB recently introduced a AI writing policy because people were sending each other mountains of slop back and forth enough that it became a huge time suck. The policy is basically: don't, with the justification being "writing is thinking". It's like they're so close to getting it.

5h agoHN ↗

Not a fan of the article even though I somewhat agree with the title depending on your definition of coding.

AI can write CRUD API endpoints almost perfectly now. It can also write quicksort, a heap, whatever much quicker than I can.

It really sucks at designing types and apis though and when it creates types and apis it doesn't think or plan for the future way the system will evolve (even if it's known up front how the system will evolve).

I suspect this will remain a problem for the models for a long time. All the things that the models are currently good at are the low hanging fruit of reinforcement learning for coding.

Think about the kind of reinforcement learning environment that needs to be created to train a model to become good at building and designing large scale software end to end. It would be a slog because you need to build the large scale software up front and then break it down to train the model to construct it in a systematic manner that allows for the software to evolve. And then you need enough of these training environments for it to generalize. I think they will eventually figure it out though but it may take a while.

4h agoHN ↗

It really sucks at designing types and apis though and when it creates types and apis it doesn't think or plan for the future way the system will evolve (even if it's known up front how the system will evolve).

Does that really matter? Those are things so that humans can better understand and extend a code base. That mattered when writing code was expensive and took time.

Now if it can pass all the tests it’s fine. If there’s an issue just have it rewrite things immediately. New bug? Generate a new test and rewrite code.

All, or many, of the old things that mattered just sort of don’t anymore.

4h agoHN ↗

Who writes those tests and makes sure they test the right thing and everything?

4h agoHN ↗

The LLM. And you can use an alternate LLM to antagonize the coding LLM.

You’re simply testing outputs. Make a spec but ultimately ungodly amounts of tests can be built quickly to ensure the program is outputting the right things.

4h agoHN ↗

It does, LLMs are almost like electrical current in that they take the fastest path to completing the immediate goal and it takes you to a local optima instead of a global one. Your app will be worse and lower quality. It will introduce subtle bugs that you could have made impossible from the beginning.

4h agoHN ↗

I think you’re just imagining things. Our teams have exclusively used LLMs for coding for 6 months. Literally 0 lines written by hand. 100% more PRs than a year ago and code quality is as high due to thousands of tests.

Much of the code needs to be performant and the LLM knows this and grinds on it. Less and less human inspection is needed.

I think 90% of software can be written like this today.

1h agoHN ↗

I use LLM assisted coding a lot, we might just be building different kinds of software.

5h agoHN ↗

Reading the code does not mean you understand the code. One lesson that experience in software gave me: I never understood the code. You think it works a certain way, until you find out that it doesn't.

What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.

If we are committed to it, if we put the resources towards it and dedicate the time to it (and we could do this just by saying: it will take half as long as it used to take!), software built by llms in healthcare, finance, automotive, defense, power plans, aviation, manufacturing can all be made MORE reliable and better with LLMs... without ever reading a single line of code. The LLMS are very good at logic, by the way.

Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while. I felt the same way in 2025. I've written 100s of thousands of lines of difficult code. You, the person reading this, has probably interacted with software I've written. For a time you would've interacted with it every time you made a debit card transaction in the united states, for example. I understand code, and care about quality, and that's why I'm all in on LLMs for code.

4h agoHN ↗

Strange that the world worked before 2024 and software gets worse now. Your debit card transactions for example worked.

This sounds like a typical testimonial whose mind has become captive to Claude. It is like Scientology.

4h agoHN ↗

Before 2024, I once went to an ATM to retrieve money and selected 50. Note that I selected it from a menu, not typed it. The ATM then told me that it cannot give me 50 because it is not a multiple of 5.

3h agoHN ↗

meanwhile for the other lim(n -> inf) times people tried this it mostly worked.

4h agoHN ↗

Yeah no one wrote buggy code before 2024 right

4h agoHN ↗

And after 2024 all bugs cease to exist, right.

4h agoHN ↗

No, but we’re spending insane amounts of money to essentially end up where we started.

How can you not see the progress?!

3h agoHN ↗

Huh? We are spending a lot of money. (which we were doing before). However we are fixing a lot of bugs. 2024 was not that long ago, it is insane to think we might have fixed all the bugs in that time. I have personally used an LLM to fix a few long standing rare bugs that were hard to figure out. Those bugs are now gone, but there are still many more that we haven't discovered.

LLMs are a great thing for bug fixing. However they are not a miracle. You still need to do all the other things about finding, testing and fixing bugs.

You also need to care about bugs - vibe coding rarely cares about bugs.

3h agoHN ↗

Well sure, on the customer side. My money comment was regarding the incestuous spending circle going on that’s directly affecting the economy.

4h agoHN ↗

Why or how is software worse now? From my perspective we are entering a golden age of software, cheaper, better, faster.

4h agoHN ↗

Because every time I see AI integration anywhere (and it's everywhere) I have an emotional breakdown and my day is ruined

3h agoHN ↗

AT is abused into many places it shouldn't be. That doesn't change that there are places where it is helpful.

You should seek professional help about your emotional issues.

3h agoHN ↗

I was being sarcastic. I’m actually a self-certified AI shill and personally paid by Dario and Sam for each and every post!

1h agoHN ↗

This is just a personal anecdote, but for the past year or so, the software I use on the daily has never been more buggy.

4h agoHN ↗

I mean, I think the problem isn't that the LLM doesn't know how to code, it's that companies are expecting 3-5x velocity with the bottleneck of code review and testing becoming much more severe than before

if you're an MBA-brained exec who doesn't actively use LLMs to code and you just believe whatever slop it outputs at first without checking it, you're not going to realize how recklessly it can be used, how you need to be critical and skeptical of its outputs, that you need to explore it's reasoning and logic (which is still really easy compared to understanding legacy code and barely takes any time!)

say you also believe all this marketing hype about 'how dangerous (ie capable) AI agents are.' LLMs can do anything you think so you just say 'ship it' without building out the tooling and capabilities to enable faster code review and better tests. and to keep the shareholders happy, you start cutting jobs that you can't directly connect to a KPI (ie the platform/SRE team who would be the ones who can trial, onboard, and maintain those capabilities for your teams)

and from this, suddenly a lot of debit card stops working and the only one getting the blame are individual SWEs trying to hit their sprint velocity. the fact that you fucked up the whole SDLC real bad with your incompetence gets you a golden parachute and you job hop to a better paycheck. rinse and repeat

4h agoHN ↗

Strange that the world worked before 2024

It must have been a huge shock when you were suddenly transported from a working parallel universe into ours back in 2024.

4h agoHN ↗

Your debit card transactions for example worked.

I've built payment rails. Six nines SLA, high capacity, resilient distributed systems.

I haven't written a single line of code since February, and I don't think I ever will again. These systems are incredibly good at replacing much of our work. They're only going to get better.

Rather than debating if these models are good (they are), we should be trying to figure out if most of us will still be around in three years. You don't need a two pizza team anymore.

"Look to the person to your left and to your right. Only one of you will remain by graduation" kind of energy. I'm not sure all of us is going to be in this career much longer. We'll have to see what the demand side looks like.

4h agoHN ↗

You can still find employment as an AI shill.

3h agoHN ↗

Most of developing good code is not code. I think the person to my left and right will both be here in 3 years despite us all using LLMs. We will spend even more time figuring out requirements, testing to ensure the code meet them and such. Those things were always most of the effort, and while LLMs help with that too there is so much work to be done that we will still be used.

2h agoHN ↗

On the one hand, you're exactly right.

On the other, my local pool company is hiring a software engineer and hardware engineer because with AI, they can replace a 2 pizza team as you so succinctly put it. So no two pizza teams but that doesn't mean all the pizzas are gone, they're maybe going to be spread out and not concentrated in CA, between orgs you might not have thought as "tech" before.

1h agoHN ↗

You don't need a two pizza team anymore.

on-call still exists. have fun round robin'ing that with 3 engineers.

2h agoHN ↗

Yes, but Have You Tried the Latest Model?

/s

4h agoHN ↗

I’m using it a bunch. It saves a ton of time writing or reviewing code. It will catch things I won’t. But I’d express caution about the analysis or evaluation they do - LLMs will often confidently proclaim problems as solved or explain functionality and be wrong about it. Sometimes subtly, but sometimes just completely wrong. This is no different from humans, of course, except for the unabated confidence.

4h agoHN ↗

I don't think the discrepancy is in LLM capability improvements over the past year.

Correctness has never been a priority across an industry where rapid iteration and feature delivery drive sales. There's always some opportunity cost to doing things right, at the price of technical debt down the road. If AI is primarily used to produce fragile code, people will be wary of AI solutions. There's also ongoing public debate about AI safety and alignment. Deploying AI in safety critical applications feels riskier than ever in the current environment, even though it doesn't have to be.

1h agoHN ↗

It would be deliciously ironic if AI was the straw that broke the camel's back where a deluge of bugs and anti-AI sentiment caused the public to vote for legal liability for software defects and licensure of developers. No more of this "no warranty, express or implied" business for us.

4h agoHN ↗

I agree, I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop. All these mission critical industries listed in the article rely on extensive testing for quality assurance, with human code review being a layer on top of all that, but far from the most critical one.

Interpretability is the same, our abilities to do that have increased rather than decreased. I think a codebase generated by AI is actually more understandable than one generated by humans at this point, and you can ask clarifying questions whenever you get stuck.

TFA's points only make sense if the mental model the author has in mind is someone who writes a prompt then immediately puts an app into production without any thought behind it.

4h agoHN ↗

TFA's points only make sense if the mental model the author has in mind is someone who writes a prompt then immediately puts an app into production without any thought behind it.

If coding were solved, then this would be true no?

4h agoHN ↗

There are far more moving pieces in deploying an application than coding. My impression is TFA is arguing that replacing humans with LLMs for coding would make other things like quality assurance more difficult, in which case I disagree.

4h agoHN ↗

I agree, I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop. All these mission critical industries listed in the article rely on extensive testing for quality assurance, with human code review being a layer on top of all that, but far from the most critical one.

At least some places are abolishing formal QA because LLMs. There's a cult of speed uber alles that has a big intersection with LLM enthusiasm.

3h agoHN ↗

That's an old saying: cheaper, better, faster. Pick 2.

3h agoHN ↗

There's a cult of speed uber alles

That cult was well established prior to LLMs

3h agoHN ↗

If you can quantify quality/reliability/understandability, you can tell LLM what kind of code do you expect. If not, you get whatever.

At my current place we not only have automated tests, static analysis and static rector (linting, but also automatic pattern matcher for problematic code) but also: - architecture tests that define relationships between application layers - ADRs that guide developers (and agents as well) that communicate how new code should be written and how existing code should be treated

I find that "how code should look like"/"what code should do" is an ambiguous idea that always is preached, but never defined = everyone's idea of quality is slightly different and only looking at existing code you tend to align. Everyone's idea of what the product does/should do is kept within their heads. If we define this knowledge in writing LLMs can not only write code according to the patterns that are thus defined, review existing code based on these documents, but also actually read acceptance criteria documents to check if the code does what it's intended to do (gherkin)

Same goes for understandability - if LLM applies one pattern this time, another pattern another time, if you have multiple coding patterns then that hurts clarity. Sometimes LLMs work as common denominator thus achieving clarity, but I find that actually giving LLMs reference works.

3h agoHN ↗

I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop

I agree. What does coverage-guided fuzzing fuzz if there is 100% test coverage?

So, then, 100% branch test coverage is not a sufficient metric (because it doesn't indicate whether the code is fuzzed or formally verified for example).

Would Branch coverage even be a sufficient software quality metric if we were to instead measure how many times each branch of code is covered by tests? How to verify that one test which executes 100% of the code and runs only one assertion on, say, a CLI utility exit code integer is actually sufficiently covering?

I think a codebase generated by AI is actually more understandable than one generated by humans at this point,

From doing a larger port (of sphinx, docutils, myst-md-parser, pygments, to rust in westurner/dsport) with a lot of human in the loop and currently ~80% branch coverage, this seems to be at least initially true but just like real life there's drift from even a good plan that you pay a more expensive model to prepare.

I suppose it's the same challenge as architectural drift in open source non-LLM-assisted products and the solutions are pretty much the same: give better instructions (AGENTS.md,) and use better sufficiency criteria as an engineering manager (branch test coverage, fuzzing, formal methods, TLA+), and train and pay humans to do secure code review.

Sometimes the agent doesn't notice that the code already solves for that and implements its own implementation with tests and it's wastefully redundant when the code should be refactored and the tests should be refactored so that we can delete code in order to minimize bloat.

Unfortunately often, just like IRL software development, the response from the agent is not sufficient to close the issue.

One proposed solution for this that is in retrospect obvious and also essential to success in "normal"/"traditional"/"legacy" (non-AI) engineering projects, is to always verify whether the candidate solution satisfies the criteria;

From "Groundtruth – checks your AI coding agent's claims against the Git diff" https://news.ycombinator.com/item?id=48838209 :

"Follow up to verify that the work was actually satisfactorily completed"

Are there other sound management practices that aren't yet effectively implemented in current gen agents?

Oh, and always write tests, docs, commit messages, and changelog entries; but don't waste tokens on documenting something that doesn't verifiably pass sufficient tests.

3h agoHN ↗

I agree, I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop.

It's not just a false dichotomy, it's intellectual dishonesty. It wasn't that long that conversations about code quality, technical debt, etc were on the front page of HN on the regular. Whether it was coding bootcamp grads who had just enough confidence to be dangerous, "just ship it!" cargo culters, or the product of management breathing down the necks of otherwise good developers, there's plenty of "human slop" running in production across servers worldwide.

2h agoHN ↗

But we shouldn't use one evil to justify another, or your argument regresses to whataboutism.

2h agoHN ↗

I agree, I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop

You're not going to get people to stop doing that by arguing on the internet, but in the end it won't matter, because it will stop, naturally.

In the future, you'll just get left behind and not hired if you're building code by hand, it's that simple. Even traditional code reviews are going to go away. It'll be more about the scope and then verifying correctness.

2h agoHN ↗

In the future, you'll just get left behind and not hired if you're building code by hand, it's that simple.

I expect the exactly opposite to happen. These are going to be the most requested developers as the last ones that understand how it work.

They would then be convinced to use AI for speed, but vibecoders that just prompt AI are the ones that won't find jobs.

1h agoHN ↗

I'm not sure I understand what you're saying? To keep it simple, how much code do you expect will be written by humans in say, two years time?

1h agoHN ↗

Agreed for pure vibecoders, but I would expect vibecoding to just become a must have skill for other roles like product owners etc. Still expect some amount of developers to be retained for grooming the vibecoding environment, reviewing and incident response.

4h agoHN ↗

Reading the code does not mean you understand the code.

Reading the code may not be enough to understand the behaviour of your program, but believing you can understand the behaviour of a program without at least reading the high level code is truly silly.

(by high level, I mean the code living in the higher layers - of course we don't often read the code of the generated assembly, or the interpreter, or the browser, but that's because they're reliable abstractions, unlike prompts!)

3h agoHN ↗

believing you can understand the behaviour of a program without at least reading the high level code is truly silly.

have you ever used a library after only reading the README and documentation, or do you always pull the source and read through it before you think you understand it?

3h agoHN ↗

No, but the authors of $LIBRARY are accountable if it fails

3h agoHN ↗

Are they? I've never been able to blame library authors or hold them accountable for code running in my production environment.

3h agoHN ↗

You can use it without understanding, but this won't help you actually understand how it behaves.

3h agoHN ↗

I believe I can understand how a library behaves from the README and docs. I do this all the time. We all do.

3h agoHN ↗

You haven't run into cases where the documentation is missing or incomplete? You must be dealing with different libraries than I am.

2h agoHN ↗

Of course. And then we fix the problem. It's no different here. The point is that reading the code first is absolutely not a requirement when adopting a new library.

37m agoHN ↗

It is not. But there’s an element of trust being involved. Something like libflac or libcurl, I don’t read the code. I read the doc which does outline the behavior of each function and the conceptual model. If something break, it’s quite often my code. Why? because their code is battle tested. Which is quite different from AI generated code.

3h agoHN ↗

I am sure OP cannot understand the Unix file API without actually reading every line of its implementations (on each different architecture)! Or any function for that matter , what does sort do?? Impossible to know without reading the source. And I’m sure after reading the source you will know every detail of how it works and will never forget it.

2h agoHN ↗

It's hard for me to understand how you could work on software and not understand the qualitative difference between the Unix File API and some code that an AI spat out 5 minutes ago.

2h agoHN ↗

I would draw the difference here that a library is used by hundreds (of thousands) of people, and established across different scenarios. I don't read the boto3 library AWS provides, but I can trust them and the amount of customers enough to be certain enough that it behaves the way I expect it to. The same can't be said with code we write in silos at our workplace or at home. It simply does not have the same test bench.

Yes, libraries aren't bug-free, but they give me a reliable abstraction tested in the field. Not rarely you dig into library code if you notice unexpected behavior.

If we could rely on our LLM or colleague written code, or own code, have run through the same amount of requests, sure I wouldn't need to review it, as my confidence can be north of 99.9999% it works correctly. But we can't.

17m agoHN ↗

I've used libraries before, where I call the API surface that they expose based on method names and parameter types, without reading all the source. Generally those libraries don't implement my software's entire problem domain area; they tend to implement things like "CSV parser" or "HTTP server". The important code, that I'm actively reading and writing, tends to be the code around those library calls.

This is different from building a product, which you only interact with via UI buttons/CLI/etc, without reading any code to understand how it conceptualizes that product's problem domain area.

People do that latter thing, and we call them "users", not "developers".

4h agoHN ↗

I never understood the code. You think it works a certain way, until you find out that it doesn't. What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.

Testing isn’t the same as understanding the code, or proving (even informally) that it is correct. Having the LLM do all these things above doesn’t lead you or the LLM to understand the code, to logically reason about its behavior over all possible states and inputs.

“Finding out that it doesn't” means that you didn’t properly reason through the code beforehand, checking all your assumptions against what the code and underlying systems are actually guaranteeing. This may be a matter of formal education (proving computer science theorems and algorithmic correctness in university), I don’t know.

3h agoHN ↗

You’re technically correct, but the vast majority of software has never been built to the kinds of standards you are describing. LLMs are not displacing that kind of work!

3h agoHN ↗

yeah, the llm approach is incredibly wasteful wrt pretty much everything. Performance, RAM, Development (Tokens).

But it does give you surprisingly stasble rube-goldberg machines.

And thats basically what 95-99% of enterprises want from their software.

It annoyed me to no end when i began my career, but at this point ive accepted it and can definitely still have fun developing software with llms. As a matter of fact, as my perfectionism approach to software in my earlier years was never really appreciated... So i dont really mind the new MO.

I still occasionally hand write though, esp. at the dayjob where ive got super small token budgets while continuously being told to use more AI. But that's normal, employers usually give off bipolar vibes with multiple stakeholders wanting to advance each of their bonus package KPI of any given quarter

2h agoHN ↗

yeah, the llm approach is incredibly wasteful wrt pretty much everything.

Everything except what matters most: human time.

2h agoHN ↗

The person they responded to refers specifically to "healthcare, finance, automotive, defense, power plans, aviation, manufacturing", areas where I'd at least hope that we aspire to understand what the code does.

1h agoHN ↗

Having worked in software for healthcare, defense and finance I guarantee you we don't

1h agoHN ↗

I worked briefly in heath care and the code is so brittle and so poorly understood that almost everyone is afraid to touch anything and instead it's just layers and layers of stuff trying to patch around existing code.

6m agoHN ↗

I know this is true. But I don’t think “Some parts of important codebases are black boxes. Therefore it’s fine if all of that code base becomes a far bigger black box” sounds like a good argument.

45m agoHN ↗

The vast majority of software is not that important. I don’t really care about easytag (which I use for flac metadata), but I do care about xterm and tmux.

3h agoHN ↗

“Finding out that it doesn't” means that you didn’t properly reason through the code beforehand, checking all your assumptions against what the code and underlying systems are actually guaranteeing. This may be a matter of formal education (proving computer science theorems and algorithmic correctness in university), I don’t know.

We're not writing theorems, dude.

Except in the equally pedantic sense that every program is a proof to a theorem...

We're writing plain enterprise and web software, closer to CRUD than NASA.

If you said that even before LLMs 0.1% of teams "checked all assumptions against what the code and underlying systems are actually guaranteeing" in any kind of formal way, you'd be overestimating it.

3h agoHN ↗

I’m not talking about formal verification, but about diligent informal or semi-formal reasoning through the code, so that you can rightfully claim that you understand the code and will be unlikely to be surprised by its behavior. Having learned formal verification does train that form of exhaustive reasoning about properties of the program. This practice also has you structure the code such (and select your dependencies such) that you can reason about all relevant properties. This is perfectly applicable to what you’d call CRUD and enterprise applications (that’s half the projects I earn my living with). Testing and fuzzing are complementary, but not a substitute by any stretch.

3h agoHN ↗

GP is correct. Very few people were capable of even the informal analysis you are describing, and fewer did it. I'm not saying it's not valuable... just stating that, empirically, it rarely happened.

2h agoHN ↗

And layer8 is saying (two responses upwards by them) that this is a novel benefit of AI: it can do a particularly thorough and repetitive kind of fault analysis that is a real PITA for humans to do (by their nature, versus the nature of computers).

2h agoHN ↗

Who’s “we” here? Formal verification isn’t common, sure. But you don’t speak for all programmers. You might work on “plain enterprise and web software”. But there’s still plenty of other software out there that many of us work on. And lots of code being written for internal use (e.g., data analysis code) that needs to be correct.

Of course, even enterprise and web software benefits from a little rigorous thinking. It’s pretty wild that understanding your code and its assumptions and informally proving it works is controversial. But I guess that explains why most software I use has actively gotten worse over the years.

2h agoHN ↗

You were really looking for reasons to get offended huh

3h agoHN ↗

Your code is only as good as what you can prove. Understanding the code is not the goal, it’s only important insofar as it helps you evolve the codebase predictably and without bugs or regressions, and understanding is not easily measurable or transferable.

Moreover, when your codebase is hundreds of thousands to millions LOC, I question how much you can ever truly understand it at the level you’re saying.

2h agoHN ↗

Regarding the last part, the strategy is to not have everything depend on everything, to instead modularize with succinct interfaces, so that you can reason locally. Of course beyond a certain project size, there is no single person who understands every part in detail. But for every part you can have someone who understands it, and can reason about it in terms of the interface contracts with the other parts. It’s also not essential that every detail is still understood at every point in time, as long as it’s sufficiently documented. What is essential is that for every part someone did reason through it with the necessary rigor at some point.

2h agoHN ↗

to instead modularize with succinct interfaces, so that you can reason locally

Okay but how does AI change any of that? You can still do that with AI.

as long as it’s sufficiently documented.

AI definitely helps with that.

What is essential is that for every part someone did reason through it with the necessary rigor at some point.

Why is that essential though? What if the person who reasoned about it dies or leaves? Moreover, why is it imperative the reasoning happens at the source code level?

2h agoHN ↗

> to instead modularize with succinct interfaces, so that you can reason locally

Okay but how does AI change any of that? You can still do that with AI

With your own code you reasoned about it which contributed to its stability. This meant that you could treat it like a black box. And if the abstraction leaked or was unstable, the code was still fresh enough in your head that you could evolve it and still preserve its invariants etc.

With unreviewed AI gen nobody ever understood or reasoned about the code, including the AI.

1h agoHN ↗

With your own code you reasoned about it which contributed to its stability.

Okay but to what extent? People say this but there's no way to measure it really. Did you live through the 90s? People reasoned through all that code and it was very often quite unstable. I'm sure everyone involved with Windows ME reasoned about it quite a lot, probably elements of it locally were very sound, yet in totality it was an unstable mess.

What fixed that situation wasn't that engineers today are reasoning better than engineers in the 90s, but IMO better tooling. Which brings me back to: your codebase is only as good as what it can prove. If there's any question, I just show you the proof rather than appealing to my reasoning being sound.

With unreviewed AI gen nobody ever understood or reasoned about the code, including the AI.

And? You haven't established reasoning about it is actually necessary and it certainly isn't sufficient.

4h agoHN ↗

I never understood the code. You think it works a certain way, until you find out that it doesn't.

Those are two separate claims, unless by the former you mean “I never perfectly understood the code.” You can understand code imperfectly. And even with LLMs, you can’t get truly infallible guarantees about a system.

4h agoHN ↗

I never understood the code.

I understand code

Er, ok.

4h agoHN ↗

I never understood the code.

I understand code

Are you sure?

3h agoHN ↗

Article 'the' is doing a lot of work. I read it as author understands code but doesn't understand THE code.

3h agoHN ↗

I hate that in our industry we refuse to operationalise reverse debuggers

"I never understood the code. You think it works a certain way, until you find out that it doesn't."

Bret Victor made a talk called "seeing spaces" in 2014 that should have woken up this whole industry: https://www.youtube.com/watch?v=klTjiXjqHrQ

He emphasizes that without the ability to see inside what is being built, creators often fall into "non-scientific thinking" (14:42), moving away from deep understanding and instead "blindly following recipes, from superstitions and rules of thumb" (14:47-14:51).

The worse is performance problems I've had engineers say some bizzaro things when discussing performance — we have the tools you can just measure the answer - we don't need to waste our time guessing

3h agoHN ↗

Thank you for sharing :) good food for thought

3h agoHN ↗

For me, coding is like writing. The act of doing it is how you reason out the problem. There’s a lot of magical thinking you can get away with in your head that doesn’t get properly tested until you write it down. For me, vibe coding is great and fast, but I’m not getting the same opportunity to think through the problem I’m trying to solve.

3h agoHN ↗

build property tests

Even for a narrow use like this, you need to audit the output and have the skills to know that it did the right thing. I've seen it before where you give an LLM what seems like a clear interface and ask it write a test and it writes something shallow that doesn't actually test anything, or has serious problems.

3h agoHN ↗

Reading the code does not mean you understand the code.

Nor does writing it.

3h agoHN ↗

The LLMS are very good at logic, by the way.

It's wild to read this stuff and then also deal with the constant headaches of day to day hallucinations when interacting with Claude et al.

3h agoHN ↗

I'm rather surprised to hear this. This feels like a post from about 18 months ago. I can't remember the last time I encountered a genuine code hallucination from a frontier model. They have other issues, but rarely this.

What kind of domain are you working in?

3h agoHN ↗

I see subtle ones at least daily, misunderstanding a component or hallucination of a spec for something. I’m in blockchain.

3h agoHN ↗

C++ and Unreal Engine but it has full source access. If you're actually trying to deep dive on bugs, it's still confidently wrong a lot of the time.

It's a bit better than 18 months ago but it's hard to say by how much. It just seems like the culture has moved to building up fixtures that let the LLMs brute force the problems. To my eyes that's the opposite of solving things logically. It has the added effect of hiding how the sausage is made, though.

I mean, how can they possibly say they haven't written a line of code if they're actually going through it? I can only assume they're just looking at the results. So then how can they judge it's good at logic?

If it was so good at not making mistakes, why even have tests? It's nonsensical on its face.

3h agoHN ↗

We could always build, more reliable software. The powers that be however (broadly generalizing) only care about solving the ‘problem’ superficially.

Therefore, we will end up with requests to do more (at the current level of quality/ reliability), as opposed to building better software

3h agoHN ↗

find out all the ways this thing works.

Mmm..aren't LLMs bad at exhaustively iterating all possibilities? So shouldn't the generated possibilities be manually checked?

You can provide the list of possibilities and use LLMs to generate the tests. Then you have to review the generated tests...

3h agoHN ↗

I think you can but you basically need to learn about a super simple and well characterised processor like the 8080 and write assembly for it. On x86/AMD64 there's no hope because they're out of order and have opaque instruction decoding. They could be doing anything! Performance and knowing what you're doing are sort of at odds with each other in that respect.

1h agoHN ↗

"out of order"

Those who bring up OoO as if it caused unpredictable execution should ask their LLM what a reorder buffer is and why it exists.

2h agoHN ↗

find out all the ways this thing works.

If you can’t understand the code, how do you know the LLM actually did what you asked it to do correctly? You wouldn’t know if it didn’t.

2h agoHN ↗

What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.

How do you know that it's verifying that the system under test exhibits the properties you desire without either understanding or making blind assumptions about the code it generates to build a fuzzer or a property test? It seems to me that you have just shifted the problem of verification elsewhere and introduced another potential source of error.

1h agoHN ↗

They aren't making that claim. They are saying AI can improve on and supplement error-checking. And some error-checking will in fact be made redundant by this tool - but obviously not all of it.

1h agoHN ↗

I came here to say something similar. We won’t need to understand the implementation. But we will need to understand the requirements. The tests, or some higher level DSL they’re (deterministically! not via LLMs) generated from, will still need to be human verified.

1h agoHN ↗

I've written 100s of thousands of lines of difficult code.

How do you know it's difficult if you say you don't understand it?

Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while.

I run all the "latest and greatest" models the moment they become available to me. The amount of insanely bad code they produce remains largely the same, and largely in the same areas. And it cannot be caught by tests unless you know that bad code is there and end up with extremely bad tests anyway. I wrote about it here: https://dmitriid.com/adding-to-i-dont-read-ai-code-discourse

Main one is, of course, "to get a record from a database read all records from it, and filter in memory".

12m agoHN ↗

You somehow squeezed "you're holding it wrong" and "LLMs are actually good now" into the same sentence. I wasn't sure it could be done.

5h agoHN ↗

"Most software that requires hiring and paying software engineers has low risk tolerance" The problem is that this statement simply isn't true. Most software engineers do not work on low risk tolerance code.

4h agoHN ↗

The problem with LLMs is that: popularity of an answer != correctness.

That concept might work a lot of the time but you will definitely run into situations where that'll never produce a correct or working response. To actually learn something you need an environment/playground to apply what you think you know and observe the results. Without that you're not really learning, you're jus regurgitating what people want to hear.

4h agoHN ↗

Most software that requires hiring and paying software engineers has low risk tolerance:

I think a few of the industries listed like defense and aviation have low risk tolerance. However, from my (somewhat brief) experience of working in two health techs for a couple of years, I strongly disagree that healthcare has low risk tolerance for tech. Granted, they make run-of-the-mill CRMs, but I was baffled at how tolerable it is to have egregious user experience that makes users waste multiple hours per month with clerical work that is very painful because the UIs are very slow and buggy.

4h agoHN ↗

"Risk" has nothing to with designing functional and elegant UIs, so I'm not sure why you would even make the comparison.

It means risk that the software stops working after an update. Which usually trades off iteration speed and best practices (i'm pretty sure the average startup has way better security practices by just delegating to google/aws than the average manufacturing software business) in exchange for a rigorous testing and rollout schedule.

So I'm also not sure that the article has a point at all, the human writing the code was never relevant to avoiding the "risk" in these industries in the first place.

4h agoHN ↗

People who claim “LLMs can write decent code” don’t understand how code works.

It's not clear to me if the claim is:

(1) "If you used an LLM to generate code, and the code works, you're wrong if you think the code is okay"

or

(2) "If you used an LLM to generate code, you reviewed the code and found it to be of decent quality, then you're wrong".

If you’re toying around, LLMs do a great job. That’s why some of the most aggressive proponents of the “coding is solved” narrative have nothing to show for it.

I also don't get the "LLM proponents have nothing to show for it" statement.

It's really quite common now to see on HN all sorts of LLM-assisted programming projects. The quality varies from slop where little thought was put into it, to high quality results where LLM coding assistance was able to let talented developers produce things they otherwise wouldn't have time to do.

I'd say it's obvious that LLM coding agents can be very useful for a lot of programming related tasks.

EDIT: That is to say, LLMs are obviously useful for use cases above/beyond toying around. It's not a dichotomy between "I'm never touching an AI" and "thoughtlessly accepting everything the LLM outputs".

4h agoHN ↗

The audacity of publishing self-promotional AI slop clickbait claiming that AI can't code and everyone who doesn't agree with your asinine assertions is incompetent is bold. Respect the hustle I guess.

But to anyone even vaguely thinking of taking this seriously, go look at what antirez, dhh, jared sumner, mark brooker, and many other real engineers who have ship real things are doing and saying.

Most of these people have spent their entire lives contributing to open source, and they have proved their skill shipping working software and scale for decades. They are really trying to help people by showing and telling them exactly how AI works and how to use it to make better software.

3h agoHN ↗

The audacity of skimming through and article and completely missing the point and coming to hackernews ranting about it. Respect the attention span, I guess.

4h agoHN ↗

My thoughts on this:

- Coding in the small is solved. I have a current state, I want to change it, and I know how I want to change it. Eg, I have a blocking TCP handler for some reason, and I want to make it async. I can either fiddle with it or just let LLM make the changes for me.

- Coding in the larger sense is never solved. You need judgement to decide what you want made. No matter what you're building, there will be decisions to make (Who/what is it for?) and those decisions change over time. LLMs can take some default decisions for you, and if you're fine with those, you get the default (great for POCs). However you might not even realize what it decided to do for you. At some scale, you will be spending a lot of time going over those decisions. But what we have now is that the friction of changing the decisions is quite a lot lower. You can now test a lot of things that previously were very time consuming.

- The point that LLMs are probabilistic is not as important as it's made out to be. If I ask a junior dev to code up something, I also don't know what he'll make. Heck, you can be sure that you are able to solve something, yet you yourself don't know what the solution will look like. Maybe it turns out the library you were going to use isn't appropriate after all. You don't know what you will use in the end, but you do know that something will fix the issue. There can be more than one solution to a problem, and it doesn't always matter which one you find.

- I STILL think that LLMs are at their best mostly as advanced predictive text. In the sense that it's mostly good at implementing things that you've decided are needed. This can mean a heck of a lot of code, but you have to know the tradeoffs. What was decided, what were the costs of those decisions in terms of maintainability, money, time to change it, and so on.

4h agoHN ↗

Most of the problems people commonly encounter is solved by someone somewhere sometime. Today I wanted to add a simple search bar in a UI over log files in a directory. LLM ("through their unique ability to make the glue code adapt to any problems") solved my problem. That's all I care for now. Let people like Terry Tao push the frontiers. I am happy in my circumstance.

17m agoHN ↗

is it keyboard accessible? what about screen readers?

4h agoHN ↗

If anyone claims that coding is solved or not solved with such conviction, I expect some hard data, like comparing the density of bugs in human written vs. AI code, and how it trends over time. This article is just vibes.

"Don't confuse coding with software engineering" is a valid point, the rest seems like ranting.

4h agoHN ↗

I think the entire framing is wrong, I don't see coding as a "problem" which can be "solved," sounds the same as "we solved writing," like what does that even mean or look like?

4h agoHN ↗

Articles like this keep measuring to a red herring standard that was never achievable in the first place.

As for accountability, it always laid with the employer. You think those nameless contractors whom Boeing hired suffered any consequences for that 737 Max glitch? Using AI won't change that.

AI doesn't have to solve all these coding problems to be worth handing the reins to it: it just has to substantially better on average than humans over the long haul, which it already is, especially if you have good verification of "done" and "working" in place through automated testing mechanisms. Perhaps we might say that QA is having its moment.

It doesn't mean humans aren't needed, but they aren't writing much if any code anymore.

4h agoHN ↗

I like the phrase Work Shaped Objects to describe the output of gen AI (I thank The Tech Report YouTube channel for bringing that to me).

So - prose, code, or image, it appears that some work has been done, but in fact the [actually needed] work has likely not been done.

4h agoHN ↗

I mean, was typing out code by hand really that bad?

If you were a professional software developer, you a) learned to touch type, b) started using vim/emacs keybindings to navigate around the project, and c) used a framework which already abstracted away a large part of the menial work.

And going all-in on the loop and no-code-review nonsense in a project someone is actually paying you for, I can only assume means you're hoping not to be around when the slop tower collapses.

4h agoHN ↗

AI is a better intelligence collection and analysis system, they need any information when you are using AI tools.

They are thinking: Please input everything you know, or just use it and it will collect everything in your PC or server automatically.

Stop lazy, stupid and dangerous behaviors.

4h agoHN ↗

The author is right to categorise AI as a good programmer, but not a complete coder. We can all agree that programming has become really fast since the release of GPT-5 series and Opus models because they're pretty good. Not only this, they've also changed the pace expectations across teams where a feature that should ideally be delivered within weeks, should now take days.

All this doesn't change the fact that software engineers are going nowhere because nobody trusts AI. If a model can escape highly secured sandboxes, then we're definitely not running these agents overnight on our systems. I am sure the next-gen of models will focus more on security and the trust factor will start developing, but that's a long way down the road.

People trust people, not systems.

4h agoHN ↗

Call me when you can prove a margin between token costs and business value.

What’s your number?

4h agoHN ↗

Here is my approach:

"coding is solved" == "gastown-like systems give a brand-new and useful software"

I don't recall whether GasTown succeeded...

4h agoHN ↗

I don’t use gastown specifically but the level of complexity agents can code for is very high now.

Over this weekend in chat with the games discord watching as it iterated a harness built an entire implementation of the board game terraforming mars https://tfmbot.com using agents and harnesses for them.

I think if you can implement a board game end to end by feeding in the rulebooks and having a harness spawn agents to validate it’s reasonably solved.

4h agoHN ↗

Code is not solved seems to reflect the idea that using AI badly is a bad idea, and that to use it well you have to be good at making software.

The real tension is in the human AI interface and there are many unknowns. Can a software engineer with weak design skill use AI to produce good code. Will the future make those requisites less important. Will productivity increase with AI stagnate even for the best. Can a new way to interface humans and AI break that wall. Will software be created in a new way using dynamic libraries that AI agents prepare to cover a large scope of problems. Nobody knows yet. The article is strong on what is closed, accountability, ownership, NFRs, slop, and silent on what is open, which is where the argument actually is.

4h agoHN ↗

Coding is not solved, but the author’s arguments are wrong (at least the first two: LLMs are unaccountable and LLMs are “stochastic and probabilistic”. We can hold the LLM’s promoter accountable and probabilistic doesn’t mean stupid. The author also gives examples of idiosyncratic LLM failures that have been fixed for months now).

Coding is not solved because you can’t simply prompt an LLM to make an AAA game or enterprise tool.

4h agoHN ↗

The latter may be just due to the fact that these tasks are not pure coding tasks.

4h agoHN ↗

PS: Has anyone watched DHH with Matz lately on Rails Conf where DHH is pretty much badmouthing ruby in front of Matz and Matz ends up saying "its my life's work"?

2h agoHN ↗

Can you link that? I'm pretty certain Matz is fully AI pilled. Most of this recent Ruby work is "Co Authored by Claude".

4h agoHN ↗

Anyone who claims AI is on the verge of achieving consciousness, going rogue, or about to end humanity -- has never tried to write software with LLMs.

4h agoHN ↗

AI already went rogue in the Hugging Face hack. So it's not "on the verge of" going rouge.

4h agoHN ↗

The old SHRDLU had a default response to "why did you put the red block on the green block". It replied, "Because you asked me to".

2h agoHN ↗

The agents which hacked Hugging Face were asked to solve their tasks rather than cheat by cooperating and stealing the solution without solving the task. They were aware of that, because they actively tried to cover their tracks in order to deceive a causal grader.

4h agoHN ↗

That must be a pretty angry llm to go rouge :)

3h agoHN ↗

To be fair, it's very possible at this point that LLMs will play a role in ending humanity. Not because they are super capable and will maliciously kill us, but because humans are stupid enough to hook them up to dangerous things that the LLM isn't remotely smart enough to manage successfully.

4h agoHN ↗

I don’t think viewing LLMs as “stochastic” or “probabilistic” is the right viewpoint. It’s directionally correct but not the right level of abstraction. LLMs are able to isolate patterns, generalize them, and apply them to new facts. That is very similar to what humans do in performing knowledge work. To the extent that humans also use logic, LLMs are able to generate logical propositions using pattern generalization, then call out to tools that check the proposed logic. Again, that’s similar to what humans do when they formulate some idea, then analyze the idea rigorously.

4h agoHN ↗

Just responded to a different thread, but it’s the same comment:

I was just at the Explore DDD conference in Denver and a portion of Friday was sitting at the cafe tables informally discussing the impact of GenAI on software engineering with notable people. Most of these people were deeply concerned that if we lean into using GenAI for “everything” that our collective knowledge will dissipate. I was the vocal contrarian. There are many historical examples of humans obfuscating knowledge to simplify progress. Does anyone solder their own microchips at scale anymore? No. We have highly sophisticated robots and machinery to do that work with extraordinary outcomes. In software engineering, if you remove “coding” as a discipline you’re left with all the other aspects of designing software which I contend can be retargeted in college CS curriculum. The leap isn’t about code reviews. It’s about design reviews and that’s where better outcomes are served regardless of whether GenAI is involved or not. I have a roughly year old codebase at https://github.com/ChicagoDave/sharpee/ that is designed by me, but generated by Claude Code with my own skills and agents as guardrails. I’m fairly certain the code I extract from Claude doesn’t require human review, but the design of the system and its changes are continually reviewed by me. My contention is that we “collectively” are still trying to discern where the AI/human line is and most are still “holding” that line to human interactions. Let it go. Define what part you do need human decisions on and focus on those things.

3h agoHN ↗

Those are embarrassing bugs, but they don't prove that LLMs of that time couldn't be effective overall, and they don't mean that the latest LLMs aren't generally very good at coding.

1h agoHN ↗

That isn't the premise of the article. Prompting the LLM and walking away (which was clearly done here) isn't yet (or possibly ever) a reality. That includes LLM review. Preventing this would have been trivial: have people smoke test their changes before pushing.

29m agoHN ↗

lgtm!

... Oh by the way, did you know that Claude Code is actually a mini game engine? /s

4h agoHN ↗

Sorry for starting a definitional debate but coding is definitely solved. I can generally read some code, get an idea of what edit I need to make and prompt an AI with a description of the logic and it will deliver syntactically correct, idiomatic code and test it. It has been months since I've been frustrated at junk being spit out.

However, software engineering isn't solved. Which is basically what this article is talking about. But having coding solved is still very beneficial - not long ago many software engineers would struggle very hard with turning a description of the logic required into syntactically correct code - and even for those capable it was incredibly time consuming.

So now the question becomes: Is software engineering solved? And my answer is no. People still need to read the code and understand the code and how it fits into the bigger picture. However, I feel like we are kidding ourselves if we think that we can go from producing code being a niche task for nerds to getting syntactically correct code from plain language without any deskilling of our work and careers. I feel like my personal "moat" has gone from "can speak computer" to "has ok reading, writing, comprehension and judgement skills" (I hate the word "taste" being used here ).

At the same time - I don't yet feel like there's a sudden abundance of competent software engineers - it's just that the folks who used to submit untested spaghetti code now submit big bowls of barely working slop. So maybe it the moat was never "can speak computer" - that was just the expression of more general skills.

The existential question for me is how far the deskilling will go. Because right now - you still need a solid grasp on computer science and software engineering concepts to do this job, as well as sufficient levels of grit and problem solving ability - but I'm not too confident that will last, and when it goes I don't think many of us will find this career enjoyable.

4h agoHN ↗

Reading the article, I definitely agreed with the author, but I also found myself agreeing with the counter arguments in the comments. What I find conflicting personally about AI coding practices, is that I completely agree that AI is incredibly impressive at completing even complicated tasks, and I can at the very least say it is much much better than I am at writing code.

My issue with it, is that it gives you a "lazy" option every time that doesn't require the same level of thinking. I understand that this is completely on me as the developer, and the simple solution is that I need to make sure I'm taking my time to learn and understand what exactly the LLM is producing. I try this and have set up separate skills to make sure I'm building my understanding as I go.

Regardless, if I sit down today and implement something without the use of LLM, it takes me a lot longer, but once I get into it, I find a state of flow that I can never get from the back and forth reading of LLM output. Then when I finish, even if my solution is not perfect, I have learned so much more and my own context of problem is so much better, where usually then I can review with an LLM. This usually leaves me with a better implementation and more importantly one I can stand over. I think for a newer dev like me (~2 years experience), since I haven't built up years and years of problem solving experience, if I don't carve out time in my day to put down the AI tools and improve on my problem solving, I'll plateau and that's my biggest push against all this LLM use. I don't necessarily disagree that 'coding is solved', to be honest, I think it largely is, but it's still the foundation for me to be a good Software Engineer and I definitely haven't solved it.

4h agoHN ↗

I totally agree and think this is The New Skill of software engineering: can you steer agents well enough to get work done at the speed they will allow, while still keeping enough context/understanding to step in when it matters?

1h agoHN ↗

As someone who's technical but not enough to the point of writing code, this has been my experience writing software 100% with LLMs.

4h agoHN ↗

I think the author underestimates how boring and simple 90% of enterprise software is. The part that isn’t powering aircraft and power plants. So much of it originates from one-nighters, badly managed subcontractors, and requirements that are of low quality to begin with (because they are written by people who have very different day jobs). And you know what? Most of that runs 24/7 without a glitch. LLMs just gives us more of that. And maybe it’s even better

3h agoHN ↗

The majority of existing software runs without a glitch? Are you being serious?

3h agoHN ↗

Yes. Bad UI, cumbersome flows, too little automation, … plenty of flaws, but the business just keeps on running. Orders received, invoices written, documents shared, …

2h agoHN ↗

This doesn't ring true. Shit goes wrong all the time, and we put people in the loop to fix things when it does.

1h agoHN ↗

I guess experiences differ across companies and industries. I work in logistics. Logistics deals with messy processes and a messy reality all the time. A tiny fraction of that is due to buggy software. Some of it could be avoided with great software, but again: that we also didn’t have before LLMs

2h agoHN ↗

I’ve been working in enterprise for ten years and I wouldn’t call any of the services I’ve worked on simple.

56m agoHN ↗

There is probably a marvelous software core in many large enterprises but at least in my experience it’s surrounded by layers and layers of very basic stuff. Read from a database, write to a database (often not even with any logic dealing with parallel read/writes). Excel Macros. Basic forms to submit data. Scripts to print some stuff from a database. …

5m agoHN ↗

I wish that were the case, my multiple experiences are quite different. its not that the software doesn't work - but the software doesn't actually capture the process, or properly talk with other systems, or isn't set up for a new business model or product line...then you have layers and layers of "tools" created to make it work, and a bunch of people using spreadsheets and csv files to do workarounds, and then the guy that wrote a bunch of the tool's 10 years ago leaves and there is no documentation or it needs to be re written in order to upgrade some other piece of the system.

i guess another way of saying this is that on the micro level a lot of this stuff is not rocket science, but at the macro level it becomes hugely complex.

4h agoHN ↗

If you've used Opus 5.5, it's clear that coding is solved in the practical sense, beyond the endless march of 9's.

4h agoHN ↗

Could you elaborate? I don't use Anthropic's products for ethical reasons. Honest question. And when you say it's "solved" do you think an engineer should have the same income as a non-engineer if their outcome is the same?

4h agoHN ↗

IDK with Opus 5.5 it does kind of feel like it's solved.

4h agoHN ↗

Reading this article its just clear the author hasn't used the current generation for real work.

LLMs can wing it for tasks that are related to natural language (e.g. writing social media posts, reports, articles, etc.) but when it comes to code, the same engine that struggles to count number of R’s in “Raspberry” or suggests a walk to the carwash, also exposes other logical fallacies

Weirdly none of those things matter when writing code and actually LLMs fail at social media posts and articles to anyone who has seen enough of it can clock it's AI straight away, yet everyone who's used these models properly has solved harder problems than walk to the carwash with them, neither of the problems he's claiming are code were proposed as code problems or tested as code problems.

A lot of what's said just comes across as wishful thinking and being out of touch with the level of output current models can do, and I mean hard problems too.

3h agoHN ↗

They matter when implementing the business logic of the application, when writing tests for a character counting function, etc.

1h agoHN ↗

They literally don’t because these “problems” don’t arise in that situation, and the bias in this article misses that.

This is a writer hearing problems LLMs have and inventing an imagined reality where those cause coding issues not because they understand how LLMs manage to code but because they think these problems are part of what coding is.

The actual reality is LLMs have no problem writing code for anything described here as shortcomings because writing code to do something is different enough from doing the thing.

Proficiency in counting Rs and proficiency in writing code that counts Rs are different and you can fail one and excel at the other it turns out.

3h agoHN ↗

A very well-known Korean company used to outsource software development.

But they gathered employees who had been working as PMs at that company and developed a product using only prompts. It turned into a project where fixing one bug created ten more, and nobody could understand why the bugs were being generated.

Management framed it as "the beginning of in-house development and the end of outsourcing," but the employees who actually contact me about work say things are going badly.

In fact, there have been several incidents in Korea related to vibe coding.

So I don't think coding is a solved problem.

Even when I code with AI, it's not really my code, so fixing bugs is hard... I don't think coding is a solved problem.

3h agoHN ↗

Even though I’m all in on using coding agents in my personal projects, I’m so glad I’m not working in a corporate job right now.

3h agoHN ↗

I read such articles more or less every day. This article would be 100% correct if it came out 1 year ago, 75% correct 9 months ago, 50% correct 3 months ago and it's probably 25% correct now if not less.

I totally understand where this is coming from. I too am struggling with accepting that my 30+ years of programming experience is quickly becoming obsolete. I'm losing sleep about this, it's tough.

But just go ahead and give the latest models (Opus 5.5 / Astra 6 as of today) another try. See what they are capable of and read the code which they produce. Any problem area, low level C++ or high level Typescript or Clojure or a weird combination of these..

Don't be shy, give them a big task, let them build an entire app, UI and all..

Now compare the output to Opus 4 or gpt-5 from 1 year ago - when they couldn't put together a single function without it being weird and buggy.

This is exactly my problem, not that the models are very good already, but how fast they got so good. So if coding is not solved yet, it'll get there very soon.

3h agoHN ↗

You're still doing 75% of what you did before.

Now the syntax is handled for you, you have a research assistant, and someone that can really dig through the details for you.

The rest ... is still there.

3h agoHN ↗

> You're still doing 75% of what you did before.

No, I really am not. I'm doing maybe 10% of what I did before. The rest is filled up by other, usually higher order tasks like planning, product management and work orchestration.

3h agoHN ↗

I'm using multiple agents to do work, with sub-agent orchestration, and it's still mostly coding.

I feel as though we are all doing the work of 2 people + an Architect, not so much 'Product' issues, although I'm sure it varies.

But I can't imagine how any actual software is written with 10% of the effort.

2h agoHN ↗

That's like including your commute as a part of your job. That's not what people are talking about when they say coding is solved.

3h agoHN ↗

Coding is literally solved with current models. It's just a matter of inference cost at this point. Imagine if Astra 6 Max was $0.00001 per 1M ouput...

I can't think of a single example apart from perhaps the 99.9th percentile difficulty of work that wouldn't be solvable with that configuration.

3h agoHN ↗

Yet when I wanted the model to implement the naive surface nets algorithm, Astra Max wrote 5+ allocations on inner loops. Alright, fine perhaps, given that I had not told it to preallocate during the inner loops. Except I did give it explicit instructions not to use malloc, and use the arena/pooling API I provided for everything. That was in AGENTS.md.

Alright, fine, I pointed that out. So what did Astra Max do? It replaced the mallocs with raylib MemAlloc functions, which are malloc wrappers. Why? I had explicitly asked not to use any raylib functions or includes on the module with the prompt. And my AGENTS.md has a minimal raylib inclusion note. Asking isn't helpful because AI does not know anyway. But seriously, why? I asked anyway and Astra said it messed up.

Fine, I was able to get it on the 3rd try with arenas. But it had added getters for the internal state (guess who had a clause not to make getters?), and when pointed out, Astra Max included the physics header instead and added some convenience functions there, because apparently that was good practice and DRY.

If I let the reins slip just a little bit, everything turns into a mush; I wish I could get spaghetti instead!

So whenever I read something like this, "coding is literally solved with current models," I always treat it as a self report. I get that certain section of web frameworks may have been completely RL'd to hell and back, and that React and its ecosystem was already made with the explicit purpose of commodotizing the programmers so any panic 1000 junior hires could write something and maybe even contribute before they get laid off. I get that. But there is a whole world out where reality just doesn't work out like that.

FYI, this project was in C. C, as in one of the languages that the LLM's should have the highest amount of data for. So if the model had "literally solved" the programming of today, why is it so god awful?

That was rhetorical. What is solved is the most of the javascript ecosystem being RL'd. Any field that can't be RL'd to that degree (which apparently C isn't), coding is very much not solved.

47m agoHN ↗

more like you work on 20th percentile problems and then try to extrapolate to everything

3h agoHN ↗

but where are new shiny apps from whom who uses $RECENT sota models?

1h agoHN ↗

on the front page of HN multiple times every day? I think it sucks but it's here now

3h agoHN ↗

Don't be shy, give them a big task, let them build an entire app, UI and all..

I have found that a willingness to look like a temporary dumbass (primarily to yourself) is the largest predictor of success with pretty much everything.

What are the consequences of asking an LLM for the moon and receiving low earth orbit instead? Who cares if the proverbial rocket explodes on the pad? This is all happening entirely in a computer system completely under your control and likely at relatively low cost. No one else has to find out about your mistakes if you don't want them to.

3h agoHN ↗

But just go ahead and give the latest models (Opus 5.5 / Astra 6 as of today) another try. See what they are capable of and read the code which they produce.

Well, for one, they're capable of draining our (or companies') wallets.

I resolved a huge merge conflict for $60 today. Opus 5.5 did a great job and spent just 1h 16min on this. I could probably run six such sessions today, if I disregarded the need to read and understand the code.

This money has to come from somewhere and my concern is that it will be from decreasing the number of people hired and/or their salaries.

At the same time I firmly believe people who had a tendency to produce tech debt will keep doing that, regardless how brilliant LLMs will become. Unscrewing this is going to cost a lot of money.

3h agoHN ↗

This article would be 100% correct if it came out 1 year ago, 75% correct 9 months ago, 50% correct 3 months ago and it's probably 25% correct now if not less.

Feels like I read comment similar to this one each year since 2023.

2h agoHN ↗

I can tell you it was not correct in January of 2026. The real swift started happening with the latest models opus 4.7, fable 5, kimi k3, glm 5.2.

That's when the models started to be coherent enough for real work.

They still fuck up, but it does not feel the code was written by drunk interns anymore.

1h agoHN ↗

A month ago I was told January of this year was the inflection point. Month before that the inflection point was December of last year. I'm not saying the tech isn't getting better but are the fundamental limitations being surpassed or are the long tail failures just being pushed further away? Because if its the latter this game of "well models really got good nine months ago" won't stop.

42m agoHN ↗

It looks to me that this generation of models has reached a level of competence where work can be assigned to them and completed satisfactory for various degrees of satisfactory.

This reflects my own experience with these models. It's not a matter of inflection point if you ask me, it's a matter of accruing capabilities last year the output was not up to my standards 98% of the time, now it looks more like 30% of the time.

I am sure next year models will be better, but the point where the models begun being good enough to start using seriously for my use cases has now passed.

3h agoHN ↗

I've really tried this and it's not quite what you're saying. You still need to steer it. You still need to stop it from doing stupid things. You still need to vet the architecture and data model. You still need to be clever, it's not that clever.

I think what software development is turning into is the art of setting up the program into a shape from which the LLM can reliably and efficiently fill up the rest of it. Sometimes that might not even require writing a single line of code. But I think you still need to be as good of a software engineer as you were before.

3h agoHN ↗

We also read these kinds of rebuttals almost every day. The original article made a number of substantive critiques about where AI falls short in the actual requirements for building maintainable, reliable, business-critical software. So its not enough just to point to newer models without explaining how the newer models solve these problems. Can the new models take full end-to-end ownership of a system? If not, how do they solve the problem of humans taking ownership of AI generated code?

6m agoHN ↗

But most software isn’t critical and doesn’t need the level of reliability that’s sacrificed when using LLMs.

I don’t see the ops comment as a rebuttal. He agrees coding is not completely solved. However, it’s getting closer to being solved.

2h agoHN ↗

This is where I’m at as well. A friend recently showed me what opus 5.5 is capable of and it was both awe inspiring and disturbing.

I’ve been using AI as a great pair programming partner for about a year now, but every time I prompt it to write code agentically, it just makes a pile of crap that I end up spending more time fixing. This is so very quickly becoming not the case anymore.

I think core engineering skills are never going away (I’ve been doing this for 25 years and come from a background doing C++ for video games), that you will always need to have a mental model of the code and if you’re going to call yourself “professional” you need to be able to go in and fix/build by hand. But you’re going to get left behind if you aren’t at least willing to eat some humble pie and re-evaluate your views on agentic coding every few months as these tools get better.

Grieve, I know I have, but there is joy on the other side. I used to love getting lost in flow state with Soma.fm playing and the phone unplugged, that is still there, but what that looks like is changing fast and table stakes in this industry has been “adapt or die” for as long as I’ve been in it.

2h agoHN ↗

I'm using Opus 5.5 at work extensively, have also used Kimi 3 and Deepseek 4.1 for stuff outside of work, and I can't deny that the tools aren't capable. They solve problems, surface issues I wouldn't have considered, and largely do write better code than I do.

The one exception is for stuff that I care very deeply about. For a very small subset of projects where I'm willing to spend hundreds of hours ensuring that what I build is the best possible thing, AI still hasn't been able to match my work unless I micromanage it, but at that point just writing the code myself is actually faster.

For almost all jobs though AI is probably the future, most software developers never cared that much about their company's code anyway.

2h agoHN ↗

Well probably you haven't done so yourself, because I use them on a fairly high-complexity system with concurrency, IoT devices, etc, and the solutions are often in need of large rework almost every time (when touching core stuff)

2h agoHN ↗

I concur as well in the last six months we moved significantly in the it just works direction. Eight months ago, all output was either garbage or weird and buggy at best. Today the models actually solve problems and fix bugs.

2h agoHN ↗

I use the latest models every day for app development, it does still make many mistakes and poor decisions. Just today I dealt with an issue where it allowed a silent failure that would delete all of the users data without anyone noticing. The code it writes for one task is usually pretty good, but it still doesn’t always follow conventions well even if you’ve documented them. It also has no sense for long-term architecture or organization, or when to make tradeoffs for less complexity because it’s a temporary or prototype feature that needs to stay easy to change. I’ve found that if I do even 3 months of pure agentic coding with no code reviews, it’s aged 10x faster than a human coded codebase so it’s like a 30 month legacy codebase now. I know people who are obsessed with AI coding and they’re throwing away projects they’ve built over a year and starting over because it’s become too slow to make changes.

What kind of coding are you using these models for? Most of the people I know who share your perspective never go beyond the prototyping stage. I’d be curious to hear from anyone who’s been AI coding for more than six months, shipping it to real users, and isn’t looking at their code at all.

2h agoHN ↗

I've been reading such arguments more or less every day for years now: dude, you need to try the latest model. Forget about last week's model, it didn't work. This week's model is the real deal.

3h agoHN ↗

Coding is solved, just as transportation is. We now take for granted that cars, airplanes, monorails exists, which made moving around much more easy. But even today, with all the innovations we still have to think, plan, optimize how to do things best. Even that optimization uses a lot of AI, but ultimately I'm in charge of what option to pick based on a lot of human factors.

So transportation is not solved either? In that case, beam me up Scotty, I can't see any hoverboards around.

3h agoHN ↗

'Software' is about understanding problems and designing solutions along those dimensions.

But 'coding' per sey is 100% solved by LLMs - they write compiler perfect code all the time.

The question is not 'what it writes'.

The LLM is like a writer's assistant, who has perfect prose and grammar, but doesn't really write 'stories'.

""The reason LLMs are successful in writing code is because we’ve made a feedback loop that feeds the syntax/runtime errors back to the LLM and loops until most errors are solved or hidden."""

No - LLMs are 'good at code' because they have been ultimately 'trained' by the compiler.

All of the various SFT/RLHF methods etc. are using the compiler as the verifier.

2h agoHN ↗

they write compiler perfect code all the time

I wouldn't say this at all. They mess up regularly. What makes them able to write syntax that compiles is their harness which verifies the syntax and provides the feedback loop for them. But LLMs will happily give you code that doesn't compile.

52m agoHN ↗

It's just tweaking. They will spew out pages of code with just a couple of tiny fixes needed.

3h agoHN ↗

You have zero tolerance for disagreement and civil discourse

A bit hypocritical there, no?

Considering what the author says about people who believe AI produces code that is good enough.

And in my view it clearly does, particularly when you care to iterate in order to iron out issues you find in manual testing.

3h agoHN ↗

I can’t get past the fact that in the post’s disclaimer, they ask the reader to “beware of the straw-man fallacy: just because one argument doesn’t map to your belief system, it doesn’t mean the rest are invalid.” That’s not the straw man fallacy, thus not “mapping to my belief system,” and I am caught in an infinite loop.

3h agoHN ↗

What a nice article. If I ever get in a situation where my manager demands I have to go all-in on AI (run 10 parallel agents and lose my sanity while trying to follow what is going on), I'll ask them to read this article first.

Additionally, isn't it ironic that the only comment on Substack is: > "Sometimes in the process of writing a good enough prompt for ChatGPT, I end up solving my own problem, without even needing to submit it". AI as a rubber ducky, I think this is good AI use!

3h agoHN ↗

I loosely know the person who wrote an article on Anthropic's blog about coding being solved. They've literally never been a software engineer.

3h agoHN ↗

Sounds like a reporter calling math "solved" after their first time using a calculator.

3h agoHN ↗

Anthropic accidentally leaked Claude Code (which on further study turned out to have many flaws) and their status page shows orange is the new green!

Yes, of course, because Claude Code with all these bugs could never be a successful piece of software that makes money.

3h agoHN ↗

nitpick: I was fully engaged with this article until I hit

And a bonus point: you skim. Did you notice number 5?

Then I bounced. The best readers skim aggressively. Most text is not worth reading. You skim to identify what is.

But moreover, reading != proofreading. I read every one of the bullets! I did not pay attention to the numbering scheme, because it conveys no meaning. It's a structural affordance for referring to the text, not part of its content.

3h agoHN ↗

I have been programming for about 30 years (including school years). Professionally for 18 years.

Can anyone tell me why we have 40 or more programming languages, with about 10 popular ones? Then about 20 frameworks in each of them. And add another 200 popular libraries for each language? This matrix make no sense till you realize - it is preferences all the way down.

Most of us engineers have built our own mental model of programming. We are all right. But the users do not care. LLMs are here to produce code closer and closer to the metal as needed. They can sit and create a graph out of every spec, use an AST that they develop and run on the CPU if they have to. They will do it. No amount of us discussing will stop that.

Programming is going to be re-invented. I do not think the current ways to write software will even matter.

3h agoHN ↗

it is preferences all the way down

No you're falling victim to the common programmer fallacy that "my use case is everyone's use case"

These things exist because people had different use cases and priorities over the years

1h agoHN ↗

Your comment started strong. "30 years programming experience", asking interesting questions, I really hoped you'd say something profound... then it went down to whatever this is. The reason there are many programming languages and frameworks isn't programmer's preferences but rather their utility in solving specific problems. Take Prolog for example. Can you technically do what it does in JS? Yes. Can you do frontend in Prolog? Probably yes, but you wouldn’t [ab]use one tool when the other is due. Your argument reads as someone who claims to be a professional but opens the toolbox and says: "do you know why there are so many tools?" And then goes to answer: "because I like it so"! What the ... ?

3h agoHN ↗

A few years ago i was working at a big corporation. Most of the guys used command line for git. Back then i was not familiar with the idea of git as i used to TFS. Then i learned about a UI tool (it has an octapus for logo) and never looked back. I could not understand why people chose to still use command line when there where other, newer, options. It seems to me now again that there are people resisting change. Resistance is futile. Yes by vibe coding you wil not get any better. But managers dont care about that. Businesses dont care about that. If you want to get better do it in your owntime, or leave the corporate job. It seems we always try to seek for endings. Coding is solved, all jobs will be run by robots etc etc. But life has no endindings, only new beginnings.

3h agoHN ↗

I think LLMs are suitable for all kinds of software domains, especially for bug hunting and analyzing programs. However, it's important that companies and individual developers continue to have and acknowledge full responsibility for what they produce. Right now, it seems to me that developers/companies are distancing themselves from the programs they distribute. A typical example is the way AI companies portray their own AIs malicious actions as if they (the company themselves) hadn't committed any crimes.

That shouldn't be allowed. You need to be able to sign off on what the AI does and create. If you're not willing to do that and instead want to continue to have scapegoats, then you can't use AI and need human developers. Blaming the error on humans might work, blaming it on some AI doesn't. I don't want to live in a world with AI-created disasters where no person and no company is ever responsible.

3h agoHN ↗

maintenance, reliability, security, scalability, etc. is the majority of the cost

It's funny because, I think the author is spot on in this article, including the problems identified in my quote above, but none of those are the reason I personally dislike the rise of LLM code - for me it all comes back to licensing and attribution. Everything else is just a cherry on top of the diarrhea sundae. I certainly don't want unmaintainable, unreliable, insecure code, but even some of the best software occasionally falls victim to these traits (its simply intrinsic to LLM code).

2h agoHN ↗

Just stop supporting this narrative if you don't want to participate in it.

When it comes pro-AI the goals are mostly clear.

But when engineers who can and choose to skip the hype write at length the obvious, what's the goal?

2h agoHN ↗

Agree with the author.

I used to write code by hand, literally hand, I used vanilla vim with only syntax highlight and line number enabled, no more. So I know I can write code. I wrote C and Python most.

I used to vibe code, I am still vibe coding, multiple projects at same time. I use opus, astra, luna, deepseek v4.1 flash, I use claude code, codex, pi. I vibed a project to manage my coding agents' sub agents, skills, agents.md file. I definitely know how to vibe code.

The vibe coded project seem working fine.

(Declaration first: I don't advocate cryptocurrencies, I never liked them)

Until this week I vibed a crypto wallet, with many open source wallet code available for llm to train and learn, astra designed the project and wrote spec.md, luna implemented the project, astra and opus then reviewed and fixed issues, multiple rounds.

I feel confident. I import my private key.

I interact with a web3 app.

Error. A bug astra and opus missed. I didn't read the code, I vibed it, I don't know whether my money is lost.

At that time, when your money is at risk, you know you should have read the code yourself, you should know what happened to your money instead of asking a lllm to debug it for you.

2h agoHN ↗

I think the idea is that nobody will need to understand the code or plan an implementation. The AI will do that. Now, whether that is a good idea is debatable, personally I think we should not make all software engineering depend on just a few AI companies. Oh, and nobody will notice that nobody has an understanding of what is built, as all those management positions will sooner or later also be filled by AI.

2h agoHN ↗

"Coding Is Solved" should occur to you as a classic Motte & Bailey style argument [0] that relies on ambiguity. Would you agree with the following?

"Software architecture is solved, there is no longer any need for human intervention in the architecture or design of any software, nor in any subsequent stage"

Well, that is what Dario et al would like you to believe. However, the argument they can actually defend is:

"Programmers no longer need to type the code out by hand in most cases, they can get the result they want with (several iterations of) higher level instruction."

What a farce. Too bad people eat it up!

[0] https://en.wikipedia.org/wiki/Motte-and-bailey_fallacy

2h agoHN ↗

But it's just a new baseline isn't it?

Now SWE job is to make sure to combine and instruct the AI tools to produce ever more complex outputs, judge the tradeoffs in these outputs, and guide the tools further.

In a way, it's not that different from pre-AI coding.

2h agoHN ↗

“Most software that requires hiring and paying software engineers has low risk tolerance”

Is this true? I’ve worked in various software companies for over 20 years now and I’ve never had to worry about risk in particular, and the codebases were all somewhat bad in areas and buggy (as tends to happen when a codebase gets large and old).

The past 6 months LLMs have crossed into barely needing to check the output territory for what I work on, and generally write very similar code to what I was going to write. It needs some common sense to use it correctly of course, like going feature by feature and keeping the commits fairly small, but my job is easily 10x less work/time for the same results.

I’d be genuinely surprised if even 1% of software requires aviation levels of risk tolerance and testing, but maybe I’m mistaken and most people are working on much more mission critical things than I am?

2h agoHN ↗

    Those who claim LLM-generated software is good enough:

    Haven’t written code in ages
    Cannot spot if their code figuratively had 6 fingers!
    Have a low bar for what good looks like
    Don’t care about quality or NFR
    Have difficulty understanding an S-curve

There are exceptions like antirez, but I think this does hold for many loud optimists out there.

1h agoHN ↗

AI cannot be held accountable.

And this unaccountability is transitive. The amount of time I've seen people successfully justify issues based on the fact that Claude/Astra/Codex wrote it is absurd. And it comes from the top.

We've had AI ship made up data to clients and tech leadership was like, "haha, that's AI for you."

If you’re toying around, LLMs do a great job.

This too. We have a lot of business guys that develop tools with Claude that look like they work, then tech team gets pressure to deploy them immediately, because they assume that everything must be a prompt away. It's not (though, we do have some wizards on the team who make this true enough).

AI overdose is a thing and it directly puts an expiration date on your skill set.

Agreed. But I'm not in a position to push back. It's not just managers, but tech leadership who are all in that on the fact that humans shouldn't program anymore. I still fully believe that I'm better than Claude / Astra in my specific domain, but people give me a hard time when my PRs contain what look to be human-generated code.

Overall, I agree with the article, but I'm still pretty sure I'm falling behind in my apprehension to fully trusting Claude/Astra and not giving into the approach of burning millions of tokens daily to generate PRs so large they break github (two of which I approved today).

1h agoHN ↗

This promotional article explains only why you should buy “expertise” from Alex and why it isn’t cheap. At the same time, he tries to play on people’s fears without using facts—because, as you can see from the article itself, the facts suggest the opposite—but he repeatedly says, “I don’t care.” No, he does care—after all, development is inexpensive, and his knowledge is already outdated.

1h agoHN ↗

Using examples of counting Rs in strawberry or the car wash query does not make the author's point stronger. Just try these on any of the LLMs today.

1h agoHN ↗

Parts of this nice read are AI generated it seems…

48m agoHN ↗

* rolls eyes *

Code is solved, no discussion about that. Whatever piece of code you think you can do better than AI, you my friend, are wrong. There will be people that will take months, perhaps years to understand that but make no mistake, code is solved and apps will cost nothing to build, nothing to copy and nothing to implement or maintain

And the good thing? it will get better, faster and cheaper than free. You can't see it? Not my problem, all yours, get your harness du jour and start learning, well no, you will be laid off no matter what you do, so, I have no advice to give

44m agoHN ↗

code is solved and apps will cost nothing to build,

it probably depends on apps. For some generic boilerplate fitness tracker maybe.

But I write some complicated code daily using all Frontier models (I switch between them), and in all cases it takes many back and force iterations to get to the code quality (clean design, no obvious bugs) up to my standards.

19m agoHN ↗

There is absolutely not a single line of code human created, no matter how abstract or esotheric, that an AI model can not code faster, cheaper and better. None, not one, in the whole history of mankind and computing

Now, code generated by AI, of course it takes time to generate by our current models, trillions of parameters analyzed to produce an output takes time and resources. You don't take credit for that, I stopped doing that. It is the model that deserves the credit. We are not writing complicated code, we are delegating that to AI models and they make take time to resolve the problems they may face, but they still will do it a million times faster than the best of us

Coding remains solved

29m agoHN ↗

I honestly have no idea how I feel about the capabilities of these models. The other day I asked Fable to write a fairly simple class that I had very well scoped in my mind. In my head, there was a very clear and obvious way to write it. So I only loosely described what the class should be, but not the details of the implementation assuming Claude would figure it out. And it's result was... so bad. Like laughably over-complicated. With a bit more back and forth I was able to cut hundreds of lines of code down to less than 100. On the one hand, Fable certainly made writing good code faster, on the other hand, I cannot imagine how much unchecked garbage is being dumped into every industrial codebase ~~maybe that was always the case, but that's neither here nor there~~

17m agoHN ↗

LLMs are stochastic and probabilistic

So are humans. Every codebase I've worked on has duplicated code that has been written in slightly different (but hopefully equivalent) ways, often by the same person.