Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005(cryptocellar.org)
    269comments
  2. Claude Opus 5.5(anthropic.com)
    194comments
  3. OpenAI is well positioned to fast-follow Jev(arcturus-labs.com)
    96comments
  4. 16-bit Intel 8088 chip(allpoetry.com)
    2comments
  5. Aging may be a program, not a breakdown(quantamagazine.org)
    discuss
  6. Writing Rust code that's fast by asking agents to make the code faster(minimaxir.com)
    14comments
  7. A WordPress vulnerability scored 9.2/10 is present in all versions since 2016(github.com/wordpress)
    2comments
  8. Apple has added persistent 'ads' to iOS, and it's driving users crazy(techradar.com)
    237comments
  9. Show HN: Drop – A rootless Linux sandbox with gVisor support(droprun.sh)
    27comments
  10. Show HN: AI·rete·RAG – a Rete rule engine decides, RAG explains why(ai-rete-rag.com)
    discuss
  11. Solitaire Alone Together(solitairealonetogether.com)
    11comments
  12. Can gzip be a language model?(nathan.rs)
    120comments
  13. MiMo v2.6(xiaomi.com)
    460comments
  14. Show HN: InstinctFlash – Run 5B world-action models in real time on Jetson Thor(github.com/general-instinct)
    discuss
  15. Spymarks, not Watermarks(brand.io)
    154comments
  16. I asked Meta’s Muse for its filesystem and it sent me 6.8GB(mouse.dev)
    83comments
  17. Meta’s Muse has a serious 0-day(arstechnica.com)
    28comments
  18. Xbox continues its “reset” with dramatic restructuring(arstechnica.com)
    12comments
  19. Vacate a drone restriction that criminalized recording immigration agents(eff.org)
    7comments
  20. MUNI Heritage Weekend in San Francisco(lawrence.lu)
    17comments
  21. Teleoperated Humans(jefftk.com)
    28comments
  22. The Economics of Open-Weight Inference(ornn.com)
    6comments
  23. Relativistic Raytracing(publish.obsidian.md)
    discuss
  24. Transformers Explained Visually(poloclub.github.io)
    84comments
  25. What Sun got wrong(dtrace.org)
    380comments
  26. I said no and Apple said yes(dbushell.com)
    516comments
  27. AMD's random number generator can't generate a 0?(flatassembler.net)
    148comments
  28. I don't want to read what you didn't write(colinbreck.com)
    387comments
  29. MiMo-v2.6-Pro: Intelligence, Performance and Price Analysis(artificialanalysis.ai)
    60comments
  30. Americans' Drinking Remains at Record Low: Gallup Poll(gallup.com)
    7comments

Writing Rust code that's fast by asking agents to make the code faster

35 pointsby 1h agominimaxir.com
14 comments
38m agoHN ↗

If it can be measured, then LLMs can optimize it.

Once I had repo commands that could dump `sample` results and a cpu profiler/trace and then a benchmark tool that let me A/A + ABBA/BAAB-test the current modified git workspace against HEAD or any commit, the LLMs could just do their thing.

And that's how my homemade terminal uses much less memory than ghostty/kitty/iterm yet has more throughput.

AI is going to increasingly unmask people and companies who don't care about correct and performant software now that it's become so trivial to guarantee both. It used to at least be expensive and time-consuming and expertise-demanding to do those things.

24m agoHN ↗

If it can be measured, then LLMs can optimize it

Then they can start attempting to optimize it. They can also spin round and round making the numbers worse because they don't actually know what to do.

23m agoHN ↗

Yeah, but that's just the scientific process of hypothesis -> evidence -> conclusion.

You need a measurement that can falsify hypotheses and reject branches that won't work.

Also, if all you have left in your project are performance issues that are hard to identify without flailing around (even with Fable/Astra) despite sampler/profiler reports, then you're doing really well.

20m agoHN ↗

The point of this post is that this is explicitly not the case. If the metric is measured, the agent finds a way eventually (around 5 total tries typically unless it gets stuck), and learns from iterations where changes caused a regression after a revert.

In one case I used a made-up metric (since I didn't know the exact name or if it existed) and it somehow optimized that too.

21m agoHN ↗

Agree it's amazing how much low-hanging performance fruit AI can trivially find. On the other hand though, once you get through the obvious no-brainer stuff, there's a lot of non-trivial tradeoffs in performance and I think that still demands a good amount of expertise to guide the AI in the right direction. Obviously AI will continue working it's way up the value chain, but I think there's a glass ceiling for AI where the right macro tradeoffs and perspectives on how software should work will bump into the hard and often articulated reality that different stakeholders want different things and often have either magical thinking or even self-deception about how those desires can co-exist with what everyone else wants.

This isn't a new problem by any means, but now that code is cheap, it means instead of getting frustrated with engineering and their pesky unimportant details, people will get frustrated with the AI and it's pesky unimportant details.

13m agoHN ↗

Yeah, the biggest example is performance optimizations that sacrifice your data model to the point that you'd never accept them.

I think it's one reason why ADRs are an important of a software project, especially with LLMs. You need a place were you can document invariants, why you have them + the rejected ideas and acceptable risks.

It helps smart agents like Fable help you decide on trade-offs and it's kind of incredible to witness that happening.

33m agoHN ↗

I have done a fair amount of low level performance optimization with Opus 5 and its reasoning is still very poor. Like why is CRC so slow and going through loops until I ask it if is using hardware instructions and it tells me it is using its own hand coded implementation poor. Reasoning about l1/l2/l3 cache hit ratios and their implications basically throwing darts at the wall, in the wrong room. If you give it a benchmark feedback loop then it might get there eventually but still massive alpha for low level systems engineers who instinctively know how this stuff works and can now automate 99% of the grind.

13m agoHN ↗

Can you give it the valgrind suite to loop over? I have yet to try that with AI, but maybe the cachegrind tool is enough to help it.

32m agoHN ↗

I've found that this is a good way to introduce bespoke code into the codebase whose perf gains don't generalize.

Unless its an easy memory/parallel/algorithmic win, its not worth it.

30m agoHN ↗

They're quite good at just iterating different "ideas" on a performance metric with an objective measure. They can use tools like `perf` and do some analysis on the output. Sometimes they go off in the weeds unproductively, and sometimes they give up because your goal was too high, but as long as you're sort of babysitting the process, you can make pretty rapid improvement to naive code.

26m agoHN ↗

"Rank the top findings/solutions by impact vs confidence" continues to be one of my best quickwins to add to all sorts of prompts. Bam, now you have a reasoned priority list.

14m agoHN ↗

I've been having great results with this kind of thing. I find that really just need a sensible framework within which the optimization can take place. Essentially just providing the measurement harness, and some sort of motivation for what I'm doing.

The great thing about LLM is that it seems to have the checklist for everything. If I rattle off a few things like "don't allocate on the hot path" and "remember to pin the cores" it will come up with a few items of its own that I might have forgotten.

Eventually, it will have gone through the whole list with me, while having documented all the measurements along the way.

But it's still guided by experience. If I see unusual numbers, I might say "hey did you forget to compile it in release mode?" and it will apologize and fix that. If I don't, it may just continue exploring without realising everything is wrong.

11m agoHN ↗

one thing nobody mentioned here, once the agent is looping against the same benchmark it will happily optimize for the benchmark itself and not the real workload. worth rerunning the win against a slightly different input shape after, just to check it did not memorize the harness instead of actually fixing anything.

6m agoHN ↗

There is an entire paragraph about benchmaxxing and how to mitigate that.