Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Training a 4B model to produce 81% faster query plans than Postgres(rohanbansal.com ↗)
    32comments
  2. Xiaomi Mimo 2.6 live post-training dashboard(xiaomi.com ↗)
    26comments
  3. macOS 27 Golden Gate – Review(arstechnica.com ↗)
    22comments
  4. Small programming tricks(will-keleher.com ↗)
    151comments
  5. Reversing Factorio's RNG(gegell.github.io ↗)
    5comments
  6. Accurate Models of AMD Matrix Cores(arxiv.org ↗)
    5comments
  7. Breaking the 1.58-bit Barrier for Ternary LLMs(arxiv.org ↗)
    discuss
  8. Vectorized and performance-portable Quicksort (2022)(googleblog.com ↗)
    24comments
  9. How good are frontier models at physics?(arxiv.org ↗)
    13comments
  10. Dream-RSI: Recursive Self-Improvement through Evolving Worlds(arxiv.org ↗)
    48comments
  11. WalShadow: Sub-second Postgres replication to ClickHouse from physical WAL(clickhouse.com ↗)
    2comments
  12. Anatomy of a Texture(agentlien.github.io ↗)
    9comments
  13. Performance Improvements in .NET 11(devblogs.microsoft.com/dotnet ↗)
    3comments
  14. Mistral X Mozilla: Private, Multilingual AI Browsing(mistral.ai ↗)
    176comments
  15. AWS says it can't restore some data from mideast facilities struck by Iran(wsj.com ↗)
    7comments
  16. The Siberian Ice Maiden and the Scythian World(patrickwyman.substack.com ↗)
    1comments
  17. Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations(github.com/arnegiacomo ↗)
    234comments
  18. Training Text-to-Image Models 3.6× Faster(linum.ai ↗)
    1comments
  19. Japan's book scene is moving from bookstores to libraries(untranslatedjp.substack.com ↗)
    1comments
  20. Anecdotally, programmers dislike "reduce"(evanhahn.com ↗)
    72comments
  21. Tell the speakers that you liked their talks(ohhelloana.blog ↗)
    69comments
  22. Show HN: Restarted – a 2026 remake of the classic 2015 startup generator(restarted.io ↗)
    1comments
  23. Show HN: AttaLambda: a language where types and data are made of untyped lambdas(attalambda.com ↗)
    discuss
  24. A warning about 'model welfare'(mustafa-suleyman.ai ↗)
    394comments
  25. Kyber (YC W23) Is Hiring a Forward Deployed Engineer(ycombinator.com ↗)
    discuss
  26. The DeepMind Institute(deepmind.com ↗)
    33comments
  27. Claude Cowork and chat are now one Claude(claude.com ↗)
    183comments
  28. How big are factorials?(thegreenplace.net ↗)
    30comments
  29. Randomized query complexity can beat certificate complexity(arxiv.org ↗)
    discuss
  30. Douglas Adams and the exterminated Doctor Who adventure(bbc.co.uk ↗)
    61comments

Vectorized and performance-portable Quicksort (2022)

151 pointsby 2h agoopensource.googleblog.com
21 comments
2h agoHN ↗

Very nice. Let's see Paul Allen's quicksort

2h agoHN ↗

(2022)

If you're curious why you would want a vectorized way to sort lists of numbers, one use case is building histograms - it's much easier to build a histogram if you've sorted all your samples first

2h agoHN ↗

(2002)

I wonder what apps have implemented this now that a few years have passed.

2h agoHN ↗

Actual title: “Vectorized and performance-portable Quicksort” (2022).

Actual sense in which it’s first:

Happily, modern instruction sets (Arm SVE, RISC-V V, x86 AVX-512) include a special instruction suitable for partitioning. Given a separate input of yes/no values (whether an element is less than the pivot), this "compress-store" instruction stores to consecutive memory only the elements whose corresponding input is "yes". We can then logically negate the yes/no values and apply the instruction again to write the elements to the other partition. This strategy has been used in an AVX-512-specific Quicksort. But what about other instruction sets such as AVX2 that don't have compress-store? Previous work has shown how to emulate this instruction using permute instructions.

We build on these techniques to achieve the first vectorized Quicksort that is portable to six instruction sets across three architectures, and in fact outperforms prior architecture-specific sorts.

2h agoHN ↗

Our implementation uses Highway's portable SIMD functions, so we do not have to re-implement about 3,000 lines of C++ for each platform.

Would they do the same thing today or have an LLM re-implement those 3000 lines of c++ ?

2h agoHN ↗

Id say that it is a lot more likely for the in-house, at least half a decade old library to be correct and performant than 3000 lines of an LLMs mediocre regurgitation of that code.

1h agoHN ↗

Guys, remember when language features allowed re-usability?

1h agoHN ↗

Barely, and rarely for C++ specifically.

I think software engineering in general is in a bit of a discoverability crisis. So many problems actually have solutions implemented... Somewhere. If you know about them. And are speaking the same vocabulary as the original implementer to realize the solution might be applicable to your problem. It's one of the reasons that jokes exist about microservice frameworks (https://www.youtube.com/watch?v=y8OnoxKotPQ) and how "We use Hadoop to store the output from our Kafka pipe, that's populated from our Traefik layer, all monitored with Grafana in front of Loki and Prometheus, of course" is a real sentence that has actual meaning and not a fever-dream.

LLMs are actually pretty impressive at being able to pull together disparate information from various domains into one place.

2h agoHN ↗

Well, it came out a while ago, so maybe we can be a bit silly:

There’s something sort of beautiful about mergesort and heapsort. Their names tell you what their main idea is, and how they work is immediately obvious.

Quicksort, on the other hand, has nothing beautiful about it and is named after it’s one redeeming feature (that it is quick for a lot of cases).

1h agoHN ↗

The beauties of quicksort are that it sorts in-place and that it is embarrassingly simple.

The in-place property can be utilized to make it very close to cache-oblivious algorithm.

1h agoHN ↗

How is it less beautiful?

Honest question, curious to hear about what that means to you.

16m agoHN ↗

To me it’s the fact that if you try to do it in real life (i.e. sorting a collection of objects in the real world), you just end up doing merge sort by accident.

So it feels like an optimization of merge sort for computers, rather than a different approach.

This is also reflected in the way that it’s usually taught. Normally merge sort is presented first, and quick sort follows from observations about what would happen if you picked random partition points instead of dividing them in half, and how you can reduce the additional space requirements.

1h agoHN ↗

Maybe if Hoare had called it Partitionsort in 1960, the name wouldn’t have been memorable enough to catch on and become so popular.

2h agoHN ↗

Definitely needs (2022) in the title, I was a bit confused!

1h agoHN ↗

only sorts numbers? wouldn't radix be much better?

1h agoHN ↗

Just imagine when candidates will get asked by pre-revenue startups to implement a vectorized version of quick-sort in person in 10 mins, just for a SWE job which they do not use this themselves.

Only the likes of MAG 7, and a couple of hedge-funds would ask to do it since this problem directly applies to them.

But certainly not pre-revenue startups.

1h agoHN ↗

this was made like 4 years back most of the current algorithms use this already !

1h agoHN ↗

I didn't like the 9 MB image file in the blog. It took me few seconds to fully render the image.

17m agoHN ↗

Just wondering: Humans can visually spot the smallest and largest items from among 1000s of items, almost instantly (usually, depending on the variance)

Could AI be used this way? Just splat a visual representation of each item on a virtual wall and have an AI "visually" pick them out?