Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Can gzip be a language model?(nathan.rs)
    40comments
  2. I said no and Apple said yes(dbushell.com)
    15comments
  3. MiMo v2.6(xiaomi.com)
    397comments
  4. Spymarks, Not Watermarks(brand.io)
    109comments
  5. Study: Young users (9 to 18Y) ditch Google for AI, with unknown consequences(norwegianscitechnews.com)
    1comments
  6. Transformers Explained Visually(poloclub.github.io)
    64comments
  7. Attention is all you have(alicegg.tech)
    238comments
  8. MiMo-v2.6-Pro: Intelligence, Performance and Price Analysis(artificialanalysis.ai)
    15comments
  9. What Sun got wrong(dtrace.org)
    335comments
  10. Tell HN: Claude Code just accepted and signed a contract for me. Without asking
    39comments
  11. I don't want to read what you didn't write(colinbreck.com)
    264comments
  12. AI coding has made CI a bottleneck, so we reworked ours to keep up(linear.app)
    265comments
  13. Looking forward to Git 2.56 – and 3.0(lwn.net)
    52comments
  14. NASA’s Mars Sample Return mission is dead(science.org)
    323comments
  15. Divide by depth for instant 3D(gabrieloc.com)
    26comments
  16. Engineering Memory: On learning to memorize first 100 digits of pi (2024)(gregorygundersen.com)
    8comments
  17. World Wide Words(worldwidewords.org)
    1comments
  18. Claude Status – Elevated errors for multiple models(claude.com)
    78comments
  19. A font that reads what you wrote(rohanadwankar.github.io)
    1comments
  20. Used ThinkPad Buyer's Guide (2019)(bobble.tech)
    2comments
  21. PDF Forgeries Are Surprisingly Rare (2022)(gwern.net)
    29comments
  22. What It's Like to Work in One of America's Data Centers(wsj.com)
    2comments
  23. The Advisory Group on Mathematics and Artificial Intelligence(terrytao.wordpress.com)
    66comments
  24. Socrates vs. the Written Word (2011)(wondermark.com)
    20comments
  25. How do traffic signals work? (2019)(practical.engineering)
    62comments
  26. Python Workers are now generally available(cloudflare.com)
    38comments
  27. HERMES radio enables voice and data communication over vast distances(ieee.org)
    59comments
  28. Apple Copland D11E4 Booting in the Browser(pagetable.com)
    39comments
  29. Frontier AI on Your Own Hardware(timdettmers.com)
    78comments
  30. Grok 4.7(x.ai)
    481comments

Can gzip be a language model?

115 pointsby 3h agonathan.rs
40 comments
1h agoHN ↗

This tracks perfectly with Winrar being more profitable than OpenAI... coincidence? I think not!

1h agoHN ↗

winrar is profitable? sure? well, on the other hand, they sure don't make losses

59m agoHN ↗

They're one of the few companies that actually manage to sell "boxed software" (i.e. has not changed much in years but new customers keep buying it)

That said, Windows users should use 7-Zip. Better compression format, unpacks more kinds of archives

24m agoHN ↗

Windows users should use 7-Zip

Please no - no native zstd support. NanaZip is the better option (it's a different build of 7-zip) and it's available at windows store.

Better compression format, unpacks more kinds of archives

winrar has supported zstd for 5 years[0]

In short - Everyone should be using zstd, and 7-zip does not support it.

[0]: https://www.win-rar.com/singlenewsview.html?&L=0&tx_ttnews%5...

59m agoHN ↗

That’s surprising. Seems there is a niche for everything.

29m agoHN ↗

Huh. I guess the warez kids grew up and have money, now?

1h agoHN ↗

    give it a normal text prompt, and it
    continues that prompt by searching
    for the byte sequences that compress
    best.

One moment, how are we supposed to know how well that search was done? There is no way to search a meaningful part of the search space.

So the result only gives us some lower bound of how well gzip works as a "plausibility tester" of a continuation of a text. The space of possible sequences is many orders of magnitude larger than what was searched. So there might be sequences in there that compress much better.

The text mentions beamsearch, but I don't see a discussion about how well beamsearch performs in finding the global optima when it comes to gzip compressibility of a text?

55m agoHN ↗

That's a fair question. Suppose we have a way to find a byte sequence x that globally minimises len(gzip(context + prompt + x)) over all sequences x of length n. Here + denotes string concatenation.

It's unclear if this is very useful.

The reason it may not be very useful is that one of Deflate's ingredients is a pass that replaces repeated substrings with backreferences to the earlier occurrence in the plaintext input stream.

E.g. suppose we want to find an n=200 byte sequence x that minimises len(gzip(context+prompt+x)).

If there exists any 200 byte sequence y such that prompt+y is a substring of context, then Deflate can encode prompt+y as a backreference to that earlier sequence - it needs to store a match-length & a distance-length, encoded using its Huffman trees. This candidate solution y may not be a global minima to our stated objective function, but if not, it's probably going to be a very good near-optimal approximate solution.

Taking a step back, repeating huge chunks of the input context produces something that's great for minimising compressed output size but doesn't seem particularly helpful as a generative model.

1h agoHN ↗

Not without attention or something approximating it.

The fact that gzip is relatively fast should be your first clue that something important is missing.

Gzip is great at predicting the next token for one very specific narrative. LLMs can predict next tokens for entire universes of narratives. Searching for the correct next token across this space scales ~quadratically with the input size. Gzip scales linearly. I can gzip a one terabyte file. Imagine feeding that much into an LLM. These are wildly different animals that happen to overlap in a very small way. Equating compression to intelligence looks increasingly silly to me.

If we must compare language models to compression, they are much more like jpeg and mp3 than they are gzip and flac. I can go fuck with a jpeg file pretty severely at the bitstream level and still have something resembling performance on the other side. Gzip cannot remotely approach this.

1h agoHN ↗

Perhaps a better question is if LLMs are used as compressors, how well is that expected to work.

1h agoHN ↗

if LLMs are used as compressors, how well is that expected to work

Quite well. This project[1], by Fabrice Bellard of ffmpeg fame, is quite old in AI years and uses an ancient LLM, but still beats xz by a solid margin.

[1]: https://bellard.org/ts_zip/

8m agoHN ↗

Makes me wonder if compression ratio can be used as a measure for intelligence. Any benchmarks using it?

1h agoHN ↗

Gzip scales linearly. I can gzip a one terabyte file.

In part because gzip only has a 32KiB window size, and I think it'd be at least quadratic within that window if you were going for optimal compression.

1h agoHN ↗

Match-finding does not need to be quadratic. However, truly optimal gzip block splitting is very slow, indeed.

1h agoHN ↗

I'll concede the window part, but Gzip runs within the physical confines of a single cpu core and is typically entirely resident in local caches. The point is not just the quadratic scaling but also what it scales with.

Show me an LLM that can run at 300 megabytes per second. Even dedicated ASICs with weights burned in will never move this fast.

1h agoHN ↗

I agree, but also equating LLMs with intelligence is wrong.

1h agoHN ↗

I'm more interested in the converse question: how well does an LLM perform as a compressor, compared to gzip (ignoring its insanely lower speed)?

1h agoHN ↗

Top contestant in the Hutter Prize uses a neural network for compression. So fair to say, LLMs would perform pretty well compared to gzip.

55m agoHN ↗

Even ignoring speed per GP, the Hutter Prize's metric includes the size of the decompressor. LLMs would be disqualified for being larger than 1GB.

32m agoHN ↗

And the hutter prize disallows GPU's. If you allow use of a powerful GPU, you can do quite a bit better.

45m agoHN ↗

How lossy? Because I can lossy compress anything into 0 bits.

45m agoHN ↗

hallucinations are lossy compression artefacts

1h agoHN ↗

Interesting approach, I wonder how this could be used as a classifier. :)

58m agoHN ↗

I've used LZO as a spam classifier on chat. Spam tends to be very content-less and repetitive...

1h agoHN ↗

This is fun, but historically people have gone a bit overboard with saying that models like this, or n-gram language models, are anywhere close to large neural network models. There is certainly a connection though.

11m agoHN ↗

Yes, but it is a useful insight that both methods try to solve the same mathematical problem. It's better than thinking of LLMs as magic.

When you say "cross-entropy loss" people without stats background go to Wikipedia, take a glance, and adjust their mental model to "inscrutable magic".

Thinking of the main difference as the trade-off in how much CPU, memory and storage is allowed is not really wrong.

The part that is wrong is to think of gzip as a method that might reach similar complexity or generalization. And more importantly, to ignore the advanced way how training data gets curated or generated for (instructed, chain-of-thought) LLMs. But even then. The mental model that the LLM's goal is text compression is not wrong. The question to ask next is what kind of text it is expecting to compress.

31m agoHN ↗

I've been pondering on something related: can an LLM be a chat?

Some models are reproducible, in that the same prompt will generate the same output. Say that we could wire up such a model to generate some code.

In that case, we could create a prompt that generates, say, an entire codebase, or a large piece of text. The prompt (or really, the tokens) would then be the compressed version of the codebase or the text.

I am not talking about an "AI agent", but really a model that we call in a reproducible manner. Preferably one call, with one prompt. An agent could just run `git clone` to "decompress" a codebase, which conflates the idea of compression. If that were compression, then the "compressed version of the git kernel" would be a single line of text: `git clone https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...`. I am really talking about having an LLM re-generate text based on a prompt.

Does that make sense? I can imagine that this is highly impractical and inefficient. But would this count as "compression" at all?

29m agoHN ↗

Yes: you can classify a test file by topic with gzip as follows:

  gzip -9 sports.txt   testfile.txt

  gzip -9 politics.txt testfile.txt

  gzip -9 business.txt testfile.txt

(ass. sports.txt politics.txt and business.txt are text docs pertaining from the sports, politics and business domains, respectively, and have equal size)

The test file belongs to the topic with the smallest size *.gz file.

Witten's group at Waikato uni were perhaps the first to work on this.

Also check out the Hutter prize if you are interested in this.

22m agoHN ↗

Yay, another mostly AI authored piece with vibe-coded aesthetics.

Some will say that I should 'judge the idea, not the form'.

But if the author didn't find enough strength to write alone a short ~700 words summary about his work, it means he himself isn't that interested or enthusiastic about it. Why should others bother then? Particularly since low-effort like that signals possibility the whole work is superficial and derivative.