- 37comments
- 60comments
- 1comments
- 405comments
- 113comments
- 94comments
- 65comments
- 245comments
- 343comments
- 19comments
- 282comments
- 270comments
- 59comments
- 29comments
- 324comments
- 10comments
- 7comments
- 83comments
- 1comments
- 66comments
- 2comments
- 34comments
- 10comments
- 22comments
- 38comments
- 60comments
- 63comments
- 40comments
- 81comments
- 488comments
LLMs are very good at lossless compression via arithmetic coding. But I didn't know that it was possible to go the reverse direction (do language modeling via a compressor). It's not super great quality, but I'm surprised it worked! Other compression algorithms (like PPMd) use variable n-grams under the hood, and should be much better (although less interesting due to already containing basic language models internally).
Reminds me of this youtube video: https://m.youtube.com/watch?v=jkdWzvMOPuo
I liked the comments explaining why this worked.