Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Claude discovers a novel enzyme system with CRISPR-like repeats(anthropic.com)
    437comments
  2. VSCode's SSH Agent Is Bananas (2025)(fly.io)
    59comments
  3. Linux support is coming to Snapdragon X2 Series(qualcomm.com)
    7comments
  4. Fixing the Portobello Police Station Clock(pointinthecloud.com)
    84comments
  5. LensVLM: Compressing long context as images, expanding only relevant pages(huggingface.co)
    2comments
  6. We just shipped support for the ugliest part of HTTP: Vary – Cloudflare Blog(cloudflare.com)
    1comments
  7. Italian parliament votes for return to nuclear energy(apnews.com)
    314comments
  8. The mystery animal on an ancient god's head(signoregalilei.com)
    8comments
  9. A brief history of Windows scroll bar shortcuts(devblogs.microsoft.com/oldnewthing)
    39comments
  10. The Curious Power of Punctuation(newyorker.com)
    2comments
  11. Jev in 25 Lines of Python(nobodywho.ai)
    193comments
  12. Gemini 3.8 text-to-speech(blog.google)
    115comments
  13. Radicle: Disclosure of Vulnerability in the Network Protocol(radicle.dev)
    42comments
  14. Tokens too cheap to meter(jyn.dev)
    174comments
  15. Show HN: An atlas of system designs with interactive architecture diagrams(atlas-sysdes.vercel.app)
    17comments
  16. Mercury 2.5 LLM hits 770 tokens per second(artificialanalysis.ai)
    3comments
  17. Making Tailscale Faster(tailscale.com)
    6comments
  18. Swap, ZRAM, Zswap and Hibernate on NixOS(matthewbrunelle.com)
    4comments
  19. Stripe's Knowledge AI Platform(stripe.dev)
    102comments
  20. I don't want the details(michaelheap.com)
    189comments
  21. Z80 REPL (2018)(abagames.github.io)
    18comments
  22. Claude Code reads AGENTS.md only when telemetry is on [fixed](szypowi.cz)
    242comments
  23. Bwbach, My Guardian Goblin(robertmay.photography)
    4comments
  24. Once Claude can measure something, it can make it faster(claude.dev)
    84comments
  25. A refined phylochronology of the second plague pandemic in Western Eurasia(pnas.org)
    discuss
  26. QuestDB (YC S20) Is Hiring a Sales Engineer(questdb.com)
    discuss
  27. 28% of job postings on company career sites have been open over 90 days(unlisted.careers)
    273comments
  28. Show HN: I built a post-mortem debugger for native Windows x64/x86 crashes(forensicdbg.com)
    2comments
  29. OpenAI breaches Medicare, Albanese reveals(smh.com.au)
    72comments
  30. UK military jamming other nations' satellites to defend itself, BBC told(bbc.com)
    174comments

LensVLM: Compressing long context as images, expanding only relevant pages

31 pointsby 4h agohuggingface.co
2 comments
1h agoHN ↗

I really like this approach! I sort of think of the vision encoder here as an expensive high fidelity RAG encoder.

The thing I’d love to do with a system like this is train it to be KV cache ordering independent (ie permutation invariant at the page level). Basically each page’s KV cache should be understandable by the model in any ordering - which would allow you to go one step further and treat the KV cache of the vision encoded page as the chunk for the model to reason over.

Then all these zoom in for more detail tricks will extend naturally.