Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Astra for Law(openai.com ↗)
    224comments
  2. Bend – A language that blocks AI mistakes via proof, on CPU and GPU(bend-lang.com ↗)
    113comments
  3. Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint(prismml.com ↗)
    31comments
  4. Hister: A private search engine for the pages you visit and the files you keep(github.com/asciimoo ↗)
    122comments
  5. Wax motor(wikipedia.org ↗)
    38comments
  6. Fujitsu launches made-in-Japan next-generation CPU FUJITSU-MONAKA(global.fujitsu ↗)
    175comments
  7. Sex, AI, and the Apocalypse(iankduncan.com ↗)
    38comments
  8. Flet 1.0 – Build cross-platform apps in Python(flet.dev ↗)
    5comments
  9. How to Write with an LLM(sockpuppet.org ↗)
    12comments
  10. Diplodocus, Long Thought Exclusively American, Turns Up in Spain(sci.news ↗)
    5comments
  11. CrowdSec Source Code Leak(crowdsec.net ↗)
    35comments
  12. Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data(arxiv.org ↗)
    26comments
  13. Rate limits on GitLab.com are changing(about.gitlab.com ↗)
    103comments
  14. How Uber Protects Against Retry Storms(uber.com ↗)
    8comments
  15. The American Religion of Self-Storage Facilities(newyorker.com ↗)
    296comments
  16. The most important product decision is what you don't build(liamnugent.me ↗)
    6comments
  17. TSMC revealing details about next gen A14 node(mapyourshow.com ↗)
    28comments
  18. Zettascale (YC S24) Is Hiring ASIC/FPGA Engineers to Build Chips for ASI(zscc.ai ↗)
    discuss
  19. How GLM built its own inference infrastructure(z.ai ↗)
    257comments
  20. Why I didn’t sign the Fields medallists’ letter(gowers.wordpress.com ↗)
    242comments
  21. More than 100k people in Japan are now aged 100 or older(bbc.com ↗)
    discuss
  22. How do we prevent mathemathics from devolving into the Medieval Era of secrecy?(mathoverflow.net ↗)
    34comments
  23. Everybody's Lost Their Minds(netmeister.org ↗)
    178comments
  24. One year of sponsored Servo development(servo.org ↗)
    138comments
  25. Canto: A speech model built for the real world(wisprflow.ai ↗)
    12comments
  26. Show HN: Snapdrop: Instantly share files between devices. No setup, no signup(snapdrop.me ↗)
    7comments
  27. Running Ubuntu on the Lenovo IdeaPad Duet(vhaudiquet.fr ↗)
    19comments
  28. CCC invites all model citizens to 40C3(ccc.de ↗)
    175comments
  29. André Weil and the Hodge Conjecture(jiahao116.github.io ↗)
    8comments
  30. Launch HN: Skillsync (YC W26) – AI chat sessions made portable across agents
    43comments

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

105 pointsby 1h agoprismml.com
31 comments
1h agoHN ↗

Love this for the folks with 16gb graphics cards - 3.8 27b has been incredible but not quite runnable on anything less than 32gb - will try loading this up on my 16gb intel b50 and see how it goes - not sure these quants can be accelerated by the XPU cores yet but maybe in time!

1h agoHN ↗

You can run the ~4 bit quant(s) on 24gb, if you're not _too_ picky on context size.

This will hopefully be better, though it'd be a _very_ surprising increase in performace at the size they say. Would love to see more about how it benchmarks.

1h agoHN ↗

I run Unsloth's UD-Q4_K_S on 20 GB of VRAM (RX 7900 XT) and I get ~90k tokens of context without quantizing KV cache. With 8-bit quantization, I get about a 134k token context window. That's with only one slot, but for me, it works pretty darn well, with 20-35 tok/s depending on how full that window is.

1h agoHN ↗

I'm not following the local mdoel scene too closely but this seems quite amazing. Is this able to be run on Apple silicon too?

1h agoHN ↗

"Ternary Bonsai 2 27B reaches up to 143 tokens/second on NVIDIA GeForce RTX 5090 and 46.8 tokens/second on M5 Max. On an RTX 4090, Ternary Bonsai 2 27B consumes just 0.714 mWh/token, making it 40% more energy-efficient than an 8B model running in full-precision."

1h agoHN ↗

Their mention of the 5090 is bit odd, since on 32 GB GPUs, Q6 fits while having better quality. Very interesting model for 16 GB GPUs though!

56m agoHN ↗

They mention 5090 with regards to speed, Q6 will not have that speed?

And speed matters a lot for many use cases

15m agoHN ↗

150 tokens per second on a ternary model implies that it’s GPU bound, I’d bet a Q6 model is even faster because it’s existed longer and seen more optimization. You’d have to be insane to not run an NVFP4 quant over a ternary quant on Blackwell if they both fit.

42m agoHN ↗

it is like "my fridge is 2mkm (millikilometer) from my desk" m=0.001 h=3600 it should be just Ws or just J

1h agoHN ↗

Their first 27B bonsai was able to run on an iphone.

1h agoHN ↗

I'd love to see a Bonsai model start with a 100B+ parameter model and get that down to <30 GB. But maybe at that point we call it Topiary?

1h agoHN ↗

I'm hoping they release an 8B v2 based on the Qwen 3.8 series in the near future - that would give us a really powerful model that could be run directly on users phones.

37m agoHN ↗

Yes! And maybe get a hf fused webgpu runner for that model so it’s fast!

1h agoHN ↗

If you want to try out out the GGUFs from https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf#th... be aware that you need Prism's llama.cpp fork to get them to work, from https://github.com/PrismML-Eng/llama.cpp/releases/tag/prism-...

This should work:

  cd /tmp

  # Get the Prism macOS runtime
  curl -fL https://github.com/PrismML-Eng/llama.cpp/releases/download/prism-b10685-7dffb15/llama-prism-b10685-7dffb15-bin-macos-arm64.tar.gz -o bonsai-runtime.tar.gz
  tar -xzf bonsai-runtime.tar.gz

  # Get the ~5.95 GB GGUF model:
  curl -fL https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/resolve/main/Ternary-Bonsai-2-27B-PTQ1_0.gguf -o Ternary-Bonsai-2-27B-PTQ1_0.gguf

  # Run the server, I used port 8331
  ./llama-prism-b10685-7dffb15/llama-server \
    -m Ternary-Bonsai-2-27B-PTQ1_0.gguf \
    --port 8331 -ngl 99 -fa on -c 32768

Then open http://localhost:8331 for the (very good) baked in llama-server web UI... or run a prompt via the API like this:

  uvx llm openai endpoint http://127.0.0.1:8331/v1 \
    --model bonsai-2-27b --responses hi

That's running at ~20 token/second for me on an M5 Pro (after a server restart I got 44 token/second, not sure why), but I'm pretty sure something isn't working right, on startup the server said "ggml_metal_device_init: - the tensor API is not supported in this environment - disabling".

14m agoHN ↗

Honestly looks pretty good except whatever is going on with its booty. Is that an ass helmet? I cannot parse what's going on there.

6m agoHN ↗

I figured them out, starting from the GGUF on Hugging Face.

If you have found better instructions and they work then use those instead!

51m agoHN ↗

I really wish people would stop saying N times smaller than something when making a comparison; that makes no sense - it's 1/9th (11.11%) the size. You don't get a smaller quantity by multiplying by a number greater than 1.0. You could instead reverse the subjects being compared - "the original model is 9x bigger than this new smaller, efficient model" or some such. That makes sense.

I keep seeing this being used when people talk about efficiency or performance gains and it's just very unintuitive language.

44m agoHN ↗

If we were talking about speed instead of size, i think it would be perfectly reasonable to say 9x faster. I'm not sure I agree that 9x smaller is unintuitive. It makes sense to me.

43m agoHN ↗

Me too!! It's a huge pet peeve. And it's so hard to get people to see how it's linguistically AND mathematically WRONG.

41m agoHN ↗

I think if they made this for Qwen3.8-Next it could fit in a single 5090?

26m agoHN ↗

Ternary Bonsai 2 27B uses ternary {−1, 0, +1} weights with FP16 group-wise scaling, for 1.76 effective bits per weight

If I recall correctly, a recent post [1] has shown that Q2 quants (with like 2.6 bpw) of the same base Qwen model sit at the edge between "noticeably worse" and Q1's "useless". I took a quick glance at Bonsai's blog posts, and don't really see them comparing themselves to "typical" quants or explaining what's the special sauce that makes them better?

https://news.ycombinator.com/item?id=49611128

19m agoHN ↗

I think the general idea is naive quantization falls apart below 4bpw but you can go lower with more sophisticated QAT-adjacent methods. Bonsai's quantization method is proprietary though.