Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Ngrx Inspector(visualstudio.com ↗)
    discuss
  2. Reflections on Trusting Trust, Revisited: Poisoning Self-Modifying AI Coding(arxiv.org ↗)
    discuss
  3. Hacker
    2comments
  4. Anderon (IBM) finalizes $1B CHIPS award for US quantum foundry R&D(ibm.com ↗)
    discuss
  5. A Bitter Lesson for Data Filtering(arxiv.org ↗)
    1comments
  6. Mojo 1.1: the compiler now accepts contributions(modular.com ↗)
    discuss
  7. Show HN: Continuity – project state your AI sessions can't silently contradict(github.com/vikcena01 ↗)
    discuss
  8. Show HN: Free Tournament Bracket Maker and Generator(snapbracket.com ↗)
    discuss
  9. macOS 27 is a lifesaver for killing leftover AI Agent processes(9to5mac.com ↗)
    2comments
  10. The Provenance Tax: How LLM Watermarking Changes AI Agent Behavior(lasso.security ↗)
    discuss
  11. Waymo in Singapore(waymo.com ↗)
    1comments
  12. The Concerns Surrounding Rapid Rise of AI and What We Should Do About It(medium.com/f9121212 ↗)
    discuss
  13. Age of Beyond [video](youtube.com ↗)
    discuss
  14. Natural General Intelligence(naturalgeneralintelligence.ai ↗)
    discuss
  15. Youper is shutting down – What kills mental health AI?(youper.ai ↗)
    discuss
  16. Show HN: Free GitHub Action that scans PR diffs for malicious code, not quality(github.com/marketplace ↗)
    discuss
  17. The Bitterest Lesson(typesafe.ai ↗)
    discuss
  18. Google announces new experimental "CC" AI agent for families(arstechnica.com ↗)
    discuss
  19. Show HN: Green Screen Remover – free in-browser chroma key, nothing uploaded(greenscreenremover.net ↗)
    discuss
  20. I built the fastest PHP webserver in the world(qbixserver.com ↗)
    discuss
  21. Data
    discuss
  22. Running OpenBAO on Kubernetes with a CloudNativePG PostgreSQL Back End(cncf.io ↗)
    discuss
  23. Show HN: Concat: the open-source CapCut replacement.(github.com/jub0t ↗)
    discuss
  24. Pre-Greek: The lost language hidden within Ancient Greek(linguisticdiscovery.com ↗)
    discuss
  25. Code Scans(devin.ai ↗)
    discuss
  26. Warez: The Infrastructure and Aesthetics of Piracy Culture(archive.org ↗)
    discuss
  27. Show HN: Jinx – a link shortener you host free on GitHub Pages(jinx.fyi ↗)
    discuss
  28. Amfly: A deterministic simulation of six fruit fly connectomes in a reward loop(github.com/aravpanwar ↗)
    discuss
  29. Alternatives to milk are losing moomentum(economist.com ↗)
    1comments
  30. Show HN: An open Add/Search evaluation framework for agent memory(agentmemoryleaderboard.ai ↗)
    discuss

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

296 pointsby 6h agoprismml.com
95 comments
6h agoHN ↗

Love this for the folks with 16gb graphics cards - 3.8 27b has been incredible but not quite runnable on anything less than 32gb - will try loading this up on my 16gb intel b50 and see how it goes - not sure these quants can be accelerated by the XPU cores yet but maybe in time!

6h agoHN ↗

You can run the ~4 bit quant(s) on 24gb, if you're not _too_ picky on context size.

This will hopefully be better, though it'd be a _very_ surprising increase in performace at the size they say. Would love to see more about how it benchmarks.

6h agoHN ↗

I run Unsloth's UD-Q4_K_S on 20 GB of VRAM (RX 7900 XT) and I get ~90k tokens of context without quantizing KV cache. With 8-bit quantization, I get about a 134k token context window. That's with only one slot, but for me, it works pretty darn well, with 20-35 tok/s depending on how full that window is.

3h agoHN ↗

7900 XT is a sleeper card. When I initially bought it, it was priced at the lowest wattage per $ per GB VRAM (not normalized for token speeds...) Although I ended up swapping for the XTX because that 4GB means everything in just increasing the context window. At 8bit KV my window is over 200k, and although qwen3.8 loves vomiting out tokens as part of its reasoning chain I trust it enough to get assigned tasks done eventually, which I could not say of any model before its release.

27m agoHN ↗

How has software/driver support been? I got burned hard by AMD last generation or the one before. Things smoother now, or do you have to baby it like hell and pick and choose software that works?

6m agoHN ↗

I don't do anything fancier than inference, and I only use llama.cpp, which supports rOCM. I've had few issues; most GGUFs I download work right out of the box. Nearly any popular model has a quant that just works. But as you can see I don't use my GPU for anything weird or nonstandard.

2h agoHN ↗

I'm doing the same with a context of about 128-150k Surprisingly, I get subjectively better results with Unsloth's 3 bit quants (UD-Q3-XL something), than their 4 bit quants (S or M)

1h agoHN ↗

You can trivially run 131k on 24GB 4bit, and there are repos with tweaks that allow you to get the full 262k but idk if there's degradation with their approach.

2h agoHN ↗

Tried it today on a B70 and couldn't get anything usable out of it.

Prism's llama.cpp fork only has the kernels for CUDA, CPU and Vulkan. No SYCL at all :(

6h agoHN ↗

I'm not following the local mdoel scene too closely but this seems quite amazing. Is this able to be run on Apple silicon too?

6h agoHN ↗

"Ternary Bonsai 2 27B reaches up to 143 tokens/second on NVIDIA GeForce RTX 5090 and 46.8 tokens/second on M5 Max. On an RTX 4090, Ternary Bonsai 2 27B consumes just 0.714 mWh/token, making it 40% more energy-efficient than an 8B model running in full-precision."

6h agoHN ↗

Their mention of the 5090 is bit odd, since on 32 GB GPUs, Q6 fits while having better quality. Very interesting model for 16 GB GPUs though!

5h agoHN ↗

They mention 5090 with regards to speed, Q6 will not have that speed?

And speed matters a lot for many use cases

5h agoHN ↗

150 tokens per second on a ternary model implies that it’s GPU bound, I’d bet a Q6 model is even faster because it’s existed longer and seen more optimization. You’d have to be insane to not run an NVFP4 quant over a ternary quant on Blackwell if they both fit.

5h agoHN ↗

it is like "my fridge is 2mkm (millikilometer) from my desk" m=0.001 h=3600 it should be just Ws or just J

6h agoHN ↗

Their first 27B bonsai was able to run on an iphone.

3h agoHN ↗

Sadly crashing on Pixel 9 Pro, but I guess phone GPU with 16GB RAM total wouldnt be enough anyway.

3h agoHN ↗

Even native AI gallery uses smaller Gemma models. I guess whatever Chrome is using for WebGPU compute on Android is just adding too much overhead.

6h agoHN ↗

I'd love to see a Bonsai model start with a 100B+ parameter model and get that down to <30 GB. But maybe at that point we call it Topiary?

4h agoHN ↗

Other names that occur to me: Orchard, Forest, Stand (of trees).

5h agoHN ↗

I'm hoping they release an 8B v2 based on the Qwen 3.8 series in the near future - that would give us a really powerful model that could be run directly on users phones.

5h agoHN ↗

Yes! And maybe get a hf fused webgpu runner for that model so it’s fast!

4h agoHN ↗

That would require Alibaba releasing a Qwen 3.8 8B first

5h agoHN ↗

If you want to try out out the GGUFs from https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf#th... be aware that you need Prism's llama.cpp fork to get them to work, from https://github.com/PrismML-Eng/llama.cpp/releases/tag/prism-...

This should work:

  cd /tmp

  # Get the Prism macOS runtime
  curl -fL https://github.com/PrismML-Eng/llama.cpp/releases/download/prism-b10685-7dffb15/llama-prism-b10685-7dffb15-bin-macos-arm64.tar.gz -o bonsai-runtime.tar.gz
  tar -xzf bonsai-runtime.tar.gz

  # Get the ~5.95 GB GGUF model:
  curl -fL https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/resolve/main/Ternary-Bonsai-2-27B-PTQ1_0.gguf -o Ternary-Bonsai-2-27B-PTQ1_0.gguf

  # Run the server, I used port 8331
  ./llama-prism-b10685-7dffb15/llama-server \
    -m Ternary-Bonsai-2-27B-PTQ1_0.gguf \
    --port 8331 -ngl 99 -fa on -c 32768

Then open http://localhost:8331 for the (very good) baked in llama-server web UI... or run a prompt via the API like this:

  uvx llm openai endpoint http://127.0.0.1:8331/v1 \
    --model bonsai-2-27b --responses hi

That's running at ~20 token/second for me on an M5 Pro (after a server restart I got 44 token/second, not sure why), but I'm pretty sure something isn't working right, on startup the server said "ggml_metal_device_init: - the tensor API is not supported in this environment - disabling".

5h agoHN ↗

Honestly looks pretty good except whatever is going on with its booty. Is that an ass helmet? I cannot parse what's going on there.

4h agoHN ↗

I like the lens effect behind the rear tire.

8m agoHN ↗

M1 Pro, same prompt, same cli options:

  32,706 tokens
  38min 19s
  14.22 t/s
5h agoHN ↗

I figured them out, starting from the GGUF on Hugging Face.

If you have found better instructions and they work then use those instead!

Personally I prefer to download models directly rather than running some `./setup.sh` script where I need to then review what it does first.

4h agoHN ↗

Yeah just wanted to mention in case it explains the 2x lower throughout you are seeing on M5. To be fair their documentation is a bit inconsistent in some spots.

Would be good to know if the release and weights from their demo repo work better. I’m trying on a 4090 and will report back.

4h agoHN ↗

It would be great to have upstream llama.cpp support for this!

2h agoHN ↗

Agreed. They always sound exciting to try out but are such a pain to get working.

1h agoHN ↗

I always just throw an agent at it. Is this the RSI I keep hearing about

3h agoHN ↗

If you want to download the gguf to your regular huggingface cache directory instead of to /tmp, you can download the model and run the server in one step:

  export HF_TOKEN=xxx # optional, speeds up the download
  
  ./llama-prism-b10685-7dffb15/llama serve \
    -hf prism-ml/Ternary-Bonsai-2-27B-gguf:PTQ1_0 \
    --port 8331 -ngl 99 -fa on -c 32768
2h agoHN ↗

Thanks for all of your exploration in public Simon.

Commenting because the fix I proposed was merged in roughly 49 commits after the PrismML Fork. The “tensor API is not supported” warning occurred because llama.cpp’s startup probe fails to compile a matmul2d kernel: Metal’s tensor headers require language version 4.0, but ggml-metal-device.m previously omitted MTLCompileOptions.languageVersion, disabling the API universally.

Here’s a link to the diff if you want to try and update that fork to take advantage of the prefill gains afforded by the hardware: https://github.com/ggml-org/llama.cpp/pull/27461/changes

5h agoHN ↗

I really wish people would stop saying N times smaller than something when making a comparison; that makes no sense - it's 1/9th (11.11%) the size. You don't get a smaller quantity by multiplying by a number greater than 1.0. You could instead reverse the subjects being compared - "the original model is 9x bigger than this new smaller, efficient model" or some such. That makes sense.

I keep seeing this being used when people talk about efficiency or performance gains and it's just very unintuitive language.

5h agoHN ↗

If we were talking about speed instead of size, i think it would be perfectly reasonable to say 9x faster. I'm not sure I agree that 9x smaller is unintuitive. It makes sense to me.

4h agoHN ↗

"Nine times" literally means multiplied by nine, but here we're dividing by nine. It's not unintelligible (because the corrupted verbiage is so commonplace) but it is needlessly awkward. Like saying "resulted in a size reduction increase of 10 megabytes."

4h agoHN ↗

"faster" relates to speed. Speed is related to time and speed of a thing is usually defined by time. 9x faster speed translates to time/9. There is an extra step of related conversion there.

Conversely, 9x [filesize/natural number] is bigger. Every time. At least in the basic maths used by most people. There is no conversion into other units.

Therefore "9x smaller" when talking about a natural number like filesize is a nonsense statement in logic terms. If you strive for unambiguous phrasing - which is a significant part of the programming experience - this logical nonsense might well perturb you.

But english language is a flexible thing and if the phrase communicates your intent to your audience then that's fine by me.

5h agoHN ↗

Me too!! It's a huge pet peeve. And it's so hard to get people to see how it's linguistically AND mathematically WRONG.

18m agoHN ↗

I disagree completely and I'd be happy to be convinced otherwise

4h agoHN ↗

I agree it makes little sense in a literal mathematical take but "it's 9x smaller" or is too much of linguistic advantage compared to "the original is 9x larger" or "it's 1/9th as large" to expect a change with. You don't have to invoke fractions, it keeps the thing in focus as the first subject, it matches the pattern of the inverse statement, and it's just plain short... so that's what people will adapt and interpret the meaning to be.

One other way to map both types of linguistic statement consistently to math is to interpret "9x" as "there is a 9 times difference between these two things" and then "smaller"/"larger" tells you which end of that separation the subject is (rather than specifying whether the multiplication builds up or down).

1h agoHN ↗

"it's 11% as large" avoids fractions but is much more clear (IMO) than "9x smaller" which my brain doesn't understand.

1h agoHN ↗

11% is also a really saying a fraction, 11 per-cent or 11/100, but there's nothing wrong with that feeling more natural to some and it is at least a nice shorthand way for the written form. "A ninth the size" is a similar alternative. All are really fine, there's always someone who has trouble with a given representation compared to another.

3h agoHN ↗

Eh, when I read smaller with an integer multiplier, I mentally switch to the reciprocal. Easier than convincing the world not to use "9x smaller". Do you feel the same way about "9x faster"? What you're actually measuring is time, and "faster" is the reciprocal of time, similarly to "smaller" being the reciprocal of size.

2h agoHN ↗

yes, actually I do think the phrase 'N times faster' is sensible and logical

if I say 'this Apple M5 chip is 3x faster' than this intel chip, it implies two things:

- the run time of most operations that runs on it is now reduced (so one quantity is smaller)

- but also: MORE WORK is being completed per unit of time compared to the intel chip (so this quantity is greater)

So yes, a greater quantity is being measured in the apple chip compared to the intel when you say apple is N times faster. i guess this a quirk with the word 'faster' - it actually measures two things, time and work performed per unit of time. the word smaller just measures size.

like, if I give a customer a cup of coffee one day, then give them a SMALLER cup of coffee the next day, for the same price, but declare "It's now 2X more space efficient!" i.e. it's now half the size, I'm certain the customer is gonna be pissed.

1h agoHN ↗

Ah, but in this case the customer is buying the coffee for the caffeine content. And you've doubled the caffeine content per volume! Twice as efficient a delivery mechanism, similar to this model!

1h agoHN ↗

I completely agree and this is a pet peeve of mine so it's nice to be validated :]

1h agoHN ↗

But we're cutting drug prices 500, 800, 1700%! Numbers nobody thought were possible.

1h agoHN ↗

It's simple

If 9 is "9 times greater" than 1 then 1 must be "9 times smaller" than 9

It'd help if you read "9 times" with the operator which is what's being flipped instead of with the number

31m agoHN ↗

They probably rephrased it from some more technical form like "we compressed the model by a factor of 9" or "we've improved the packing efficiency of the model by 9x". Where these are measurements of the transformation the model is undergoing, not measurements of the resulting model.

19m agoHN ↗

This sounds like the same kind of error as writing "0.10 cents" because it's less than a dollar when the number is in dollars regardless of how big or small it is

16m agoHN ↗

This is widespread usage.

I don't see what makes it hard to understand.

5h agoHN ↗

I think if they made this for Qwen3.8-Next it could fit in a single 5090?

4h agoHN ↗

Came to ask the same. From my really rough understanding, it seems like Unsloth's method allows a slightly higher precision at a higher file size, while PrismML's uses a different approach to achieve a smaller size (and presumably less precision).

1h agoHN ↗

Oh, wow, they think it's just a smidge below the q4? That's crazy good if true.

1h agoHN ↗

The benchmarks they chose are rather cherry picked to not include long context or difficult ones that involve long horizon work or many agent turns, as I suspect this is where the model shows more differences compared to the full fat one

5h agoHN ↗

Ternary Bonsai 2 27B uses ternary {−1, 0, +1} weights with FP16 group-wise scaling, for 1.76 effective bits per weight

If I recall correctly, a recent post [1] has shown that Q2 quants (with like 2.6 bpw) of the same base Qwen model sit at the edge between "noticeably worse" and Q1's "useless". I took a quick glance at Bonsai's blog posts, and don't really see them comparing themselves to "typical" quants or explaining what's the special sauce that makes them better?

https://news.ycombinator.com/item?id=49611128

5h agoHN ↗

I think the general idea is naive quantization falls apart below 4bpw but you can go lower with more sophisticated QAT-adjacent methods. Bonsai's quantization method is proprietary though.

4h agoHN ↗

That's correct. I think there can certainly be issues even with 4bpw with naive quantization (IE you'll notice far better results from a QAT 4bpw vs a naive 4bpw).

One such method that I've been meaning to look into further is Tencent's AngelSlim QAT/PTQ approach. They did a Hy4 preview release thats an STQ_1_0 at 2.38 bpw:

https://huggingface.co/AngelSlim/Hy4-preview-GGUF https://arxiv.org/abs/2602.21233

Of course, it's still 213g of VRAM I'd need so it's somewhat out of the range of what I can run locally. In contrast, this new Bonsai is nice because the original was already exciting for making use of low VRAM devices. Could breath new life into some of the older GPUs that were previously close to top of the line just quite VRAM constrained by modern standards and still quite cost effective for now.

3h agoHN ↗

Could've been better if GGUF implemented QTIP format. GGUF representation is a major limitation for llama.cpp quantization performance

3h agoHN ↗

They use their own llama fork anyway, so that shouldn't matter.

4h agoHN ↗

1.76 bpw number is kinda misleading if you compare it directly to IQ2/Q2. The encoding is ternary, but the quantization procedure is way more sophisticated than "round Qwen weights to {-1,0,+1}."

They rotate the weights into a quantization-friendly basis first, then ternarize with per-group scales and error compensation.

4h agoHN ↗

Cautiously optimistic. The V1 was noticeably weak on world knowledge but here the 3.8 base model is geared more towards reasoning than world knowledge anyway so might not matter as much

4h agoHN ↗

Running at about 7-8 tok/s (~60 tok/s prefill) on a Mac Mini M2 16GB.

So far feels smarter than Bonsai 1 27B, it’s slightly larger than the Q1_0 quant. Super exciting stuff :)

4h agoHN ↗

GPT Astra did some benchmarking on the DGX Spark. Speed: 34.38 tokens/sec for generation.

Seems like we don't have a drafter model yet so it could not test with speculative decoding on. ngram speculative decoding did not help too much either - not enough accepted tokens.

Smaller size I suppose does not mean better performance in this case - we maybe limited by Spark's low memory bandwidth.

3h agoHN ↗

450 with PTQ_01 and 900 with the other PQ2_0.

1h agoHN ↗

That's a rough place to land on a spark. It seems unlikely to be memory bandwidth at this model size, but maybe just lack of tuned kernels? The chip is missing some CUDA features but with tuning you should be able to hit way more than that even without a drafter.

4h agoHN ↗

I wonder how their talks with Apple went. Having this run on the TPU opposed to just the GPU, which drains a significant amount of battery life by comparison, is what I'm really interested in.

4h agoHN ↗

What I'd love to see is this done for DS4.1 Flash.

That would bring it down to the point where it can fit in 128GB on things like the Spark or Strix Halo.

3h agoHN ↗

RAM requirements? My current rule of thumb is “a byte per parameter”, but I doubt this runs in 1/9th that (~ 3GiB).

Also, perf speedup?

3h agoHN ↗

I'm seeing around 7.9GB of ram, 120 tokens/s on a 6000 pro blackwell.

3h agoHN ↗

They need to make a Big Bonsai, something at the enterprise levels that can compete with DSV4 Flash etc.

2h agoHN ↗

What is never totally clear with a lot of these releases is the scope of what it's good at. Models that can run with good speed on affordable consumer hardware for coding only is the dream. I am never going to use this for writing, images, or "general knowledge". Coding only

2h agoHN ↗

I tried their WebGPU version and it immediately started looping. Yeah "near lossless" my ass. Plus the reasoning that it looped on was clearly wrong and unlike the non quantized 27B

1h agoHN ↗

Never heard of Bonsai before, but that looks great and promising for local on-device inference.

Yet, seems like there is still another year for improvements.

I like local models (but not mainly using them) for offline needs.

59m agoHN ↗

LLM quants seem to eerily converge to modern/not so modern graphics techniques. You wouldn't think it would apply but it's obvious in hindsight. In fact mining graphics ideas is probably a good inspiration for efficient LLM architecture.

For example, the Hadamard activation transform used here feels a lot like multiplying Fourier basis ala DFT; strong parallels to how image codecs work to make the residuals more compressible (especially discrete block codecs like are used in GPU compressed textures).

I thought I was being clever suggesting that you could even abuse texture decode units to efficiently sample compressed LLMs with hardware; turns out Apple foundation models are already doing this [1].

[1] https://arxiv.org/abs/2507.13575

53m agoHN ↗

I tried it on my 6GB GPU and got 0.67 tokens per second. Need more than 6GB to run it well.

8m agoHN ↗

Testing on a MBP m4 pro 24gb

~100t/s prefill, ~15t/s, dropping to ~10t/s later with 64k context.

The issue is I have yet to find a useful agentic local llm that I can run on this machine.

Just given a relatively simple task on a swift app, took 25 minutes, brainstorming like crazy but can not decide on what to do. Eventually I killed it. GPT 5.6 sol-medium took 3 minutes to complete the same task for reference.