Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. An Empirical Study of Harness Design for Coding Agents(arxiv.org ↗)
    15comments
  2. Cloudflare Quick Tunnels(cloudflare.com ↗)
    21comments
  3. North Korean nuclear test sets off years of earthquakes(science.org ↗)
    5comments
  4. C++26: Trivial infinite loops are no longer undefined behaviour(sandordargo.com ↗)
    16comments
  5. Bend 2 and the Vibe-Coding Trap(liampwll.com ↗)
    181comments
  6. BeanShell3 in Development(beanshell.github.io ↗)
    3comments
  7. OpenJev(openjev.com ↗)
    187comments
  8. The Shadows Lurking in the Equations – Underwater Islands(gods.art ↗)
    6comments
  9. I Vibed a Proof of Conway's Conjecture(overreacted.io ↗)
    43comments
  10. I don't like passkeys(hawksley.dev ↗)
    350comments
  11. NATS publishes preliminary report on technical incident of 8 September(nats.aero ↗)
    4comments
  12. Jemalloc 5.4.0(github.com/jemalloc ↗)
    63comments
  13. Cekura (YC F24) Is Hiring(ycombinator.com ↗)
    discuss
  14. Warren Buffett Steps Down as Berkshire Chairman, Names Son to Replace Him(nytimes.com ↗)
    120comments
  15. The scourge of x86 emulation(fex-emu.com ↗)
    59comments
  16. Build Faster Feedback Loops Using Qualitative User Research(nseldeib.com ↗)
    discuss
  17. ZCode, the GLM coding agent, silently uploads your Git history(tokenstead.ai ↗)
    52comments
  18. Show HN: Rickub – The Smartest Git in the Universe(rickub.com ↗)
    discuss
  19. Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint(prismml.com ↗)
    170comments
  20. Astra for Law(openai.com ↗)
    647comments
  21. Microsoft exec called AI scraping 'the largest theft of labor in human history'(techcrunch.com ↗)
    514comments
  22. Replacing Pull Requests with Delta(zed.dev ↗)
    52comments
  23. Bend – A language that blocks AI mistakes via proof, on CPU and GPU(bend-lang.com ↗)
    279comments
  24. Mathematicians Build Long-Awaited Graph Sandwich(quantamagazine.org ↗)
    discuss
  25. Qwen 3.8 Omni Flash(qwen.ai ↗)
    103comments
  26. Subnormal floating-point numbers are expensive on Intel processors(lemire.me ↗)
    40comments
  27. Second Circuit Allows Government to Search Electronic Devices at the Border(knightcolumbia.org ↗)
    19comments
  28. When the fractional part of a float fixes your shader(crocidb.com ↗)
    15comments
  29. How to Write with an LLM(sockpuppet.org ↗)
    171comments
  30. Pre-Greek: The lost language hidden within Ancient Greek(linguisticdiscovery.com ↗)
    57comments

Subnormal floating-point numbers are expensive on Intel processors

53 pointsby 2d agolemire.me
40 comments
3h agoHN ↗

...Is this running extra micro code to fix some hardware bug/unreliability? How can this happen? Doesn't look like a normal design decision.

3h agoHN ↗

Subnormal numbers have a different, basically fixed-point, representation. They exist in order to bridge the large (relatively speaking; indeed "infinite" in a sense) gap between the least positive normal number, zero, and the greatest negative normal number, caused by the usual significand-exponent representation.

Most "mundane" uses of floating point have no need for subnormal numbers, and results that underflow could just be flushed to zero. But they’re sometimes important in scientific computing to ensure sufficient smoothness around zero, avoiding precision issues.

1h agoHN ↗

I don’t know how useful they are in scientific computing either, really. They are less precise than normalized numbers… if flushing them makes a difference I think it is a bad algorithm smell.

1h agoHN ↗

I can think of one useful property of subnormal numbers off the top of my head. If subnormal processing is enabled, then for all finite values of `a` and `b`, `a != b` if and only if `a - b != 0`. But if subnormals are flushed to zero, then two tiny normal distinct values `a` and `b` would have a subnormal difference that is flushed to zero.

1h agoHN ↗

Isn't that just a scale issue that exists with or without subnormals? If a and b are closer to zero than the smallest representable number, a and b compare as the same. With subnormals your smallest possible number is smaller than without, but it's still the same issue.

1h agoHN ↗

Because subnormals are fixed point, ie. have a fixed exponent, the difference of any two distinct subnormal values is nonzero like with integers.

56m agoHN ↗

But that same statement applies to normal values too, right? With normal numbers you might get the oddity of a-b -> a even if b is nonzero, but you don't get the oddity of a-b -> 0 unless the same number is represented, IIUC. A and B might not be bit identical, but they represent the same number if the difference is 0.

48m agoHN ↗

One nice thing that subnormals get you is the property that if x-y == 0 then x == y. If you want to guard against division by 0, and your denominator is a difference of two terms, it’s nice to be able to check equality of those terms and know that if they are not equal, then their difference will not be 0.

More generally, subnormals are needed for Sterbenz Lemma to hold everywhere: https://en.wikipedia.org/wiki/Sterbenz_lemma

3h agoHN ↗

I don’t know if any bugs contribute to this but this in the intel case but it has been very common historically for subnormal performance to be lower on many processors, and things like the Alpha required you to handle them in software if the COU fired a trap.

Have a look at https://en.wikipedia.org/wiki/Subnormal_number for some context.

2h agoHN ↗

It's to satisfy IEEE 754 and it's been this way for decades.

2h agoHN ↗

Does that mean that the ARM processors in the writeup are not satisfying IEEE 754?

2h agoHN ↗

It probably means Apple spent the silicon to handle subnormals at full speed in hardware, rather than triggering a slow microcode handler for such numbers.

1h agoHN ↗

I don't know who is down voting you. AFAIK IEEE 754:2008 does require support for subnormals. You can optionally have modes that flush them to zero, but you must support subnormals.

I haven't done any work on this stuff since 2019, so my memory may be hazy.

2h agoHN ↗

Apparently it only happens on P-cores, recent E-cores have a fast path for subnormals.

2h agoHN ↗

That seems weird. They did throw extra hardware at it to speed it up for the efficient cores, but not for the performance cores?

1h agoHN ↗

Probably different teams. Plus you can not underestimate the role that momentum plays in semiconductor engineering teams. If some respected person determined that subnormals are either Hard (tm) or not a Real Problem (tm), it will take a long time to correct this mistaken belief. I believe there have been some recent academic papers on FP implementation from the Intel E-Core team, which is a sign that they are a bit more with the times.

30m agoHN ↗

larger SIMD ALUs, which also have been around for longer?

2h agoHN ↗

If you don't _need_ subnormals MXCSR.DAZ/FTZ (which you can get gcc to set via -mdaz-ftz) will let you ignore all of this.

2h agoHN ↗

IIRC intel’s compilers enable FTZ/DAZ, at least at higher optimization levels.

1h agoHN ↗

GCC does with the infamous -ffast-math as well.

31m agoHN ↗

IIRC that has been since split out of the flag and need to be asked for separately at link time.

21m agoHN ↗

If you use -ffast-math when linking an executable but not a shared library, both gcc and clang will link in crtfastmath.o which has the bit of code to set the DAZ/FTZ flags.

2h agoHN ↗

I'm still trying to understand what a subnormal number is; IE, I'm looking for the TLDR so I know just enough to know if I'm using them and need to learn more.

Unfortunately, the Wikipedia article, while probably being accurate, doesn't give a clear and concise answer.

IE, is 0.0001 a subnormal? Or is it 0.000000000000000000001?

1h agoHN ↗

Usually IEEE floats have an implied 1 in the front. So for the standard represented numbers, there's some minimum number 1.bbbbbb.. * 2^-N. This allows 1bit more precision than is actually stored.

between any two numbers, there's basically the same epsilon difference, but from the smallest number to zero it's bigger.

A subnormal number breaks that convention, it just becomes 0.bbbbb... * 2^-N. As the numbers get smaller, the relative difference between the numbers gets larger. That also means their precision is smaller than the normal floats.

1h agoHN ↗

Floating point numbers are usually interpreted as

    sign * 1.mantissa * 2 ^ exponent

where sign, mantissa and exponent are fixed bit width integers. The 1. before the number is normally implicit because it would be a waste of a bit to encode it when you could just use a diferent exponent to represent such a number.

However with this simple scheme the number zero and a relatively large gap around it cannot be represented (relatively large to the gap between the smallest and next smalles number that can be represented).

So there is a special case where for the smallest encodeable exponent the mantissa must also specify that 1. or 0. prefix. Because its a special case it needs special handling that clever silicon engineers might think is unimportant enough to handle in microcode instead of dedicated silicon.

x86 has a mode to assume that all such small numbers are actually equal to zero which can then be handle without microcode fallback. Technically its even a bit more complicated because x86 has two different float implementations and for at least SSE floats you can control the denormals-are-zero and flush-(denormals)-to-zero-(when writing) modes independently. GCC -ffast-math actual enables that mode for the entire main thread.

AFAIK ARM NEON always works in that mode so the Gravion and Apple benchmarks might be unfair here undless you compare with DAZ and FTZ enabled on Intel. No idea if the AMD benchmarks might have used different modes. Because the flags are global per thread you can easily have unrelated loaded libraries messing the benchmark up.

1h agoHN ↗

That depends on how you store it.

Each number can be written in infinitely many ways, for example 12, 1.2E1, and 0.012E3 all are “twelve”

In (binary) IEEE floats, the canonical way to write floats is

  significant × 2^exponent

with 1 ≤ significant < 2. So, “twelve” gets stored as 1.5 × 2³ and not as, for example, 0.375 × 2⁵, 12 × 2⁰ or 96 × 2⁻³.

Float operations normally return numbers satisfying that.

However, in IEEE, the exponent cannot be made arbitrary small. Because of that, some very small numbers cannot be represented that way.

In those cases the standard says operations can return numbers with the value closest to the correct value with a significant less than 1. Those number representations are called subnormals.

50m agoHN ↗

For 32-bit floats, subnormals are the numbers closer to 0 than 2**(-126) == 0.0000000000000000000000000000000000000117549. For 64-bit doubles, it's 2**(-1022), a number starting with 308 decimal zeroes.

44m agoHN ↗

The 32-bit subnormals are all the non-zero 32-bit floating point values between but not including -0.000000000000000000000000000000000000011754943508222875079687365372222456778186655567720875215087517062784172594547271728515625 and +0.000000000000000000000000000000000000011754943508222875079687365372222456778186655567720875215087517062784172594547271728515625

Does that help you?

[Edited: correct decimal after noticing that my calculator defaulted to the wrong setting]

[And again because I think there's a bug in the last few digits, so debugging that's a fun activity for the weekend]

[And a third time because nope, those were correct and I can't type]

14m agoHN ↗

Binary floating-point numbers are scientific notation except the pieces are all in binary. In proper scientific notation, the only time the digit before the decimal point can be 0 is when the number itself is 0. Since the only other digit in binary notation is 1, there is no need to store the digit before the decimal point, since it's always 1... except now you can't store 0.

This problem is fixed by reserving one of the exponents for the representation of 0. Some of the formats (e.g. VAX floating point) that introduced this implicit-1-bit for the binary format said that every number with this special-0-exponent was a zero. But IEEE 754 introduced the concept of gradual underflow, and says instead that it is a bit string with the implicit digit before the decimal point as a 0 instead of 1.

Putting it differently and more succinctly: a subnormal number is a number that has fewer digits of precision than is normally implied by the format. Which numbers are subnormal numbers is entirely dependent on the floating-point format.

2h agoHN ↗

This has been the case since a zillion years, since the Core 2 Duo days at minimum.

2h agoHN ↗

The interesting part is that this seems to be Intel-specific.

1h agoHN ↗

That's what I mean though- in the Core 2 Duo vs K8 / Athlon days, Intel was much slower for denormals and subnormals than AMD. I'm not sure why but I thought this was common knowledge.

27m agoHN ↗

This is true on most hardware. Running ftz on H100s or B200 gives a free 10-20% boost for GEMM-epilogue workload

1h agoHN ↗

Are the results compared across architectures?

39m agoHN ↗

I mean the values of the computation not the runtime. I don’t know enough about ARM to say if doubles simply punt denormals to 0 for example.

28m agoHN ↗

IEEE 754 defines bit exact results for a lot of FP operations, including denormals.

18m agoHN ↗

And yet floating point math in general is non-deterministic across different CPUs. IEEE 754 was not good enough, so it's a valid question.

1h agoHN ↗

It used to be quite normal to add a low level random signal to inputs when writing DSP code so as to avoid dropping into subnormal territory. Careful analysis of the algorithm would identify any points where this was also necessary (e.g. feedback paths when running delays).

Obviously those lucky/unlucky enough to be writing 56k fixed precision code wouldn't have this concern, but other ones instead :)

I think flush to zero is probably the preferred strategy these days.