Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Platform-Independent SIMD in Go (go.dev)
    86comments
  2. Git-bug: Distributed, offline-first bug tracker embedded in Git (github.com/git-bug)
    62comments
  3. First Principles Thinking (sunilsadasivan.com)
    40comments
  4. U.S. appeals court upholds designation of Anthropic as supply chain risk (cnbc.com)
    228comments
  5. Classified Estimates Show the NSA Is Paying Billions to Test AI Models (washingtonsun.com)
    58comments
  6. Pentium II at 600Mhz with Voodoo 3 Emulated on 86Box with M6 Mac Mini (nyaa.sh)
    91comments
  7. F-Droid 2.0 (f-droid.org)
    399comments
  8. Ink and Switch Interactive Homepage (inkandswitch.com)
    21comments
  9. Dutch governments builds alternative for Microsoft based on NixOS (dawo.community)
    488comments
  10. Factorio that you can touch (factorio.com)
    22comments
  11. Gravity Seems Holographic. What Does That Mean for Reality? (quantamagazine.org)
    46comments
  12. Show HN: Make cursed fonts like Times New Bastard (mitpit.com)
    116comments
  13. Amiga Screens: A Primer (datagubbe.se)
    18comments
  14. Jevmem – automatic project memory for Claude Code, built on Jev (github.com/avinash-jetwani)
    33comments
  15. Allow Carriers on Planes (jefftk.com)
    182comments
  16. Boards of Casio (ambionix.com)
    20comments
  17. Show HN: Whiteboard (YC W26) – An open-source IDE for thoughtful software design (github.com/devdotfast)
    126comments
  18. CVE-2025-13032: Entering and Breaking the Avast Antivirus Sandbox Part 2 (safateam.com)
    26comments
  19. What About Rails? (jardo.dev)
    139comments
  20. Why is the liver so weirdly regenerative? (dynomight.substack.com)
    262comments
  21. Microsoft Abandons Personal AI Chatbot Race with Copilot Reboot (bloomberg.com)
    36comments
  22. The Test (tante.cc)
    43comments
  23. 2DWillNeverDie (2dwillneverdie.com)
    76comments
  24. Rails World 2026 Opening Keynote [video] (youtube.com)
    442comments
  25. Opus 5.5 is good at explainer videos (launchvideo.io)
    201comments
  26. Fearless SIMD v1.0 (linebender.org)
    48comments
  27. Toyota is taking the Corolla electric (electrek.co)
    747comments
  28. My weird new hobby: Wandering around Tokyo on Google Maps (ahmedhossamdev.com)
    171comments
  29. Two-tier encryption in the UK (macanorak.com)
    449comments
  30. Using LLMs to trace alchemical knowledge and decode 17th century letters (resobscura.substack.com)
    38comments

Platform-Independent SIMD in Go

237 pointsby 6h agogo.dev
86 comments
5h agoHN ↗

This feature opens many doors for optimizing low-level performance in Go projects, that are already running multicore. IIRC there aren’t a lot of languages with built-in std lib support for SIMD and variants. Love the way Go is trying new stuff lately.

5h agoHN ↗

Vectorizing computations has been Matlabs secret sauce.

4h agoHN ↗

Does matlab these days do stuff like JIT operator fusing to avoid memory roundtrips and take advantage of FMAs?

4h agoHN ↗

Besides the usual C and C++, we have Java, .NET, D, Zig, Julia, Swift, Rust.

So yeah, also appreciate having Go in the group instead of manually having to write Assembly.

However not many languages adopt ways to manually write SIMD, because most of us have no idea how to write good SIMD code in first place, I surely don't.

4h agoHN ↗

Even with languages that adopt ways to manually write SIMD, it’s mostly left to library maintainers rather than application developers.

I work for a C++ timeseries database startup that leverages SIMD about as much as we possibly can, and except for some extremely rare places we just use libraries.

4h agoHN ↗

Yeah, that is what I have heard from some NVidia folks as well, like Bryce Adelstein, use the libraries as much as possible, and leave the kernels for experts.

However even then, it depends on how the libraries API surface looks like.

4h agoHN ↗

With AI I'm pretty sure SIMD will be easier to integrate when necessary.

4h agoHN ↗

But it’s not necessary at all, the whole point is that these utility libraries bring you more elegant code that work on all platforms without having to pollute your codebase with SIMD intrinsics.

Unless this was tongue in cheek, because this is in fact a problem with AI that it degrades your codebase in these types of ways.

2h agoHN ↗

In 2026 if you are not doing A with AI you are doing it wrong /s

5h agoHN ↗

Oh this is great, it was one of my biggest bugbears about Go since you almost always have to link C/C++ code to get the appropriate performance.

The one negative I'd say is that often autovectorisation is 'good enough' and this doesn't really tackle that gap.

5h agoHN ↗

As a first step, it might be possible to write a linter rule that rewrites suitable numeric loops to SIMD. There are already rules to rewrite several loop types, so that should be doable.

4h agoHN ↗

The poor Assembler and the unsafe package forgotten in the corner.

While reaching out to CGO is the easier way, it doesn't mean it is the only tool available in Go.

4h agoHN ↗

FWIW, there is some pretty substantial autovectorization work that is already in-flight for the Go compiler.

There's a CL stack here:

https://go.dev/cl/791740

It's hard to make predictions with an open source project, but my personal guess is some flavor of it will land (including it is already demonstrating good results without an enormous level of code complexity in the compiler and without overly slowing down compile speeds), but I guess we'll see.

It's being driven by an external contributor who has landed some good changes in the past to the Go compiler. (I think the autovectorization work might be part of their PhD or other academic research, but not sure.)

4h agoHN ↗

The problem with Go isn't performance but with the C/C++ interop overhead, even with the "30% less overhead" from a few updates ago which isnt true for 99% of cases, it isnt enough

3h agoHN ↗

Why is that the case? I don’t know low level programming so why is Go limited in interop with C?

2h agoHN ↗

its not limited but it has overhead because of the memory model of go doesnt match the C one so there has to be some sort of rerodering being done, that's what i understood atleast, and theres also the go concurrency

2h agoHN ↗

Use Assembly instead of CGO, isn't that scary, back in the 8 bit days we were coding Assembly aged 10, on our Spectrum, C64, Atari, Apple, Acorn, MSX,....

4h agoHN ↗

Already using this for foreground estimation of cutouts in my project, around 30% speedup over non-SIMD, but the algorithm is probably not very optimised yet.

3h agoHN ↗

This is why I love Go. Nobody was asking for this, but they took the time to do it right and continue to Push go as a memory safe, high-level systems language.

2h agoHN ↗

That does not make sense to me. Go is memory-safe, but it does not guarantee data-race freedom.

So whats your point here? Haskell?

2h agoHN ↗

Probably Rust, that's always Rust with this kind of comments...

1h agoHN ↗

Or basic software engineering and understanding of what memory safe means.

1h agoHN ↗

There is no memory safety without freedom from data races. One is a prerequisite of the other. This is why languages like C# throw exceptions on unsynchronized concurrent access to some container types, and treat all property accesses as atomic.

1h agoHN ↗

Are you saying any language that does not promise data-race freedom is memory unsafe? That would rule out almost every programming language.

1h agoHN ↗

I am, but you’d be surprised. All of the single-threaded languages are fine, for example. Very few languages are actually low-level enough to allow data races. C# and Java go to great lengths to avoid it.

Race conditions in general are another matter, and aren’t generally considered a requirement (though you can certainly create nasty bugs).

1h agoHN ↗

I have not written a line of Java in 15 years. But im pretty sure Java has threads? Once you have threads, you pretty much have data races.

        Thread a = new Thread(() -> x++);
        Thread b = new Thread(() -> x++);

        a.start();
        b.start();
1h agoHN ↗

From the former head of the Go security team [1]:

I have never seen real Go code (i.e. not code written purposefully to be exploitable) that was exploitable due to a data race.

And from tptacek in that same discussion [2]:

The fact is that Go doesn't admit memory corruption vulnerabilities, and the way you know that is the fact that there are practically zero exploits for memory corruption vulnerabilities targeting pure Go programs, despite the popularity of the language.

[1] https://news.ycombinator.com/item?id=44672003

[2] https://news.ycombinator.com/item?id=44672371

1h agoHN ↗

Memory safety isn’t really about vulnerabilities. This thread is about Go, so I won’t go further into it here.

1h agoHN ↗

I suppose that if you create a map in one thread, and then access it from another thread, then that might cause segmentation faults? Because map is a type that is implemented in C.

1h agoHN ↗

That does not make sense to me.

You said that because you assumed Go is memory-safe in all conditions.

Go is memory-safe

Yes, but only if there's no data race.

Go is not like Java. Java doesn't guarantee no data race, but when it happens, it's still memory-safe.

1h agoHN ↗

Neither Java or Go guarantees data-race freedom. A data race does not by itself make ordinary Java or Go code memory-unsafe in the C/C++ sense.

I fail to see how a racy Java program is more memory safe than a racy Go program?

1h agoHN ↗

Java arrays are thin pointers to Array objects, which contain the length and data in the same place (on the far side of the pointer). Since thin pointers cannot tear and Array objects cannot be resized, Arrays themselves are always memory-safe in safe code.

Go slices are fat pointers to undecorated memory. The slice itself is a 3-tuple of pointer, length, and capacity. If you append to a slice that's already at capacity, the Go runtime will allocate new memory for you and return a new 3-tuple. If you assign that result to a variable that's also being accessed by another goroutine, the latter can observe the slice in an inconsistent state. It can, for example, see the old pointer but with the new length, allowing out-of-bounds access. None of this requires unsafe code.

The same issue applies to string and interface variables, which are also fat pointers.

50m agoHN ↗

Indeed.

Shared Go slices are a bad mix in concurrent code. This is a given. But its also not a fair comparison, you should instead compare java arrays to go arrays, not slices.

This goes for slices, strings and maps. Those a usually wrapped in a mutex, or used with sync primitives like sync.Map.

54m agoHN ↗

From what I understand a data race on a simple built in feature like an interface pointer can result in a bad address / type pair, which can cause memory safety issues on any future access. I don't think you can get the JVM itself confused about what type a pointer points to.

1h agoHN ↗

Yes, but in practice they are extremely hard to exploit. It has been discussed extensively here on HN and in other forums.

1h agoHN ↗

That's not what "memory safe" means. "Memory safe" is a term of art meaning "not susceptible to memory corruption exploits", like stack and heap overflows, UAFs, and type confusion. Last I checked, there are essentially no non-contrived memory corruption exploits for Go programs; the best you get are people demonstrating register control on contrived programs.

The definition I'm giving is the same as the ISRG's definition at MemorySafety.org. It's the thing everybody is talking about when they talk about memory safety.

The claim being made here is "big if true", because it would imply a lot more languages than Go "aren't memory safe", despite decades without memory corruption exploits.

1h agoHN ↗

Yeah, we get it, you performatively hate go.

2h agoHN ↗

Go is broadly considered to be a memory safe language.

See for example comments from tptacek like:

https://news.ycombinator.com/item?id=43335748

https://news.ycombinator.com/item?id=46028232

https://news.ycombinator.com/item?id=44672371

(The gist: memory safety is a term of art coined by security practitioners. Go, Python, Rust, Java, others: memory safe. C/C++: memory unsafe. Periodically, people in different slices of industry or academia come up with new definitions of memory safety that declare Rust or Go or other languages to be memory unsafe, but that is not by the broadly accepted definition across industry.)

2h agoHN ↗

Rust does allow you to overflow buffers, confuse types, and duplicate mutable pointers in safe code. See cve-rs.

1h agoHN ↗

No, Rust does not allow that. The current Rust compiler does, but that’s a bug that is being fixed.

At some point in the future, a fully backwards compatible Rust compiler will report an error when you try to compile cve-rs.

1h agoHN ↗

Since we are not at some point in the future where that correct compiler exists and there is only one official compiler, the distinction you make is practically meaningless!

54m agoHN ↗

Isn't one of the bugs around ten years old, now? Isn't ten years enough to call something a feature of the language rather than a bug?

I like Rust, but with this bug existing for so long, I personally no longer think of it as memory-safe.

18m agoHN ↗

I define undefined behaviour as a bug in C++. Now C++ is memory-safe!

btw, it's not actually that hard to write correct code in C++, easier than in C because you have all the container types. The problem is that nothing will tell you when you write incorrect code - there's no guarantee.

1h agoHN ↗

He’s very wrong about this. Just because ‘tptacek posts a lot and did security once upon a time does not make him “broad consideration”.

2h agoHN ↗

No idea why you're getting downvoted for true statement. Without a ? like in C# you're always at risk of a nil pointer being dereferenced

2h agoHN ↗

like in C# you're always at risk of a nil pointer being dereferenced

That throws a NullReferenceException

2h agoHN ↗

Cool, very clever. But at least the compiler warns me of a potential exception, whereas in Go no such op even exists.

2h agoHN ↗

you can dereference nil in Go and it panics

2h agoHN ↗

And then use Go's pseudo exception handling and recover.

2h agoHN ↗

This is not generally considered part of the "memory safety" contract. You can not lift a nil pointer exception into a replacement for Go's "unsafe" library.

When we finally rid ourselves of C and C++ is so larded over with extensions and additions and features that we can finally plausibly say the C subset is just not in use anymore, we can perhaps consider as a community expanding what "memory safe" means, but in the meantime it has some very important meanings and we should not try to augment the term. Memory safety doesn't mean anything like "forcing exhaustiveness into sum type deconstructions" or "never has a race condition" (though it does mean said race condition shouldn't be something that allows you to escape out of an array or forcibly change the type on something in a way the language doesn't normally permit) or any of several other things that may be very nice to have indeed, but are not part of the definition of "memory safe".

Memory safe is a very old concept, and almost everything is memory safe now. But not quite, and as such the term still has use. And also zig for some reason gave it up so it won't be disappearing as soon as I'd like.'

2h agoHN ↗

You can write unsafe code in Go (import unsafe), but then, you can do the same in Rust. Unsafe code is not the default, and in day to day Go i rarely see the use of the unsafe package.

2h agoHN ↗

What he probably means is data-races in go can result in memory/type unsafe accesses -- I suspect, likely due to slice types -- not sure if that is true/false.

1h agoHN ↗

Sure, but a data race is, IMHO not the same as memory safety. A data race, can be 100% memory safe, but just cause a logic bug in some program. I often see people mixing memory safety with racing. Go has bounds checks so you end up with a panic either way. Not UB.

As an (outside go) example, Ocaml (5) promises strong memory safety, but not to be data race free. A data race is not something we can prevent, because its usually not bound by code, but by time and the race-source rarely in source-code.

This means we have data races in http, database inserts etc. The source is usually not a concurrent task in source code-land.

1h agoHN ↗

That’s a race condition, not a data race.

1h agoHN ↗

Now you are pushing pixels. A data-race IS a kind of race condition.

My point is "races" happen all over. In concurrent code, databases, http and pretty much anywhere where you have some kind of timing, not scoped to a unit.

55m agoHN ↗

A race condition is an application invariant violation under concurrency, so it has no application-independent definition. A data race is unsynchronized access by two concurrent threads to the same memory location where at least one access is a write. Ergo, the definition of a data race has nothing to do with application logic. Is that distinction so hard to understand?

2h agoHN ↗

Mostly safe, contrary to other safer languages, Go memory model doesn't prevent data tearing.

1h agoHN ↗

It's been discussed for a long time, and the related proposals were heavily upvoted, including various older proposals.

As I understand it, part of the reason it took a while is that the core Go team was generally of the opinion that doing user-facing SIMD APIs the right way was to design a high-level, cross-platform API that would stand the test of time, and that was then punted a few times given its complexity and need to do other things.

Part of what helped the current approach take off was switching to a philosophy of designing a lower-level architecture-dependent API first (the 'simd/archsimd' package), and then later doing a higher-level portable API (the 'simd' package, which is topic of this blog post).

That two-level approach I think also gave some additional freedom for the design and implementation of the friendlier / high-level 'simd' package, including because the lower-level 'simd/archsimd' package is available for people who need or want to drop down.

It's a nice design.

1h agoHN ↗

Layered API design is a great way to resolve ergonomics/performance tradeoffs.

3h agoHN ↗

The interface conversion and type switch look like they should be inefficient, but the compiler-side implementation of simd specializes code and optimizes away the type switch.

I don’t understand this - how is it able to if the same go binary might run on unknown types? I’m assuming what it means is that the switch is implemented efficiently due to CPU branch prediction? I know fearless SIMD is doing cool stuff with static dispatch so that the feature set is checked just once at program start - is that what it means it’s doing under the hood? Very unclear.

3h agoHN ↗

It creates multiple versions of functions referencing SIMD and lifts the dispatch switching cost to their callers.

The AST rewrite creates multiple specialized copies of functions, variables, and types that mention simd types, where simd types are replaced with references to size-specialized types in simd/internal/bridge. Each of these bridge types is defined as an archsimd type, but with a restricted set of methods. The specialized functions, variables, and types acquire a suffix of the form @simdNNN, where NNN is either a vector length (128, 256, or 512) or 0, indicating emulation. Functions that mention simd internally, but not in their signature, are converted to wrappers that switch on the SIMD level detected at program start, and call the appropriate specialized version of that function. Specialized functions call other specialized functions directly without dispatch overhead (and perhaps with inlining). This rewrite strategy was chosen as a compromise between code duplication and SIMD performance; the overhead is hoisted as high as necessary to avoid dispatch within SIMD computations, but not higher. If SIMD dispatch appears “too low” in a computation, a gratuitous mention of a simd type will move it upwards, as in this example:

3h agoHN ↗

Go 1.26 and 1.27 include experimental APIs for Single Instruction Multiple Data (SIMD) operations.

You'd think these people would know the meaning of API, no?

2h agoHN ↗

One wonders what overly-narrow definition of API you're stuck on.

2h agoHN ↗

They could have used third party packages or Assembly directly.

This naturally is an easier way.

3h agoHN ↗

C++ is getting std::simd in the latest version and I am all aboard writing the vectorization with the least amount of intrinsic builtins I am able to. Even if not optimal, it's far better than the scalar ops.

1h agoHN ↗

Seconded!! This doesn’t really help the well established codebases much that are already doing this on a platform specific path but in general this is much appreciated for the future.

1h agoHN ↗

Write it once with N errors, not N*M errors :)

2h agoHN ↗

https://imjasonh.github.io/playground/palette-swap/ swaps colors in a provided image in wasm, entirely locally in your browser, to benchmark portable SIMD vs non-portable archsimd vs non-SIMD.

Portable SIMD is ~11% slower than non-portable SIMD in this case, but both are ~5x faster than non-SIMD.

1h agoHN ↗

I hope portable simd will be stabilized some time in rust :/

1h agoHN ↗

Rust kind of seems to have overtaken Go in momentum recently. I wonder if Go will do well in, say, two years from now on.

51m agoHN ↗

Just want to say among many portable SIMD solutions I’ve seen recently (e.g. Fearless SIMD), this is the first that makes non-fixed vectors like SVE and RISC-V vector (RVV) easier to support. Glad to see they made this decision

6m agoHN ↗

We pioneered this in Highway and shared some advice on the API. Great to see this decision taken :D

48m agoHN ↗

The new simd package hides these differences by removing fixed-size vectors from the type system, and by only supporting those operations that are in the intersection of all the different platforms, and fills gaps in the intersection with efficient emulation in terms of other SIMD instructions.

The intersection would be the operations supported by all platforms and so would not have gaps.

39m agoHN ↗

I did some testing with the experimental SIMD on a project I was doing to make speech-to-text and text-to-speech models run natively in Go (with CGO_ENABLED=0, so no C depenencies), and testing non-SIMD w/ SIMD.

I don't have formal benchmarks for that, but I can anecdotally say the SIMD work made a measurable improvement in the performance of the calculations vs. just plain Go. I'm very optimistic about how these improvements will help make the Go runtime an even better target for more of these types of work going forward, especially since it is cross-platform.