95 comments

[ 0.29 ms ] story [ 30.8 ms ] thread
This feature opens many doors for optimizing low-level performance in Go projects, that are already running multicore. IIRC there aren’t a lot of languages with built-in std lib support for SIMD and variants. Love the way Go is trying new stuff lately.
Besides the usual C and C++, we have Java, .NET, D, Zig, Julia, Swift, Rust.

So yeah, also appreciate have Go in the group instead of manually having to write Assembly.

However not many languages adopt ways to manually write SIMD, because most of us have no idea how to write good SIMD code in first place, I surely don't.

Even with languages that adopt ways to manually write SIMD, it’s mostly left to library maintainers rather than application developers.

I work for a C++ timeseries database startup that leverages SIMD about as much as we possibly can, and except for some extremely rare places we just use libraries.

Yeah, that is what I have heard from some NVidia folks as well, like Bryce Adelstein, use the libraries as much as possible, and leave the kernels for experts.

However even then, it depends on how the libraries API surface looks like.

With AI I'm pretty sure SIMD will be easier to integrate when necessary.
But it’s not necessary at all, the whole point is that these utility libraries bring you more elegant code that work on all platforms without having to pollute your codebase with SIMD intrinsics.

Unless this was tongue in cheek, because this is in fact a problem with AI that it degrades your codebase in these types of ways.

In 2026 if you are not doing A with AI you are doing it wrong /s
You probably weren't looking for a tutorial about SIMD, but just in case you were interested, Mitchell [0] did one recently that got on HN [1]

[0] Mitchell Hashimoto: "Everyone Should Know SIMD" https://mitchellh.com/writing/everyone-should-know-simd

[1] https://news.ycombinator.com/item?id=49010648

I definitely was so... thanks!
Thanks, pretty much appreciated, always worth at least to skim through.
Oh this is great, it was one of my biggest bugbears about Go since you almost always have to link C/C++ code to get the appropriate performance.

The one negative I'd say is that often autovectorisation is 'good enough' and this doesn't really tackle that gap.

The poor Assembler and the unsafe package forgotten in the corner.

While reaching out to CGO is the easier way, it doesn't mean it is the only tool available in Go.

FWIW, there is some pretty substantial autovectorization work that is already in-flight for Go.

It's being driven by an external contributor who has landed some good changes in the past to the Go compiler. (I think the autovectorization work might be part of their PhD or other academic research, but not sure.)

There's a CL stack here:

https://go.dev/cl/791740

It's hard to make predictions with an open source project, but my personal guess is some flavor of it will land (including it is already demonstrating good results without an enormous level of code complexity in the compiler and without overly slowing down compile speeds), but I guess we'll see.

The problem with Go isn't performance but with the C/C++ interop overhead, even with the "30% less overhead" from a few updates ago which isnt true for 99% of cases, it isnt enough
Why is that the case? I don’t know low level programming so why is Go limited in interop with C?
its not limited but it has overhead because of the memory model of go doesnt match the C one so there has to be some sort of rerodering being done, that's what i understood atleast, and theres also the go concurrency
The main problem is that Goroutine stacks are small, starting at 2KiB. When you're calling a C function, the Go runtime can't know how much stack space the C function will use, so it has to defensively expand the stack in a lot of cases
Use Assembly instead of CGO, isn't that scary, back in the 8 bit days we were coding Assembly aged 10, on our Spectrum, C64, Atari, Apple, Acorn, MSX,....
Already using this for foreground estimation of cutouts in my project, around 30% speedup over non-SIMD, but the algorithm is probably not very optimised yet.
This is why I love Go. Nobody was asking for this, but they took the time to do it right and continue to Push go as a memory safe, high-level systems language.
Go is in no way automatically memory safe. It's up to the programmer to write memory safe code with it.
Go is broadly considered to be a memory safe language.

See for example comments from tptacek like:

https://news.ycombinator.com/item?id=43335748

https://news.ycombinator.com/item?id=46028232

(The gist: memory safety is a term of art coined by security practitioners. Go, Python, Rust, Java, others: memory safe. C/C++: memory unsafe. Periodically, people in different slices of industry or academia come up with new definitions of memory safety that declare Rust or Go or other languages to be memory unsafe, but that is not by the broadly accepted definition across industry.)

Rust does allow you to overflow buffers, confuse types, and duplicate mutable pointers in safe code. See cve-rs.
No, Rust does not allow that. The current Rust compiler does, but that’s a bug that is being fixed.

At some point in the future, a fully backwards compatible Rust compiler will report an error when you try to compile cve-rs.

Since we are not at some point in the future where that correct compiler exists and there is only one official compiler, the distinction you make is practically meaningless!
I mean… no? It matters whether something is a part of the language or not, because it matters if you can write code relying on this behavior. Since this is a compiler bug, you cannot - the code will stop compiling the moment the bug is fixed.

There are no known instances of this bug being encountered in the wild, and if you look into it, you will see how extremely unlikely such code is.

Isn't one of the bugs around ten years old, now? Isn't ten years enough to call something a feature of the language rather than a bug?

I like Rust, but with this bug existing for so long, I personally no longer think of it as memory-safe.

> Isn't ten years enough to call something a feature of the language rather than a bug?

I suppose it depends on who is doing the classifying? From the developer's standpoint I'd imagine intent is all that matters: a bug is something that does not match developer intent and that is (eventually) expected to be changed to match the intent, while a feature is something that does match developer intent. From a user's standpoint I'd imagine it's a combination of developer intent and the user's reliance on said behavior, but IIRC in this particular case there's no known non-demo code that has organically run into this particular bug so there's little weight in favor of calling the bug a feature despite the devs' stance.

Also as GP said I think one needs to be careful to distinguish between the compiler and the language. IIRC the devs have known for basically this entire time exactly in what manner the Rust compiler fail to implement the rules of Rust the language, but a general fix has been blocked on long-running projects that have only recently been approaching the finish line [1].

[0]: https://news.ycombinator.com/item?id=40431444

[1]: https://blog.rust-lang.org/2026/08/21/enabling-next-solver-o...

I define undefined behaviour as a bug in C++. Now C++ is memory-safe!

btw, it's not actually that hard to write correct code in C++, easier than in C because you have all the container types. The problem is that nothing will tell you when you write incorrect code - there's no guarantee.

UB is part of the C++ standard. Surely you can understand the difference between the C++ standard and bugs in compilers implementing the C++ standard. This is that.
Which part of the Rust standard does cve-rs violate?
Rust does not have an ISO standard, but it does have a language design, and if you knew the first thing about cve-rs (including what’s on its own Github page), you would know that this is an extremely confirmed soundness bug.

The cve-rs repo is not meant to be the toxic gotcha aimed at Rust language maintainers you seem to think it is. It’s a repro case.

So the detailed spec is "whatever the compiler does". And the compiler allows cve-rs, so it does not violate the detailed spec.
No, and you are clearly trolling, and I’ll engage in no further interaction with you.
> I define undefined behaviour as a bug in C++. Now C++ is memory-safe!

I mean, sure, insofar as such a thing would also imply that a) the standard would need quite a bit of cleanup/clarification work to not contradict your definition, and b) the main optimizing C++ compilers are miscompiling code, analogous to how cve-rs is a rustc miscompilation rather than an issue with Rust itself.

(Fil-C might be an interesting exception here, though IIRC its definition of memory safety is slightly different)

And yet, there's a C compiler (fil-c) that doesn't allow that.
He’s very wrong about this. Just because ‘tptacek posts a lot and did security once upon a time does not make him “broad consideration”.
Well, you're right about one thing: the fact that I've spent my career in software security doesn't make me "broad consideration". The cites I give on what "memory safe" means, though, do.

My argument has never been "memory safe means what I say it does because I say so", but I get how that's a much more convenient argument to knock down than the ISRG site built specifically to talk about this.

ISRG is wrong too, definition-wise. Go is not memory safe. It is much safer, and I wouldn’t fault you for taking a C codebase and porting it to Go to avoid memory safety problems, but that does not make it memory safe in the same way Rust et al are memory safe. This is the same way that MTE does not thwart all memory corruption but it stops a lot of them. I accept your premise that Go has brought memory safety over the line to where it is apparently easier to find logic bugs than exploit memory corruption, which is laudable since C(++) has never been able to do this and likely never will, but in line with the pedantry that started this whole chain of comments, it’s not memory safe.
No idea why you're getting downvoted for true statement. Without a ? like in C# you're always at risk of a nil pointer being dereferenced
You can write unsafe code in Go (import unsafe), but then, you can do the same in Rust. Unsafe code is not the default, and in day to day Go i rarely see the use of the unsafe package.
What he probably means is data-races in go can result in memory/type unsafe accesses -- I suspect, likely due to slice types -- not sure if that is true/false.
Sure, but a data race is, IMHO not the same as memory safety. A data race, can be 100% memory safe, but just cause a logic bug in some program. I often see people mixing memory safety with racing. Go has bounds checks so you end up with a panic either way. Not UB.

As an (outside go) example, Ocaml (5) promises strong memory safety, but not to be data race free. A data race is not something we can prevent, because its usually not bound by code, but by time and the race-source rarely in source-code.

This means we have data races in http, database inserts etc. The source is usually not a concurrent task in source code-land.

That’s a race condition, not a data race.
Russ Cox points out that races are the one place in Go besides unsafe where Go lacks memory safety: https://research.swtch.com/gorace

I do think the nitpicking about this is mostly from people that want to say “my favorite language is safer than Go” which stupid and annoying.

Mostly safe, contrary to other safer languages, Go memory model doesn't prevent data tearing.
People were definitely asking for it.
It's been discussed for a long time, and the related proposals were heavily upvoted, including various older proposals.

As I understand it, part of the reason it took a while is that the core Go team was generally of the opinion that doing user-facing SIMD APIs the right way was to design a high-level, cross-platform API that would stand the test of time, and that was then punted a few times given its complexity and need to do other things.

Part of what helped the current approach take off was switching to a philosophy of designing a lower-level architecture-dependent API first (the 'simd/archsimd' package), and then later doing a higher-level portable API (the 'simd' package, which is topic of this blog post).

That two-level approach I think also gave some additional freedom for the design and implementation of the friendlier / high-level 'simd' package, including because the lower-level 'simd/archsimd' package is available for people who need or want to drop down.

It's a nice design.

This is kind of the opposite of Go. Not giving people what they are asking for.

There are pros and cons of course. You don't have 17 different ways to iterate over an array, so that's nice. But you also went 13 years without generics, despite them being one of the most requested features, because the designers didn't want that complexity inside Go.

Overall I think Go is better for this philosophy but there are times where the language is clearly written more for its maintainers than it's users.

Some of the concerns around generics and why it took so long were for the users as well. One of the biggest draws to Go has always been that you get the performance of a compiled language and yet compile times are so low that it can feel like you're developing with an interpreted language. The design of generics needed to maintain the compile time advantage or else it wouldn't feel like Go any more.
Languages like CLU, Ada, Delphi, Standard ML, OCaml, D, were having Go like fast compilation times, with generics, some of them decades before Go was created on much weaker hardware.
> But you also went 13 years without generics

Go shipped with generics (aka bounded parametric polymorphism), but only for built-in types: slices, arrays, and maps. That, with subtyping via interfaces, handled most demand for generics. The most common pain point was custom containers.

Go was first released in November 2009. Russ Cox posted "The Generic Dilemma" [1] in December 2009. The comments show the generics debate raging from the earliest days.

As a fun side note, I forgot I posted a comment on that post pointing to Ada's generics. I was in college, and Ada was our intro language.

[1]: https://research.swtch.com/generic

> the designers didn't want that complexity inside Go.

Yes, with some nuance. Go's goal of writing server programs didn't require the type-system complexity and run-time hit of user-defined generics. [2]

> Go was intended as a language for writing server programs [...] Polymorphic programming did not seem essential [...] so was initially left out for simplicity. > > Generics are convenient but they come at a cost in complexity in the type system and run-time. It took a while to develop a design that we believe gives value proportionate to the complexity.

[2]: https://go.dev/doc/faq#beginning_generics

Out of curiosity, I collected all proposals for Go's journey to generics. https://gist.github.com/jschaf/eaa7aff1af14ea7276a18a1b7370d...

> The interface conversion and type switch look like they should be inefficient, but the compiler-side implementation of simd specializes code and optimizes away the type switch.

I don’t understand this - how is it able to if the same go binary might run on unknown types? I’m assuming what it means is that the switch is implemented efficiently due to CPU branch prediction? I know fearless SIMD is doing cool stuff with static dispatch so that the feature set is checked just once at program start - is that what it means it’s doing under the hood? Very unclear.

[delayed]
One wonders what overly-narrow definition of API you're stuck on.
I suspect you stopped reading at web services on your link, but API is indeed the correct word to describe a set of functions from a library (built-in or not). If you disagree perhaps you should share your preferred term here?

> The term API is often used to refer to web APIs, which allow communication between computers that are joined by the internet. There are also APIs for programming languages, software libraries, computer operating systems, and computer hardware.

[delayed]
“API” meant “library functions” long before the first web server responded with JSON.
This will welcome more database/warehouses to be written in Go.

Personally I will implement it in github.com/viggy28/streambed

They could have used third party packages or Assembly directly.

This naturally is an easier way.

You're right. I could have but this encourages me to seriously consider it.
C++ is getting std::simd in the latest version and I am all aboard writing the vectorization with the least amount of intrinsic builtins I am able to. Even if not optimal, it's far better than the scalar ops.
Seconded!! This doesn’t really help the well established codebases much that are already doing this on a platform specific path but in general this is much appreciated for the future.
https://imjasonh.github.io/playground/palette-swap/ swaps colors in a provided image in wasm, entirely locally in your browser, to benchmark portable SIMD vs non-portable archsimd vs non-SIMD.

Portable SIMD is ~11% slower than non-portable SIMD in this case, but both are ~5x faster than non-SIMD.

I hope portable simd will be stabilized some time in rust :/
Rust kind of seems to have overtaken Go in momentum recently. I wonder if Go will do well in, say, two years from now on.
I personally use both, and keep using both. There are much more Golang job in the market now. Noone planning to ditch Go in my network or unhappy with it. Highload E-commerce, logistics, etc are way easier to write in Go IMHO.

Discover a very good niche for Rust - geo spatial analytics. Would not do it Go or Python. LLMs gave a huge boost to Rust too. Claude produce a very high code ... if designed right. Lot of feature complete libraries now.

Both will do fine

Just want to say among many portable SIMD solutions I’ve seen recently (e.g. Fearless SIMD), this is the first that makes non-fixed vectors like SVE and RISC-V vector (RVV) easier to support. Glad to see they made this decision
How so? I imagine you'd still want to constrain the length to the maximum vector size supported by the lowest platform you want to support or you lose the portability and actually end up with code that performs much worse than the scalar alternative on some platforms.

Mojo has an even more portable simd[1] type that isn't just generic over length but also over type. In my opinion it is almost always better to specialize for each platform and use portable implementation as fallback. It's a shame that just very few languages support Zig-like comptime, because it would be excellent for specializations without introducing runtime penalties.

1. https://mojolang.org/docs/std/simd/SIMD/

> constrain the length to the maximum vector size supported by the lowest platform you want to support or you lose the portability and actually end up with code that performs much worse than the scalar alternative on some platforms.

Or, put a dynamic factor into your vector size and design everything around it. Such that every platforms can plug in their own factor and _scale_ the size of vectors. This is basically what LLVM IR does for SVE and RVV: `<vscale x 4 x i32>` where vscale is the said dynamic factor. Though the exact value of vscale is only known during runtime, it doesn't matter -- we still can design compiler optimizations and lowering around it. The generated binaries can then be portable across platforms with different vscale values.

> The new simd package hides these differences by removing fixed-size vectors from the type system, and by only supporting those operations that are in the intersection of all the different platforms, and fills gaps in the intersection with efficient emulation in terms of other SIMD instructions.

The intersection would be the operations supported by all platforms and so would not have gaps.

I did some testing with the experimental SIMD on a project I was doing to make speech-to-text and text-to-speech models run natively in Go (with CGO_ENABLED=0, so no C depenencies), and testing non-SIMD w/ SIMD.

I don't have formal benchmarks for that, but I can anecdotally say the SIMD work made a measurable improvement in the performance of the calculations vs. just plain Go. I'm very optimistic about how these improvements will help make the Go runtime an even better target for more of these types of work going forward, especially since it is cross-platform.

Very neat, and comes pretty close to how Mojo handles portable SIMD.

It's great to see two of my favorite languages finally making SIMD easy to use. It's such low-hanging fruit for performance, yet somehow languages have ignored it for years. Portable SIMD, even with some performance penalty, still beats scalar computation whenever vector operations are needed. Yet language implementations always seemed to assume that hardware-specific SIMD APIs were the only way to go. That did nothing but make SIMD unusable excepting special cases where performance is absolutely critical, rather than just something anyone can use in day to day programming.

For God sake, add a syntax highlighting on the official page! Otherwise this is awesome
Rob Pike seems to dislike syntax highlighting:

> Syntax highlighting is juvenile. When I was a child, I was taught arithmetic using colored rods (http://en.wikipedia.org/wiki/Cuisenaire_rods). I grew up and today I use monochromatic numerals.

https://groups.google.com/g/golang-nuts/c/hJHCAaiL0so/m/kG3B...

He is definitely not alone in thinking this, for example (but for different reasons): https://www.linusakesson.net/programming/syntaxhighlighting/

Not convinced at all by these reasons. Rob Pike has long stepped down from the Go team so it shouldn't matter? This is a website not a language change.
I like Linus's work, but that post smells of rotting strawmen. All the examples look manipulated into throwing the opposition under the bus.
Wow, that second link does exactly what I do when I'm talking to people about skittle highlighting; demonstrating a skittle highlighted snippet of prose to show how absolutely harmful it is. I only started doing that maybe three years ago, so they've got me beat by 15 years or so! Glad I'm not alone on this crusade, at least. It is genuinely my belief that skittles have cost humankind millions of manhours of productivity, at minimum.
Rob Pike doesn't disclose that he has color blindness when he says he doesn't see the value in syntax highlighting. Classy.
Curious how much of the emulation ends up in hot paths before SVE and the feature variants land.
Now with Mojo and WebAssembly I notice 3 different approaches for platform-independent SIMD. For examples operating on an array of floats:

- WebAssembly: 4 float32s

- Mojo: N float32s where N is a compile-time parameter

- Go simd: vector of float32s