<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<title>Sergey &quot;Shnatsel&quot; Davidoff</title>
	<subtitle>Rust, security, and making computers go brrr</subtitle>
	<link href="https://shnatsel.github.io/atom.xml" rel="self" type="application/atom+xml"/>
	<link href="https://shnatsel.github.io"/>
	<generator uri="https://www.getzola.org/">Zola</generator>
	<updated>2026-09-25T00:00:00+00:00</updated>
	<id>https://shnatsel.github.io/atom.xml</id>
	<entry xml:lang="en">
		<title>The state of SIMD in Rust in 2026</title>
        <author>
            <name>Sergey &quot;Shnatsel&quot; Davidoff</name>
        </author>
		<published>2026-09-25T00:00:00+00:00</published>
		<updated>2026-09-25T00:00:00+00:00</updated>
		<link href="https://shnatsel.github.io/state-of-simd-rust-2026/"/>
		<link rel="alternate" href="https://shnatsel.github.io/state-of-simd-rust-2026/" type="text/html"/>
		<id>https://shnatsel.github.io/state-of-simd-rust-2026/</id>
        <summary type="html">&lt;p&gt;A lot of progress was made since last year, and I made some of it!&lt;&#x2F;p&gt;</summary>
		<content type="html">&lt;p&gt;A lot of progress was made since last year, and I made some of it!&lt;&#x2F;p&gt;
&lt;span id=&quot;continue-reading&quot;&gt;&lt;&#x2F;span&gt;
&lt;p&gt;After &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;shnatsel.medium.com&#x2F;the-state-of-simd-in-rust-in-2025-32c263e5f53d&quot;&gt;last year&#x27;s survey&lt;&#x2F;a&gt; I started contributing to the SIMD library that seemed the most promising. One thing led to another, and now I&#x27;m a maintainer of Fearless SIMD.&lt;&#x2F;p&gt;
&lt;p&gt;To avoid a conflict of interest, I invited authors of other libraries (&lt;code&gt;std::simd&lt;&#x2F;code&gt;, &lt;code&gt;wide&lt;&#x2F;code&gt;, &lt;code&gt;pulp&lt;&#x2F;code&gt;, &lt;code&gt;macerator&lt;&#x2F;code&gt;) to review and provide feedback on a draft of this article. However, I retained editorial control, and all mistakes are my own.&lt;&#x2F;p&gt;
&lt;p&gt;This year&#x27;s survey is more in-depth than my previous one. So buckle up, and let&#x27;s take it... &lt;em&gt;from the top!&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;h2 id=&quot;what-s-simd-why-simd&quot;&gt;What’s SIMD? Why SIMD?&lt;a class=&quot;zola-anchor&quot; href=&quot;#what-s-simd-why-simd&quot; aria-label=&quot;Anchor link for: what-s-simd-why-simd&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;Hardware that does arithmetic is cheap, so any CPU made this century has plenty of it. But you still only have one instruction decoding block and it is hard to get it to go fast, so the arithmetic hardware is vastly underutilized.&lt;&#x2F;p&gt;
&lt;p&gt;To get around the instruction decoding bottleneck, you can feed the CPU a batch of numbers all at once for a single arithmetic operation like addition. Hence the name: “single instruction, multiple data,” or SIMD.&lt;&#x2F;p&gt;
&lt;p&gt;Instead of adding two numbers together, you can add two batches or “vectors” of numbers and it takes about the same amount of time as doing just one addition.&lt;&#x2F;p&gt;
&lt;p&gt;On recent x86 chips these batches can be up to 512 bits in size, so in theory you can get an 8x speedup for math on &lt;code&gt;f64&lt;&#x2F;code&gt; or a 64x speedup on &lt;code&gt;u8&lt;&#x2F;code&gt;. In practice it can run both &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Advanced_Vector_Extensions#Downclocking&quot;&gt;slower&lt;&#x2F;a&gt; and &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Instruction-level_parallelism&quot;&gt;faster&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;instruction-sets&quot;&gt;Instruction sets&lt;a class=&quot;zola-anchor&quot; href=&quot;#instruction-sets&quot; aria-label=&quot;Anchor link for: instruction-sets&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;Historically, SIMD instructions were added after the CPU architecture was already designed, so SIMD is an extension with its own marketing name on each architecture.&lt;&#x2F;p&gt;
&lt;p&gt;ARM calls theirs “NEON”, and all 64-bit ARM CPUs have it.&lt;&#x2F;p&gt;
&lt;p&gt;WebAssembly doesn’t have a marketing department, so they just call theirs “WebAssembly 128-bit packed SIMD extension”.&lt;&#x2F;p&gt;
&lt;p&gt;64-bit x86 shipped with one called “SSE2” which has basic instructions for 128-bit vectors, but &lt;em&gt;later&lt;&#x2F;em&gt; they added a whole menagerie of extensions on top of that, with SSE 4.2 adding more operations, AVX and AVX2 adding 256-bit vectors and AVX-512 adding 512-bit vectors and even more operations.&lt;&#x2F;p&gt;
&lt;p&gt;The word “later” in the above paragraph creates a problem.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;does-this-cpu-have-that-instruction&quot;&gt;Does this CPU have that instruction?&lt;a class=&quot;zola-anchor&quot; href=&quot;#does-this-cpu-have-that-instruction&quot; aria-label=&quot;Anchor link for: does-this-cpu-have-that-instruction&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h3&gt;
&lt;p&gt;If you’re running a program on an x86_64 CPU, it’s not a given that the CPU has any particular SIMD extension. So by default the compiler isn’t allowed to use instructions beyond SSE2 because that won’t work on all x86_64 CPUs.&lt;&#x2F;p&gt;
&lt;p&gt;There are two ways around this problem.&lt;&#x2F;p&gt;
&lt;p&gt;If you work for a company that only ever runs their binaries on their own servers or on a public cloud, you can just assert that they’re all recent enough to at least have AVX2 that was introduced over 10 years ago, and have the program crash or misbehave if it ever runs on anything without AVX2:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;plain&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;RUSTFLAGS=&amp;#39;-C target-cpu=x86-64-v3&amp;#39; cargo build --release&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;However, if you are distributing the binaries for other people to run, that’s not really an option.&lt;&#x2F;p&gt;
&lt;p&gt;Instead you can do something called &lt;strong&gt;function multiversioning:&lt;&#x2F;strong&gt; compile the same function multiple times for different SIMD extensions, and when the program actually runs, check what features the CPU supports and select the appropriate version based on that.&lt;&#x2F;p&gt;
&lt;p&gt;Fortunately, this problem only exists on x86.&lt;&#x2F;p&gt;
&lt;p&gt;ARM made NEON mandatory on its 64-bit CPUs and hasn&#x27;t really added useful SIMD extensions after that (more on that later).&lt;&#x2F;p&gt;
&lt;p&gt;WebAssembly makes you compile two different binaries, one with SIMD and one without, and use JavaScript to check if the browser supports SIMD.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;how-do-i-simd&quot;&gt;How do I SIMD?&lt;a class=&quot;zola-anchor&quot; href=&quot;#how-do-i-simd&quot; aria-label=&quot;Anchor link for: how-do-i-simd&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;There are three ways to leverage SIMD:&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;Automatic vectorization: &lt;code&gt;&amp;amp;[i32].sum()&lt;&#x2F;code&gt;&lt;&#x2F;li&gt;
&lt;li&gt;Portable SIMD abstractions: &lt;code&gt;i32x4 + i32x4&lt;&#x2F;code&gt;&lt;&#x2F;li&gt;
&lt;li&gt;Platform-specific intrinsics - hang on, we&#x27;re gonna need a bigger code block:&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;rust&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;#&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;[&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;cfg&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;all&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;any&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;target_arch &lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;=&lt;&#x2F;span&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt; &amp;quot;x86&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; target_arch &lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;=&lt;&#x2F;span&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt; &amp;quot;x86_64&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;),&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;target_feature &lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;=&lt;&#x2F;span&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt; &amp;quot;sse2&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;))]&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;_mm_add_epi32&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;__m128i&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; ,&lt;&#x2F;span&gt;&lt;span&gt; __m128i&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;#&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;[&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;cfg&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;all&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;target_arch &lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;=&lt;&#x2F;span&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt; &amp;quot;aarch64&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; target_feature &lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;=&lt;&#x2F;span&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt; &amp;quot;neon&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;))]&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;vaddq_u32&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;int32x4_t&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; int32x4_t&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Let&#x27;s look at what each one entails and what the state of each programming model is.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;automatic-vectorization&quot;&gt;Automatic vectorization&lt;a class=&quot;zola-anchor&quot; href=&quot;#automatic-vectorization&quot; aria-label=&quot;Anchor link for: automatic-vectorization&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;Just write plain Rust and let the compiler heuristics do the work!&lt;&#x2F;p&gt;
&lt;p&gt;You can get it to work quite well, if you are careful to write code in a way that the compiler can reliably(ish) vectorize. This usually involves iterating over &lt;code&gt;&amp;amp;[i32].as_chunks()&lt;&#x2F;code&gt; instead of &lt;code&gt;&amp;amp;[i32]&lt;&#x2F;code&gt; and benchmarking or staring at the assembly to verify it worked. See &lt;strong&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;matklad.github.io&#x2F;2023&#x2F;04&#x2F;09&#x2F;can-you-trust-a-compiler-to-optimize-your-code.html&quot;&gt;Can You Trust a Compiler to Optimize Your Code?&lt;&#x2F;a&gt;&lt;&#x2F;strong&gt; for details.&lt;&#x2F;p&gt;
&lt;p&gt;This is the easiest option to use, requires no dependencies, and automatically supports all instruction sets the compiler supports, no matter how obscure.&lt;&#x2F;p&gt;
&lt;p&gt;The downside is that this method is not very reliable. The larger and more complex your function is, the greater is the chance that the compiler will not be able to vectorize it. Performance can also swing wildly depending on the compiler version or due to changes to the surrounding code.&lt;&#x2F;p&gt;
&lt;p&gt;Floating-point types also need special care.&lt;&#x2F;p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Floats are weird.&lt;&#x2F;strong&gt; Even something as trivial as summing an array of floats with reasonable precision gets surprisingly involved, see &lt;strong&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;orlp.net&#x2F;blog&#x2F;taming-float-sums&#x2F;&quot;&gt;Taming Floating-Point Sums&lt;&#x2F;a&gt;&lt;&#x2F;strong&gt;.&lt;&#x2F;p&gt;
&lt;&#x2F;blockquote&gt;
&lt;p&gt;Previously automatic vectorization didn&#x27;t work with floating-point types because it would change the precision of the result (often for the better, but the compiler is not permitted to change any observable results).&lt;&#x2F;p&gt;
&lt;p&gt;This changed in &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;doc.rust-lang.org&#x2F;stable&#x2F;releases.html#version-1980-2026-08-20&quot;&gt;Rust 1.98&lt;&#x2F;a&gt; which stabilized algebraic ops such as &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;doc.rust-lang.org&#x2F;stable&#x2F;core&#x2F;primitive.f32.html#method.algebraic_add&quot;&gt;&lt;code&gt;algebraic_add()&lt;&#x2F;code&gt;&lt;&#x2F;a&gt; that let the compiler change the observable result, like a &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;codingnest.com&#x2F;files&#x2F;Fun,%20Safe,%20Math%20Optimizations.pdf&quot;&gt;less dangerous &lt;code&gt;-ffast-math&lt;&#x2F;code&gt;&lt;&#x2F;a&gt;. You still have to rewrite your code to use them for it to be eligible for vectorization in most cases.&lt;&#x2F;p&gt;
&lt;p&gt;And you still need to get multiversioning somehow. So while we&#x27;re at it...&lt;&#x2F;p&gt;
&lt;h3 id=&quot;the-multiversion-crate&quot;&gt;The &#x27;multiversion&#x27; crate&lt;a class=&quot;zola-anchor&quot; href=&quot;#the-multiversion-crate&quot; aria-label=&quot;Anchor link for: the-multiversion-crate&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h3&gt;
&lt;p&gt;The all-in-one SIMD crates discussed below also provide multiversioning, but let&#x27;s take a look at &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;multiversion&quot;&gt;&lt;code&gt;multiversion&lt;&#x2F;code&gt;&lt;&#x2F;a&gt; real quick since it&#x27;s most useful for automatic vectorization.&lt;&#x2F;p&gt;
&lt;p&gt;It&#x27;s very easy to use: you add the &lt;code&gt;#[multiversion(targets = &quot;simd&quot;)]&lt;&#x2F;code&gt; annotation to your function and that&#x27;s it.&lt;&#x2F;p&gt;
&lt;p&gt;But that ease hides an undocumented pitfall: calling a function annotated with &lt;code&gt;#[multiversion]&lt;&#x2F;code&gt; has a little bit of overhead. It is very small - under a dozen instructions, but it shows up as significant overhead if the function you put it on is itself tiny.&lt;&#x2F;p&gt;
&lt;p&gt;As a rule of thumb, if your function has a loop in it, add &lt;code&gt;#[multiversion]&lt;&#x2F;code&gt;; if it processes a handful of values add &lt;code&gt;#[inline(always)]&lt;&#x2F;code&gt;, so long as there is &lt;code&gt;#[multiversion]&lt;&#x2F;code&gt; somewhere up the call chain.&lt;&#x2F;p&gt;
&lt;p&gt;The other crates listed below don&#x27;t have this pitfall and don&#x27;t make you think about the sizes of functions, at the cost of more boilerplate.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;code&gt;multiversion&lt;&#x2F;code&gt; is the only crate that allows you to list the exact CPU extensions you require, as opposed to opting in to a predefined SIMD level. So if your code happens to benefit from some very recent instruction, you can opt in to it. But in my experience this hardly ever comes up for autovectorized code.&lt;&#x2F;p&gt;
&lt;p&gt;For AVX-512 &lt;code&gt;multiversion&lt;&#x2F;code&gt; checks if it&#x27;s present, not whether it&#x27;s actually fast, which may hurt performance in practice (more on that below). You can work around that at the cost of boilerplate - you have to put this on every function:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;rust&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;#&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;[&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;multiversion&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;multiversion&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;targets&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt;    &amp;quot;x86_64+cmpxchg16b+popcnt+sse3+sse4.1+sse4.2+ssse3&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt; &#x2F;&#x2F; x86_64-v2&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt;    &amp;quot;x86_64+avx+avx2+bmi1+bmi2+cmpxchg16b+f16c+fma+lzcnt+movbe+popcnt+sse3+sse4.1+sse4.2+ssse3+xsave&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt; &#x2F;&#x2F; x86_64-v3&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt;    &amp;quot;x86_64+fxsr,adx,avx512bitalg,avx512bw,avx512cd,avx512dq,avx512f,avx512ifma,avx512vbmi,avx512vbmi2,avx512vl,avx512vnni,avx512vpopcntdq,bmi1,bmi2,cmpxchg16b,fma,gfni,lzcnt,movbe,pclmulqdq,popcnt,vpclmulqdq,xsave,xsavec,xsaveopt,xsaves&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt; &#x2F;&#x2F; Ice Lake and later&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)]]&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;&lt;h2 id=&quot;portable-simd-abstractions&quot;&gt;Portable SIMD abstractions&lt;a class=&quot;zola-anchor&quot; href=&quot;#portable-simd-abstractions&quot; aria-label=&quot;Anchor link for: portable-simd-abstractions&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;There are several production-ready ones. The desirable features are:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Fixed-width vectors:&lt;&#x2F;strong&gt; write code in terms of &lt;code&gt;f32x4&lt;&#x2F;code&gt;, &lt;code&gt;u8x16&lt;&#x2F;code&gt;, etc (known size)&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Hardware-width vectors:&lt;&#x2F;strong&gt; use the largest vector size the hardware supports, without knowing it in advance&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Generic over element type:&lt;&#x2F;strong&gt; write code that works on both &lt;code&gt;f32x4&lt;&#x2F;code&gt; and &lt;code&gt;f64x2&lt;&#x2F;code&gt;&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Generic over vector width:&lt;&#x2F;strong&gt; write code that works on all of &lt;code&gt;f32x4&lt;&#x2F;code&gt;, &lt;code&gt;f32x8&lt;&#x2F;code&gt;, &lt;code&gt;f32x16&lt;&#x2F;code&gt;&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;The TL;DR table:&lt;&#x2F;p&gt;
&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;&lt;&#x2F;th&gt;&lt;th style=&quot;text-align: center&quot;&gt;std::simd (nightly)&lt;&#x2F;th&gt;&lt;th style=&quot;text-align: center&quot;&gt;fearless simd&lt;&#x2F;th&gt;&lt;th style=&quot;text-align: center&quot;&gt;wide&lt;&#x2F;th&gt;&lt;th style=&quot;text-align: center&quot;&gt;pulp&lt;&#x2F;th&gt;&lt;th style=&quot;text-align: center&quot;&gt;macerator&lt;&#x2F;th&gt;&lt;&#x2F;tr&gt;&lt;&#x2F;thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;multiversioning&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;📦&#x2F;🛠️&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;❌&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;fixed-width vectors&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;☑️&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;❌&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;hardware-width vectors&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;📦&#x2F;🛠️&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;🛠️&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;generic over element type&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;🛠️&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;🛠️&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;generic over vector width&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;🛠️&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;safe access to intrinsics&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;🛠️&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;☑️&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;🛠️&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;trigonometry&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;📦&#x2F;🛠️&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;🛠️&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;☑️&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;🛠️&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;🛠️&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;&#x2F;tbody&gt;&lt;&#x2F;table&gt;
&lt;ul&gt;
&lt;li&gt;✅ Yes&lt;&#x2F;li&gt;
&lt;li&gt;☑️ Yes, with caveats&lt;&#x2F;li&gt;
&lt;li&gt;📦 Yes, with a third-party crate&lt;&#x2F;li&gt;
&lt;li&gt;🛠️ Build it yourself&lt;&#x2F;li&gt;
&lt;li&gt;❌ Absolutely not&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;And the instruction set support:&lt;&#x2F;p&gt;
&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th style=&quot;text-align: center&quot;&gt;&lt;&#x2F;th&gt;&lt;th style=&quot;text-align: center&quot;&gt;std::simd&lt;&#x2F;th&gt;&lt;th style=&quot;text-align: center&quot;&gt;fearless simd&lt;&#x2F;th&gt;&lt;th style=&quot;text-align: center&quot;&gt;wide&lt;&#x2F;th&gt;&lt;th style=&quot;text-align: center&quot;&gt;pulp&lt;&#x2F;th&gt;&lt;th style=&quot;text-align: center&quot;&gt;macerator&lt;&#x2F;th&gt;&lt;&#x2F;tr&gt;&lt;&#x2F;thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td style=&quot;text-align: center&quot;&gt;SSE2&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;☑️&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;🐌&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td style=&quot;text-align: center&quot;&gt;SSE4.x&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;☑️&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td style=&quot;text-align: center&quot;&gt;AVX2&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td style=&quot;text-align: center&quot;&gt;AVX-512&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td style=&quot;text-align: center&quot;&gt;NEON&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td style=&quot;text-align: center&quot;&gt;WASM&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td style=&quot;text-align: center&quot;&gt;All the rest&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;✅&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;🐌&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;🐌&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;🐌&lt;&#x2F;td&gt;&lt;td style=&quot;text-align: center&quot;&gt;🐌*&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;&#x2F;tbody&gt;&lt;&#x2F;table&gt;
&lt;ul&gt;
&lt;li&gt;✅ Has optimized routines&lt;&#x2F;li&gt;
&lt;li&gt;☑️ Implemented but not used. Requires writing a custom dispatch to opt in.&lt;&#x2F;li&gt;
&lt;li&gt;🐌 Reliant on autovectorization, often slow&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;* macerator also supports LoongArch because the author was, and I quote, &quot;bored&quot;.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;std-simd&quot;&gt;std::simd&lt;a class=&quot;zola-anchor&quot; href=&quot;#std-simd&quot; aria-label=&quot;Anchor link for: std-simd&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h3&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;doc.rust-lang.org&#x2F;std&#x2F;simd&#x2F;index.html&quot;&gt;std::simd&lt;&#x2F;a&gt; is not a complete solution for SIMD. It&#x27;s more of a set of building blocks that absolutely has to be in the standard library, while everything else is left up to the ecosystem crates.&lt;&#x2F;p&gt;
&lt;p&gt;The largest drawback is that it&#x27;s nightly-only, and still undergoes infrequent breaking API changes. So one day you update the compiler and your code stops compiling, and you have to go and fix it. But so long as you&#x27;re OK with that, and only need fixed-width vectors and maybe multiversioning, it&#x27;s pretty great!&lt;&#x2F;p&gt;
&lt;p&gt;&lt;code&gt;std::simd&lt;&#x2F;code&gt;&#x27;s &lt;em&gt;raison d&#x27;être&lt;&#x2F;em&gt; is that it sits directly on top of LLVM and can target any platform LLVM can target, including weird CPUs that only large banks use or that only the Chinese government uses. On the flip side, if LLVM doesn&#x27;t have a perfectly matching operation inside it for &lt;code&gt;std::simd&lt;&#x2F;code&gt; to make use of, there is no plan B and &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;shnatsel.github.io&#x2F;improving-std-simd-swizzle-dyn&#x2F;&quot;&gt;no SIMD is actually used&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;This happens disturbingly often. Its &lt;code&gt;sin()&lt;&#x2F;code&gt; could not be more apt: shipping scalar implementations in a SIMD guise is the cardinal sin. And &lt;code&gt;reduce_sum()&lt;&#x2F;code&gt; is somehow the worst case for &lt;em&gt;both&lt;&#x2F;em&gt; performance and accuracy. So don&#x27;t bother using any non-trivial functions on floats.&lt;&#x2F;p&gt;
&lt;p&gt;The closest thing we have to proper trigonometry is the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;sleef&quot;&gt;sleef&lt;&#x2F;a&gt; crate, a partial port of &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;shibatch&#x2F;sleef&quot;&gt;SLEEF&lt;&#x2F;a&gt; to &lt;code&gt;std::simd&lt;&#x2F;code&gt; that&#x27;s only &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;burrbull&#x2F;sleef-rs&#x2F;issues&#x2F;43&quot;&gt;a little&lt;&#x2F;a&gt; &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;burrbull&#x2F;sleef-rs&#x2F;issues&#x2F;44&quot;&gt;buggy&lt;&#x2F;a&gt;. And that&#x27;s the best trigonometry I have in this whole article!&lt;&#x2F;p&gt;
&lt;p&gt;&lt;code&gt;std::simd&lt;&#x2F;code&gt; is uniquely flexible when it comes to multiversioning. You can use the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;multiversion&quot;&gt;multiversion&lt;&#x2F;a&gt; crate or the multiversioning from any other SIMD crate in this section. All the other crates work with their own built-in multiversioning only.&lt;&#x2F;p&gt;
&lt;p&gt;Its &lt;code&gt;Simd&amp;lt;T, N&amp;gt;&lt;&#x2F;code&gt; API looks like it would be very elegant and work great if you could just do math on &lt;code&gt;N&lt;&#x2F;code&gt;, but &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;rust-lang.github.io&#x2F;project-const-generics&#x2F;documents&#x2F;min_const_generics_plan.html&quot;&gt;you cannot&lt;&#x2F;a&gt;. That feature is very incomplete even on nightly. Without it using &lt;code&gt;Simd&amp;lt;T, N&amp;gt;&lt;&#x2F;code&gt; to get hardware-sized vectors is doable, but &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;gist.github.com&#x2F;Shnatsel&#x2F;edc642125ac73fa7c365216c3a938802&quot;&gt;a lot uglier&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;While you can use &lt;code&gt;std::simd&lt;&#x2F;code&gt; directly in many cases and have it perform okay, disparately tacking on features through third-party crates only gets you so far. As an example, the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;sleef&quot;&gt;&lt;code&gt;sleef&lt;&#x2F;code&gt; crate&lt;&#x2F;a&gt; doesn&#x27;t work with the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;multiversion&quot;&gt;&lt;code&gt;multiversion&lt;&#x2F;code&gt; crate&lt;&#x2F;a&gt;, you have to fork &lt;code&gt;sleef&lt;&#x2F;code&gt; and mate them yourself. Third-party extensions work in isolation but don&#x27;t compose.&lt;&#x2F;p&gt;
&lt;p&gt;What you need is an all-in-one solution where all the parts work together. Speaking of which...&lt;&#x2F;p&gt;
&lt;h3 id=&quot;fearless-simd&quot;&gt;fearless_simd&lt;a class=&quot;zola-anchor&quot; href=&quot;#fearless-simd&quot; aria-label=&quot;Anchor link for: fearless-simd&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h3&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;fearless_simd&quot;&gt;Fearless SIMD&lt;&#x2F;a&gt; is an all-in-one solution where all the parts work together.&lt;&#x2F;p&gt;
&lt;p&gt;Just look at that beautiful column of green check boxes that makes in the tables!&lt;&#x2F;p&gt;
&lt;p&gt;Beyond the tables, the features unique to &lt;code&gt;fearless_simd&lt;&#x2F;code&gt; are:&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;Orders of magnitude less &lt;code&gt;unsafe&lt;&#x2F;code&gt; code under the hood than other crates &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;shnatsel.github.io&#x2F;safe-simd-in-rust-even-on-the-inside&#x2F;&quot;&gt;thanks to a clever design&lt;&#x2F;a&gt;.&lt;&#x2F;li&gt;
&lt;li&gt;Multiversioning that Just Works, even for tiny functions. Just slap &lt;code&gt;#[simd]&lt;&#x2F;code&gt; on a function and you&#x27;re done. A manual mode is available if you hate procedural macros.&lt;&#x2F;li&gt;
&lt;li&gt;Multiversioning is controlled by whoever builds the final binary. You can &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;linebender&#x2F;fearless_simd&#x2F;tree&#x2F;main&#x2F;fearless_simd#multiversioning-on-x86&quot;&gt;configure it&lt;&#x2F;a&gt; without patching the libraries.&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;p&gt;The main drawback is boilerplate: instead of&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;rust&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;fn&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; my_func&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt;a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; A&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; B&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;) {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;you have to write&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;rust&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;#&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;[&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;simd&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;]&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;fn&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; my_func&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;lt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;S&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Simd&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt;simd&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; S&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; A&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; B&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;) {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;which is a mouthful.&lt;&#x2F;p&gt;
&lt;p&gt;AVX-512 is only used on recent-ish CPUs where it doesn&#x27;t hurt performance (see the hardware section below). You can manually configure &lt;code&gt;multiversion&lt;&#x2F;code&gt; to behave like this, but it&#x27;s not the default and requires a lot of boilerplate (see above). All the other SIMD abstraction crates just check if AVX-512 is present or not.&lt;&#x2F;p&gt;
&lt;p&gt;It recently shipped v1.0, with a &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;linebender&#x2F;fearless_simd&#x2F;blob&#x2F;main&#x2F;fearless_simd&#x2F;SECURITY.md&quot;&gt;security policy&lt;&#x2F;a&gt; and everything.&lt;&#x2F;p&gt;
&lt;p&gt;The biggest gap is trigonometry. There just isn&#x27;t a port of anything like SLEEF to &lt;code&gt;fearless_simd&lt;&#x2F;code&gt; machinery yet.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;wide&quot;&gt;wide&lt;a class=&quot;zola-anchor&quot; href=&quot;#wide&quot; aria-label=&quot;Anchor link for: wide&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h3&gt;
&lt;p&gt;&lt;code&gt;wide&lt;&#x2F;code&gt; has a lot going for it: good platform coverage, lots of implemented operations, and it&#x27;s v1.0 already. It even has trigonometric functions, although their precision is &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;docs.rs&#x2F;wide&#x2F;1.7.0&#x2F;wide&#x2F;struct.f32x16.html#method.sin&quot;&gt;explicitly left unspecified&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;The biggest downside is that it&#x27;s fundamentally incompatible with multiversioning. This is fine if you&#x27;re not targeting x86, or if you always build with &lt;code&gt;-C target-cpu=&lt;&#x2F;code&gt; for known hardware, but cripples performance otherwise. The only workaround is &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;ronnychevalier&#x2F;cargo-multivers&#x2F;&quot;&gt;&lt;code&gt;cargo multivers&lt;&#x2F;code&gt;&lt;&#x2F;a&gt;, but it only works for long-running programs, otherwise its startup costs dwarf the performane gains from SIMD.&lt;&#x2F;p&gt;
&lt;p&gt;The other downside is not supporting any kind of generics, either over element types or vector widths. However, you can work around that using macros. Instead of making a function generic, wrap it in &lt;code&gt;macro_rules!&lt;&#x2F;code&gt; and write &lt;code&gt;$type::from_slice&lt;&#x2F;code&gt; instead of &lt;code&gt;T::from_slice&lt;&#x2F;code&gt;. It adds a bit of boilerplate, but removes the boilerplate for generic bounds, so win some lose some. I&#x27;ve done it, it&#x27;s not too bad, especially if you pull in something like the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;paste&quot;&gt;paste&lt;&#x2F;a&gt; crate.&lt;&#x2F;p&gt;
&lt;p&gt;If you&#x27;re only targeting a handful of types, e.g. &lt;code&gt;f32&lt;&#x2F;code&gt; and &lt;code&gt;f64&lt;&#x2F;code&gt;, you might want to use the macro approach regardless, even in libraries with generics, because it also allows you to have arrays of &quot;generic&quot; sizes. Actual generic array sizes are a nightly-only and incomplete feature. But you can also work around that with the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;generic-array&quot;&gt;generic-array&lt;&#x2F;a&gt; crate or just by making an array of the largest possible SIMD size.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;pulp&quot;&gt;pulp&lt;a class=&quot;zola-anchor&quot; href=&quot;#pulp&quot; aria-label=&quot;Anchor link for: pulp&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h3&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;pulp&quot;&gt;pulp&lt;&#x2F;a&gt; was built to power the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;faer&quot;&gt;faer&lt;&#x2F;a&gt; linear algebra library. This informs its priorities: the implemented operations are mostly math (e.g. no swizzles), and the API is geared towards native-width vectors.&lt;&#x2F;p&gt;
&lt;p&gt;Fixed-width vectors are technically possible, but completely undocumented and quite awkward to use. I&#x27;ve contributed some &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;sarah-quinones&#x2F;pulp&#x2F;pull&#x2F;36&quot;&gt;fixes&lt;&#x2F;a&gt; for them while I was researching them, including for &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;sarah-quinones&#x2F;pulp&#x2F;pull&#x2F;37&quot;&gt;a soundness bug&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;There is no native support for being generic over the element type, but the macro trick I described for &lt;code&gt;wide&lt;&#x2F;code&gt; should work fine here too.&lt;&#x2F;p&gt;
&lt;p&gt;Its multiversioning is &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;docs.rs&#x2F;pulp&#x2F;latest&#x2F;pulp&#x2F;#manual-vectorization-example&quot;&gt;the most verbose I&#x27;ve ever seen&lt;&#x2F;a&gt;. There&#x27;s &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;sarah-quinones&#x2F;pulp&#x2F;#less-boilerplate-using-pulpwith_simd&quot;&gt;a macro to reduce boilerplate&lt;&#x2F;a&gt; but even that is rather verbose compared to the alternatives.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;macerator&quot;&gt;macerator&lt;a class=&quot;zola-anchor&quot; href=&quot;#macerator&quot; aria-label=&quot;Anchor link for: macerator&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h3&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;macerator&quot;&gt;macerator&lt;&#x2F;a&gt; is a relative of &lt;code&gt;pulp&lt;&#x2F;code&gt; with a similar design. It was built to power the CPU backend for &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;burn&quot;&gt;burn&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;Compared to &lt;code&gt;pulp&lt;&#x2F;code&gt; it adds support for code generic over element type, but removes safe access to intrinsics and most of the documentation. There is no attempt at fixed-width vectors.&lt;&#x2F;p&gt;
&lt;p&gt;It also enables SSE4.2 by default and adds optimized codepaths for LoongArch.&lt;&#x2F;p&gt;
&lt;p&gt;This is the only library other than &lt;code&gt;std::simd&lt;&#x2F;code&gt; with some portable operations on &lt;code&gt;f16&lt;&#x2F;code&gt; data, albeit the list of supported operations is very limited. Using it with AVX-512 requires a nightly compiler, while NEON works on stable. It is still rather awkward because the standard library&#x27;s &lt;code&gt;f16&lt;&#x2F;code&gt; is nightly-only and this crate has to get by without it.&lt;&#x2F;p&gt;
&lt;p&gt;It isn&#x27;t used by anything on crates.io other than &lt;code&gt;burn&lt;&#x2F;code&gt;.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;others&quot;&gt;Others&lt;a class=&quot;zola-anchor&quot; href=&quot;#others&quot; aria-label=&quot;Anchor link for: others&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h3&gt;
&lt;p&gt;I&#x27;m excluding SIMD crates made for a single specific project (e.g. &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;jxl_simd&quot;&gt;jxl_simd&lt;&#x2F;a&gt;, &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;pathfinder_simd&quot;&gt;pathfinder_simd&lt;&#x2F;a&gt;) since they are not intended for a general audience. I&#x27;m also excluding crates whose development is primarily AI-driven (e.g. &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;magetypes&quot;&gt;magetypes&lt;&#x2F;a&gt;, &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;simdeez&quot;&gt;simdeez&lt;&#x2F;a&gt;, &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;thermite&quot;&gt;thermite&lt;&#x2F;a&gt;) because I cannot recommend them for production use, especially since the latter two are &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;rust&#x2F;comments&#x2F;1vlx6cg&#x2F;thermite_simd_melt_your_cpu_020_release_complete&#x2F;p37pe28&#x2F;&quot;&gt;disconcertingly&lt;&#x2F;a&gt; &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;arduano&#x2F;simdeez&#x2F;issues&#x2F;131&quot;&gt;buggy&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;safe-access-to-intrinsics&quot;&gt;Safe access to intrinsics&lt;a class=&quot;zola-anchor&quot; href=&quot;#safe-access-to-intrinsics&quot; aria-label=&quot;Anchor link for: safe-access-to-intrinsics&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;Portable SIMD is good, but sometimes you want a very specific instruction that only a certain instruction set has. In that case you have to use intrinsics directly.&lt;&#x2F;p&gt;
&lt;p&gt;Rust v1.87+ allows safely calling platform-specific intrinsics:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;rust&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;#&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;[&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;target_feature&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;enable &lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;=&lt;&#x2F;span&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt; &amp;quot;avx2&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)]&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;fn&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; add_avx2&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt;a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; __m256&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; __m256&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; -&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; __m256&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;    _mm256_add_ps&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt; &#x2F;&#x2F; this is an avx2 intrinsic&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;}&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;There are two caveats:&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;Intrinsics to load data from memory or store it are still &lt;code&gt;unsafe&lt;&#x2F;code&gt; because they operate on raw pointers&lt;&#x2F;li&gt;
&lt;li&gt;The function we defined, &lt;code&gt;add_avx2&lt;&#x2F;code&gt;, still requires an &lt;code&gt;unsafe&lt;&#x2F;code&gt; block to call from a function not annotated with &lt;code&gt;#[target_feature(enable = &quot;avx2&quot;)]&lt;&#x2F;code&gt;&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;p&gt;But there are established solutions for both:&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;Use safe wrappers for loads&#x2F;stores that add bounds checks, which the optimizer then &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;shnatsel.medium.com&#x2F;how-to-avoid-bounds-checks-in-rust-without-unsafe-f65e618b4c1e&quot;&gt;trivially removes&lt;&#x2F;a&gt; from machine code so no performance is lost.&lt;&#x2F;li&gt;
&lt;li&gt;Check if a CPU feature is available at runtime and &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;shnatsel.github.io&#x2F;safe-simd-in-rust-even-on-the-inside&#x2F;#lemma-cpu-feature-tokens&quot;&gt;encode it in a type-level token&lt;&#x2F;a&gt;, then use that to call functions requiring those features safely.&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;p&gt;Everything on this list is various implementations of these two ideas.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;archmage&quot;&gt;archmage&lt;a class=&quot;zola-anchor&quot; href=&quot;#archmage&quot; aria-label=&quot;Anchor link for: archmage&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h3&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;archmage&quot;&gt;archmage&lt;&#x2F;a&gt; provides the CPU feature tokens and uses the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;safe_unaligned_simd&quot;&gt;safe_unaligned_simd&lt;&#x2F;a&gt; crate for safe load&#x2F;store wrappers.&lt;&#x2F;p&gt;
&lt;p&gt;Its centerpiece is the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;docs.rs&#x2F;archmage&#x2F;0.9.28&#x2F;archmage&#x2F;attr.arcane.html&quot;&gt;&lt;code&gt;#[arcane]&lt;&#x2F;code&gt; procedural macro&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;You get a selection of predefined SIMD levels: the usual suspects of SSE2&#x2F;SSE4.2&#x2F;AVX2, and there are two different levels of AVX-512: the early slow implementations, and Ice Lake and later which is actually useful, at your option. On ARM there&#x27;s baseline NEON plus a couple of extension levels.&lt;&#x2F;p&gt;
&lt;p&gt;No support for 32-bit x86 (you always get the scalar fallback), but that&#x27;s not a big deal in 2026.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;fearless-simd-1&quot;&gt;fearless_simd&lt;a class=&quot;zola-anchor&quot; href=&quot;#fearless-simd-1&quot; aria-label=&quot;Anchor link for: fearless-simd-1&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h3&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;fearless_simd&quot;&gt;fearless_simd&lt;&#x2F;a&gt; gives you basically the same tools as &lt;code&gt;archmage&lt;&#x2F;code&gt; via its &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;docs.rs&#x2F;fearless_simd&#x2F;latest&#x2F;fearless_simd&#x2F;macro.kernel.html&quot;&gt;&lt;code&gt;kernel!&lt;&#x2F;code&gt; macro&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;This is a declarative macro, not a procedural one. This improves build times, but unlike &lt;code&gt;archmage&lt;&#x2F;code&gt; it doesn&#x27;t support annotating generic or const-generic functions with it. Intrinsics and generics don&#x27;t gel anyway, so it&#x27;s usually not a big deal.&lt;&#x2F;p&gt;
&lt;p&gt;It doesn&#x27;t bundle &lt;code&gt;safe_unaligned_simd&lt;&#x2F;code&gt; since safe loads can be done through its portable SIMD abstraction, but you can pull it yourself if you really want to spell loads as &lt;code&gt;_mm256_loadu_epi64()&lt;&#x2F;code&gt;, usually for porting existing code written like that.&lt;&#x2F;p&gt;
&lt;p&gt;The SIMD levels are the same as for the portable SIMD abstraction. So you don&#x27;t get to opt in to early, slow AVX-512 if you really know what you&#x27;re doing, or access NEON&#x27;s non-baseline extensions like &lt;code&gt;aes&lt;&#x2F;code&gt; or &lt;code&gt;bf16&lt;&#x2F;code&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;On the upside, you can easily &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;linebender&#x2F;fearless_simd&#x2F;blob&#x2F;main&#x2F;fearless_simd&#x2F;examples&#x2F;srgb.rs&quot;&gt;mix and match portable SIMD and intrinsics&lt;&#x2F;a&gt;. This lets you write most of the algorithm in portable SIMD, and use a handful of intrinsics only where they are really needed.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;pulp-1&quot;&gt;pulp&lt;a class=&quot;zola-anchor&quot; href=&quot;#pulp-1&quot; aria-label=&quot;Anchor link for: pulp-1&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h3&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;pulp&quot;&gt;pulp&lt;&#x2F;a&gt; is deceptively powerful in this regard.&lt;&#x2F;p&gt;
&lt;p&gt;If you want to use intrinsics that aren&#x27;t part of any SIMD level, such as &lt;code&gt;_mm_aesenc_si128&lt;&#x2F;code&gt; from the &lt;code&gt;aes&lt;&#x2F;code&gt; feature, this is the best (and only) way to do it safely without rolling your own SIMD feature tokens.&lt;&#x2F;p&gt;
&lt;p&gt;Unfortunately it &lt;strong&gt;does not document how to do that.&lt;&#x2F;strong&gt;&lt;&#x2F;p&gt;
&lt;p&gt;If you look up the docs, you&#x27;ll find structs named after various CPU features, e.g. &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;docs.rs&#x2F;pulp&#x2F;latest&#x2F;pulp&#x2F;core_arch&#x2F;x86&#x2F;struct.Avx512ifma.html&quot;&gt;Avx512ifma&lt;&#x2F;a&gt;, with a way to construct it and with the intrinsics corresponding to that CPU feature on it. So you&#x27;d think you just construct it and call the function, right? &lt;strong&gt;Wrong.&lt;&#x2F;strong&gt;&lt;&#x2F;p&gt;
&lt;p&gt;You can do that, and it works, but performance is awful. You are calling an intrinsic that requires extra CPU features from a function that isn&#x27;t guaranteed to have them (remember, the check for the CPU feature can fail), so the intrinsic has to be in its own separate function. And now you are paying function call overhead - several instructions - to call a single instruction. &quot;Several instructions&quot; is a lot more than one, so the function call overhead dominates and performance plummets.&lt;&#x2F;p&gt;
&lt;p&gt;What you have to do instead is create a context with all required features enabled in it, and then call a bunch of intrinsics from that context. Like this:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;rust&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #DF8E1D;&quot;&gt;pulp&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;simd_type!&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;    pub struct&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Ifma&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;        pub&lt;&#x2F;span&gt;&lt;span&gt; ifma&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt; &amp;quot;avx512ifma&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;    }&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;}&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;if let&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Some&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;isa&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Ifma&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;try_new&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;() {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    isa&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;vectorize&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;        #&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;[&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;inline&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;always&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)]&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #179299;&quot;&gt;        ||&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;            &#x2F;&#x2F; Put the entire hot loop here.&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;            &#x2F;&#x2F; isa.ifma._mm512_madd52lo_epu64(...)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;        },&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;    );&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;}&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;See &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;Shnatsel&#x2F;pulp-intrinsic-access-example&quot;&gt;here&lt;&#x2F;a&gt; for a more complete example you can actually run.&lt;&#x2F;p&gt;
&lt;p&gt;For completeness, I should mention that a similar feature was proposed for the &lt;code&gt;fearless_simd&lt;&#x2F;code&gt; repository as a separate, independent crate. It was &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;linebender&#x2F;fearless_simd&#x2F;pull&#x2F;108&quot;&gt;fully implemented&lt;&#x2F;a&gt;, but nobody stepped up to actually maintain it, so it was never merged. If something irks you about &lt;code&gt;pulp&lt;&#x2F;code&gt;, try that instead and see if you&#x27;re willing to take it over.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-state-of-intrinsics&quot;&gt;The state of intrinsics&lt;a class=&quot;zola-anchor&quot; href=&quot;#the-state-of-intrinsics&quot; aria-label=&quot;Anchor link for: the-state-of-intrinsics&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;SIMD intrinsics underpin all SIMD code except for &lt;code&gt;std::simd&lt;&#x2F;code&gt; and autovectorization. They&#x27;re quite straightforward, too: intrinsics are supposed to clearly map to specific CPU instructions. Given how simple and important they are, you&#x27;d expect them to work really well.&lt;&#x2F;p&gt;
&lt;p&gt;They don&#x27;t. Not in Rust, not in C++, not in C.&lt;&#x2F;p&gt;
&lt;p&gt;There is a fundamental tension between &quot;give me this exact instruction&quot; and compiler optimizations. If you have the compiler treat SIMD intrinsics as pure black boxes, you end up with inefficiencies elsewhere.&lt;&#x2F;p&gt;
&lt;p&gt;For example, a real bug I&#x27;ve run into on ARM is that &lt;code&gt;u32x4::from([1,2,3,4])&lt;&#x2F;code&gt; was slow. This is literally loading a constant, and &lt;code&gt;u32x4&lt;&#x2F;code&gt; has the exact same memory layout as an array of four &lt;code&gt;u32&lt;&#x2F;code&gt;, so it should be &lt;em&gt;really&lt;&#x2F;em&gt; cheap - just a single load.&lt;&#x2F;p&gt;
&lt;p&gt;It turns out that the underlying ARM load intrinsic, &lt;code&gt;vld1_u32_x4&lt;&#x2F;code&gt;, was implemented as a black-box operation in the compiler, so all LLVM saw was a black-box operation on some on-stack value. The generated assembly first loaded the constant into registers, then placed it onto the stack, and then loaded it back into registers through the black-box &lt;code&gt;vld1_u32_x4&lt;&#x2F;code&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;The fix was to &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;stdarch&#x2F;pull&#x2F;2004&quot;&gt;drop the black-box implementation for &lt;code&gt;vld1_u32_x4&lt;&#x2F;code&gt;&lt;&#x2F;a&gt; and make it into a compatibility wrapper for regular loads that the compiler can properly optimize. Many thanks to &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;folkertdev&quot;&gt;Folkert de Vries&lt;&#x2F;a&gt;, a Rust stdarch maintainer, for helping investigate this and implementing the fix.&lt;&#x2F;p&gt;
&lt;p&gt;So let&#x27;s just turn all intrinsics into wrappers for regular compiler ops, right? I wish.&lt;&#x2F;p&gt;
&lt;p&gt;That &lt;code&gt;vld1_u32_x4&lt;&#x2F;code&gt; isn&#x27;t really a black box. It&#x27;s a hardware operation that nobody has written optimization passes for yet. And if you want to make some exotic operation into basic blocks comprehensible to the compiler (or just &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;discourse.llvm.org&#x2F;t&#x2F;rfc-ir-ability-to-shuffle-vectors-with-dynamic-mask&#x2F;91282&quot;&gt;abstract over common behavior of slightly different hardware instructions&lt;&#x2F;a&gt;), the operation needs to be made up of &lt;em&gt;several&lt;&#x2F;em&gt; basic blocks. This is really attractive for Rust because it makes supporting backends other than LLVM easier, but Clang has also been moving in this direction.&lt;&#x2F;p&gt;
&lt;p&gt;But then to emit the desired operation from several building blocks, you need the optimizer to recombine them into a single instruction. This can be easily messed up by unrelated optimization patterns that e.g. reorder these blocks and break the pattern-matching. So these optimizations sometimes &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;rust&#x2F;issues&#x2F;159831&quot;&gt;work on simple test cases but break in real-world code&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;So &lt;strong&gt;SIMD intrinsics are stuck in an endless tug of war&lt;&#x2F;strong&gt; between lowering into the expected instructions and working with the expected compiler optimizations.&lt;&#x2F;p&gt;
&lt;p&gt;And on top of the fundamental limitations, there are also compiler instruction selection bugs. I&#x27;ve run into LLVM seeing a 512-bit vector shuffle operation with constant indices and going &quot;oh, I know, I can optimize this!&quot; except its &quot;optimization&quot; uses SSE4.2-era operations and &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;rust&#x2F;issues&#x2F;156891&quot;&gt;is far, far slower&lt;&#x2F;a&gt; than just running the actual shuffle instruction I asked for. (That one&#x27;s fixed in LLVM 23, following my report). Or lowering an intrinsic whose sole purpose is efficient encoding &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;rust&#x2F;issues&#x2F;156946&quot;&gt;into a less efficient encoding&lt;&#x2F;a&gt;. I literally have &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;linebender&#x2F;fearless_simd&#x2F;issues&#x2F;281&quot;&gt;a list of such compiler bugs&lt;&#x2F;a&gt; I&#x27;ve found. The Rust-specific ones got fixed after I reported them, but there is a bunch of LLVM bugs affecting all of Rust, C++ and C still unfixed.&lt;&#x2F;p&gt;
&lt;p&gt;And in case you&#x27;re wondering - no, this isn&#x27;t just an LLVM problem. GCC has similar issues, and MSVC is noticeably worse at this than either of the major open-source compilers.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;Bonus fun fact:&lt;&#x2F;strong&gt; Intel &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;rust&#x2F;issues&#x2F;158196&quot;&gt;forgot to include some AVX-512 instructions&lt;&#x2F;a&gt; into their searchable &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;www.intel.com&#x2F;content&#x2F;www&#x2F;us&#x2F;en&#x2F;docs&#x2F;intrinsics-guide&#x2F;index.html&quot;&gt;Intrinsics Guide&lt;&#x2F;a&gt;, so most compilers didn&#x27;t implement them, and now you can&#x27;t reach those instructions from high-level languages. I&#x27;ve &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;stdarch&#x2F;pull&#x2F;2197&quot;&gt;contributed them to rustc&lt;&#x2F;a&gt; but they haven&#x27;t shipped on stable yet.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;inline-assembly&quot;&gt;Inline assembly&lt;a class=&quot;zola-anchor&quot; href=&quot;#inline-assembly&quot; aria-label=&quot;Anchor link for: inline-assembly&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h3&gt;
&lt;p&gt;You&#x27;d think you could outsmart the compiler that way and bypass all the issues with intrinsics. And you kinda sorta can, except now your inline assembly block is a real honest-to-goodness black box.&lt;&#x2F;p&gt;
&lt;p&gt;Not only are all the optimization issues back with a vengeance, but entering and exiting the inline assembly block has some overhead.&lt;&#x2F;p&gt;
&lt;p&gt;This works okay if you want to write a large-ish function in it and are willing to sacrifice compiler optimizations, but isn&#x27;t profitable if you just want to use an instruction or two.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;compiler-feature-wishlist&quot;&gt;Compiler feature wishlist&lt;a class=&quot;zola-anchor&quot; href=&quot;#compiler-feature-wishlist&quot; aria-label=&quot;Anchor link for: compiler-feature-wishlist&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;I&#x27;ll keep this brief, in the order of importance:&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;code&gt;min_generic_const_args&lt;&#x2F;code&gt; would allow using arrays in conjunction with hardware-width SIMD vectors. There are workarounds (just use a huge array, use &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;generic-array&quot;&gt;&lt;code&gt;generic-array&lt;&#x2F;code&gt; crate&lt;&#x2F;a&gt;, or &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;smu160&#x2F;PhastFT&#x2F;blob&#x2F;7bbbfa5bbac8681af7d1abf6fb02990d8eacb552&#x2F;src&#x2F;algorithms&#x2F;bravo.rs#L76-L294&quot;&gt;use macros instead of generics&lt;&#x2F;a&gt;) but all are partial and&#x2F;or ugly.&lt;&#x2F;li&gt;
&lt;li&gt;With the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;rfcs&#x2F;pull&#x2F;3525&quot;&gt;Struct Target Features RFC&lt;&#x2F;a&gt;, &lt;code&gt;fearless_simd&lt;&#x2F;code&gt;&#x2F;&lt;code&gt;pulp&lt;&#x2F;code&gt;&#x2F;&lt;code&gt;macerator&lt;&#x2F;code&gt; would no longer need &lt;code&gt;#[simd]&lt;&#x2F;code&gt; annotations on functions. It&#x27;s less boilerplate, but most importantly you can no longer accidentally forget to put them there and cause performance to drop.&lt;&#x2F;li&gt;
&lt;li&gt;We need a way to make iterators &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;linebender&#x2F;fearless_simd&#x2F;issues&#x2F;380&quot;&gt;not conflict with multiversioning&lt;&#x2F;a&gt;. The &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;rfcs&#x2F;pull&#x2F;3525&quot;&gt;Struct Target Features RFC&lt;&#x2F;a&gt; would solve this too. Alternatively an equivalent of &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;gcc.gnu.org&#x2F;onlinedocs&#x2F;gcc-12.5.0&#x2F;gcc&#x2F;Common-Function-Attributes.html&quot;&gt;GCC&#x27;s &lt;code&gt;__attribute__((flatten))&lt;&#x2F;code&gt;&lt;&#x2F;a&gt; should do the trick, but that requires either even more boilerplate or proc macros.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;code&gt;generic_const_args&lt;&#x2F;code&gt; (not &lt;code&gt;min&lt;&#x2F;code&gt;) would &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;gist.github.com&#x2F;valadaptive&#x2F;b0eedd611749fae45993f2e73036d25f&quot;&gt;significantly improve build times&lt;&#x2F;a&gt; when wrapping certain intrinsics into portable abstractions, and make using &lt;code&gt;std::simd&lt;&#x2F;code&gt; much nicer.&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;p&gt;In the standard library I&#x27;d love to see &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;rust&#x2F;issues&#x2F;159831#issuecomment-5552538552&quot;&gt;crater-like verification of changes to intrinsics&lt;&#x2F;a&gt;, and &lt;code&gt;std::simd&lt;&#x2F;code&gt; available on stable so that ecosystem crates would delete most of their code and gain support for all the obscure platforms.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;a class=&quot;zola-anchor&quot; href=&quot;#conclusion&quot; aria-label=&quot;Anchor link for: conclusion&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;Support for SIMD in the Rust ecosystem has matured a great deal.&lt;&#x2F;p&gt;
&lt;p&gt;While some things could still be improved (notably trigonometry), recent advances in compiler features and the library ecosystem made Rust attractive for SIMD code even when memory safety is not a hard requirement.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;bonus-round-the-state-of-hardware&quot;&gt;Bonus round: The state of hardware&lt;a class=&quot;zola-anchor&quot; href=&quot;#bonus-round-the-state-of-hardware&quot; aria-label=&quot;Anchor link for: bonus-round-the-state-of-hardware&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;After writing portable SIMD code for 5 different instruction sets, I have &lt;em&gt;opinions.&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;h3 id=&quot;x86&quot;&gt;x86&lt;a class=&quot;zola-anchor&quot; href=&quot;#x86&quot; aria-label=&quot;Anchor link for: x86&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h3&gt;
&lt;p&gt;It&#x27;s a mess, and we have Intel to thank for it.&lt;&#x2F;p&gt;
&lt;p&gt;AVX2 is simultaneously the most common and the most cursed instruction set I&#x27;ve ever had to work with.&lt;&#x2F;p&gt;
&lt;p&gt;Whenever I try to implement a simple, straightforward SIMD operation, half the time AVX-512 and NEON have it natively, but AVX2 needs complex and slow emulation. The fact that instead of a proper 256-bit ISA it&#x27;s more like two 128-bit execution units smushed together really doesn&#x27;t help.&lt;&#x2F;p&gt;
&lt;p&gt;But it gets worse. Despite CPUs with AVX2 launching in 2013, the most recent Intel CPU without AVX2 &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Tremont_(microarchitecture)&quot;&gt;launched in 2021&lt;&#x2F;a&gt;! So you can&#x27;t even count on having AVX2, good luck making do with SSE4.2 from... &lt;em&gt;checks notes...&lt;&#x2F;em&gt; 2008!&lt;&#x2F;p&gt;
&lt;p&gt;That&#x27;s how you get 15% of x86 CPUs in the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;firefoxgraphics.github.io&#x2F;telemetry&#x2F;#view=system&quot;&gt;Firefox hardware survey&lt;&#x2F;a&gt; still not having AVX2 in 2026. Have fun writing and maintaining SSE4.2 codepaths just for them! &lt;em&gt;What year is this?!&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;p&gt;Intel launched AVX-512 in 2015, which fixed much of the insanity of AVX2, and then... just didn&#x27;t put it into any CPUs? It was only really present on the server, everyone else was stuck with AVX2 or even just SSE4.2. So Intel has &lt;strong&gt;three completely different SIMD extensions&lt;&#x2F;strong&gt; all existing at the same time!&lt;&#x2F;p&gt;
&lt;p&gt;But wait, it gets even worse!&lt;&#x2F;p&gt;
&lt;p&gt;Running AVX-512 instructions on early Intel CPUs with AVX-512, even on a single core, &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;stackoverflow.com&#x2F;a&#x2F;56861355&#x2F;585725&quot;&gt;reduces CPU frequency of &lt;em&gt;all&lt;&#x2F;em&gt; cores&lt;&#x2F;a&gt;. An AVX-512 workload anywhere hurts performance of the entire rest of the chip! Ironically, AVX-512 only appeared in high-end CPUs with lots of cores where this kind of fallout is &lt;em&gt;especially&lt;&#x2F;em&gt; bad!&lt;&#x2F;p&gt;
&lt;p&gt;You&#x27;d think you could still benefit from AVX-512 on these CPUs if you run it on all cores at once for a long time, but then you end up bottlenecked on memory anyway, and whatever the CPU is doing becomes irrelevant. So on those CPUs AVX-512 doesn&#x27;t actually give you any performance and often hurts it, except in artificial microbenchmarks that don&#x27;t touch memory.&lt;&#x2F;p&gt;
&lt;p&gt;This is such a shame, because AVX-512 is such a big improvement on AVX2 otherwise. Forget the 512-bit width, just give me the sane set of supported operations!&lt;&#x2F;p&gt;
&lt;p&gt;This downclocking behavior was only fixed in 2019, in the Ice Lake architecture. Not fully, but enough to make AVX-512 profitable overall. This is why &lt;code&gt;fearless_simd&lt;&#x2F;code&gt; only supports AVX-512 on Ice Lake and later, and on AMD which never had downclocking issues to begin with.&lt;&#x2F;p&gt;
&lt;p&gt;AMD showed how badly Intel messed this up by releasing Zen 4, which didn&#x27;t even have hardware 512-bit operations. It mapped most 512-bit operations to 256-bit execution units, and &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;www.phoronix.com&#x2F;review&#x2F;zen4-avx512-7700x&quot;&gt;still smoked&lt;&#x2F;a&gt; Intel&#x27;s native 512-bit hardware in benchmarks. Zen 5 with its native 512-bit hardware sealed the deal. No wonder Intel is struggling recently.&lt;&#x2F;p&gt;
&lt;p&gt;Not that AMD is blameless. They&#x27;re the reason we can&#x27;t use scatter&#x2F;gather instructions because in Zen they&#x27;re not implemented natively in hardware, and end up being slower than issuing lots of small loads. Intel made scatter&#x2F;gather slower than scalar loads in early AVX2 CPUs too, but they got their act together eventually, sort of; AMD didn&#x27;t even try.&lt;&#x2F;p&gt;
&lt;p&gt;LLVM sometimes emits scatter&#x2F;gather instructions for AVX-512 when autovectorizing code, which &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;llvm&#x2F;llvm-project&#x2F;issues&#x2F;70259&quot;&gt;hurts performance by 1.75x on Intel&lt;&#x2F;a&gt; and &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;llvm&#x2F;llvm-project&#x2F;issues&#x2F;91370&quot;&gt;by 4x on AMD&lt;&#x2F;a&gt;, so I&#x27;m not even sure why LLVM even bothers. I believe you need to pass &lt;code&gt;-C target-cpu=&lt;&#x2F;code&gt; to hit this, but I haven&#x27;t extensively tested it.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;store.steampowered.com&#x2F;hwsurvey&#x2F;Steam-Hardware-Software-Survey-Welcome-to-Steam&quot;&gt;Steam hardware survey&lt;&#x2F;a&gt; shows that 23.9% of systems have AVX-512. This is skewed towards high-end&#x2F;gaming systems; for example, the 15% of systems with only SSE4.2 from the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;firefoxgraphics.github.io&#x2F;telemetry&#x2F;#view=system&quot;&gt;Firefox graphics survey&lt;&#x2F;a&gt; are at only 2% here. It also shows a breakdown by AVX-512 optional features, and from them we can infer that 23.95% (&lt;code&gt;avx512vnni&lt;&#x2F;code&gt;) minus 23.90% (baseline &lt;code&gt;avx512f&lt;&#x2F;code&gt;) equals -0.05% of systems with awfully slow AVX-512. Your guess on why this percentage is negative is as good as mine.&lt;&#x2F;p&gt;
&lt;p&gt;So at least on desktop, the broken AVX-512 is nonexistent, which means you don&#x27;t have to worry about it. Therefore setting Ice Lake as a requirement for AVX-512 loses you nothing and gains some useful instructions, but using the baseline AVX-512 isn&#x27;t awful either.&lt;&#x2F;p&gt;
&lt;p&gt;There is no public data on the prevalence of Skylake servers, where AVX-512 is present but degrades performance. Using Ice Lake as a requirement for AVX-512 so that Skylake uses AVX2 should prevent that degradation.&lt;&#x2F;p&gt;
&lt;p&gt;Intel was &lt;em&gt;this&lt;&#x2F;em&gt; close to messing things up again by replacing AVX-512 with AVX10, which is AVX-512 but with either 256-bit or 512-bit vectors, and you don&#x27;t know which ones. But AMD&#x27;s clearly superior design that just maps 512-bit vectors onto 256-bit hardware averted this disaster and forced Intel back into a sane programming model. Whew. Thanks, AMD.&lt;&#x2F;p&gt;
&lt;p&gt;This year, at long last, AVX-512 is becoming mandatory in upcoming Intel CPUs via its rebranding into AVX 10.2. Which is what we wanted all along.&lt;&#x2F;p&gt;
&lt;p&gt;Well, not the rebranding.&lt;&#x2F;p&gt;
&lt;p&gt;Also, doing math on floating-point values very close to zero &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;gitlab.in2p3.fr&#x2F;CTA-LAPP&#x2F;COURS&#x2F;GRAY_SCOTT_REVOLUTIONS&#x2F;GrayScottRevolution&#x2F;-&#x2F;wikis&#x2F;uploads&#x2F;4-ComputingPrecision&#x2F;Subnormal.pdf&quot;&gt;makes performance plummet&lt;&#x2F;a&gt;. AMD is about 2x slower on those, but on Intel you get a 30x slowdown.&lt;&#x2F;p&gt;
&lt;p&gt;Somehow, every time I learn something horrifying about SIMD, it&#x27;s always Intel&#x27;s fault.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;arm&quot;&gt;ARM&lt;a class=&quot;zola-anchor&quot; href=&quot;#arm&quot; aria-label=&quot;Anchor link for: arm&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h3&gt;
&lt;p&gt;Everything that AVX2 got wrong, 64-bit NEON gets right.&lt;&#x2F;p&gt;
&lt;p&gt;With only 128-bit vectors you&#x27;d think it is an equivalent of SSE4.2, but it&#x27;s actually closer to AVX2.&lt;&#x2F;p&gt;
&lt;p&gt;NEON has the same register space as AVX2, which is often the limiting factor in practice. And instead of making you deal with two 128-bit execution units side by side explicitly, beefy ARM cores transparently run 128-bit operations in parallel via &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Instruction-level_parallelism&quot;&gt;instruction-level parallelism&lt;&#x2F;a&gt;, while cheap power-constrained cores can still execute them one by one.&lt;&#x2F;p&gt;
&lt;p&gt;NEON also adds just enough instructions larger than 128 bits to make common operations Just Work. In my tests with SIMD base64 decoding, the same algorithm runs 1.5x to 2x faster on NEON than on AVX2 (but still 2x slower than AVX-512).&lt;&#x2F;p&gt;
&lt;p&gt;And all of this in a single, simple programming model instead of several different ones. And it&#x27;s mandatory in 64-bit ARM chips, with no need for multiversioning!&lt;&#x2F;p&gt;
&lt;p&gt;The only criticism I can level at Aarch64 NEON is that a single chip has two kinds of cores (&quot;performance&quot; and &quot;efficiency&quot;) with completely different execution characteristics, so an instruction sequence that is fast on performance cores is slow on efficiency cores, and vice versa. So even if you know a specific CPU you&#x27;re targeting, you can&#x27;t really select an optimal implementation, it&#x27;s all trade-offs! And when you consider the diversity of ARM CPUs out there, it only gets worse. Fortunately, NEON has enough operations implemented directly as hardware instructions with reasonable performance to prevent this from turning into a total nightmare.&lt;&#x2F;p&gt;
&lt;p&gt;Meanwhile SVE is pretty much useless. SVE2 is now mandatory in ARM CPUs, but it&#x27;s implemented at 128-bit width even in high-end server chips, so it&#x27;s just an awkward NEON with extra steps. Technically there was 256-bit SVE in a single generation of server ARM chips for the cloud, but that&#x27;s not SVE2, and in the cloud you just use AVX-512 instead anyway. Maybe we need to wait another decade or so to see its genius, when 256-bit SVE2 hardware becomes widespread, but for now - don&#x27;t bother.&lt;&#x2F;p&gt;
&lt;p&gt;ARM doesn&#x27;t have an answer to AVX-512, with its 4x larger register space and four 512-bit execution units for crunching through 2048 bits at once. But considering their target markets, and that good AVX-512 &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;www.numberworld.org&#x2F;blogs&#x2F;2024_8_7_zen5_avx512_teardown&#x2F;&quot;&gt;ends up bottlenecked by memory bandwidth anyway&lt;&#x2F;a&gt;, I&#x27;m not convinced ARM needs one. The main benefit of AVX-512 is a much wider range of supported operations, which NEON already has.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;risc-v&quot;&gt;RISC-V&lt;a class=&quot;zola-anchor&quot; href=&quot;#risc-v&quot; aria-label=&quot;Anchor link for: risc-v&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h3&gt;
&lt;p&gt;RISC-V vectors (RVV) are completely irrelevant because vector-capable RISC-V hardware is completely irrelevant. RISC-V is dominating in cheap microcontrollers, but the performance&#x2F;price ratio for vector-capable RISC-V hardware in 2026 is abysmal. Maybe Tenstorrent will change that in 2028 or so when they actually tape out some silicon, but you definitely don&#x27;t have to worry about it in 2026.&lt;&#x2F;p&gt;
&lt;p&gt;Even if decent hardware existed, the way the spec is written makes certain crucial instructions &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;riscv&#x2F;riscv-profiles&#x2F;issues&#x2F;187&quot;&gt;unusable to compilers&lt;&#x2F;a&gt;. This alone degrades performance to ridiculous levels unless you mess with obscure compiler flags.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;others-1&quot;&gt;Others&lt;a class=&quot;zola-anchor&quot; href=&quot;#others-1&quot; aria-label=&quot;Anchor link for: others-1&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h3&gt;
&lt;p&gt;The remaining SIMD-capable architectures are so obscure that you shouldn&#x27;t bother thinking about them unless someone is paying you to work on them - IBM and&#x2F;or banks in case of POWER and s390x, or the Chinese government in the case of LoongArch. &lt;code&gt;std::simd&lt;&#x2F;code&gt; still runs CI on 64-bit SPARC, but nobody&#x27;s going to pay you to support that.&lt;&#x2F;p&gt;
</content>
	</entry>
	<entry xml:lang="en">
		<title>Implementing FMA and finding bugs in C and Rust standard libraries</title>
        <author>
            <name>Sergey &quot;Shnatsel&quot; Davidoff</name>
        </author>
		<published>2026-08-20T00:00:00+00:00</published>
		<updated>2026-08-20T00:00:00+00:00</updated>
		<link href="https://shnatsel.github.io/implementing-fma-finding-bugs-in-std/"/>
		<link rel="alternate" href="https://shnatsel.github.io/implementing-fma-finding-bugs-in-std/" type="text/html"/>
		<id>https://shnatsel.github.io/implementing-fma-finding-bugs-in-std/</id>
        <summary type="html">&lt;p&gt;This is a story of how I tried to compute &lt;code&gt;a * b + c&lt;&#x2F;code&gt;, and found that Rust and musl libc get it subtly wrong.&lt;&#x2F;p&gt;</summary>
		<content type="html">&lt;p&gt;This is a story of how I tried to compute &lt;code&gt;a * b + c&lt;&#x2F;code&gt;, and found that Rust and musl libc get it subtly wrong.&lt;&#x2F;p&gt;
&lt;span id=&quot;continue-reading&quot;&gt;&lt;&#x2F;span&gt;
&lt;p&gt;&lt;strong&gt;Fused multiply-add&lt;&#x2F;strong&gt; (FMA) computes &lt;code&gt;a * b + c&lt;&#x2F;code&gt; with only one rounding error instead of two. It&#x27;s an important building block for e.g. trigonometric functions like &lt;code&gt;sin(x)&lt;&#x2F;code&gt; and &lt;code&gt;cos(x)&lt;&#x2F;code&gt; if you want to implement them accurately.&lt;&#x2F;p&gt;
&lt;p&gt;It&#x27;s a basic primitive that is usually implemented in hardware, but there is still some hardware out there that doesn&#x27;t have it. You&#x27;d think it would be something like cheap phones, but no, it&#x27;s Intel. Cheap ARM phones have it and it&#x27;s been required in 64-bit ARM since the very beginning, but Intel has been launching new parts without AVX2 or fused multiply-add &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Tremont_(microarchitecture)&quot;&gt;as recently as 2021&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;(I&#x27;ve come to learn that whenever something is holding SIMD back, it&#x27;s usually Intel.)&lt;&#x2F;p&gt;
&lt;p&gt;15% of machines in the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;firefoxgraphics.github.io&#x2F;telemetry&#x2F;#view=system&quot;&gt;Firefox hardware survey&lt;&#x2F;a&gt; don&#x27;t have AVX2 and hardware FMA that comes with it, so it has to be emulated for precise algorithms built on top of it to work correctly.&lt;&#x2F;p&gt;
&lt;p&gt;On machines without hardware FMA, Rust&#x27;s &lt;code&gt;std::simd&lt;&#x2F;code&gt; gives up and runs scalar FMA on each &lt;code&gt;f32&lt;&#x2F;code&gt; in &lt;code&gt;[f32; 4]&lt;&#x2F;code&gt; individually, which is slow. I wanted to do better in &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;linebender&#x2F;fearless_simd&quot;&gt;fearless_simd&lt;&#x2F;a&gt; and provide an actually vectorized implementation.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;emulating-fma-with-simd&quot;&gt;Emulating FMA with SIMD&lt;a class=&quot;zola-anchor&quot; href=&quot;#emulating-fma-with-simd&quot; aria-label=&quot;Anchor link for: emulating-fma-with-simd&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;Since FMA works on three &lt;code&gt;f32&lt;&#x2F;code&gt; values, the total number of possible inputs is 2 to the 96th power. It would take the world&#x27;s largest supercomputer only 500 years to try them all. We&#x27;ve come a long way! But I need working FMA later this year, so exhaustive verification isn&#x27;t really on the cards. The best I could do is some known values plus some random tests.&lt;&#x2F;p&gt;
&lt;p&gt;I followed the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;guillaume.melquiond.fr&#x2F;doc&#x2F;08-tc.pdf&quot;&gt;2008 paper&lt;&#x2F;a&gt; &quot;Emulation of FMA and correctly-rounded sums: proved algorithms using rounding to odd&quot; by Sylvie Boldo and Guillaume Melquiond, which has a formal proof of correctness in Coq. That way I don&#x27;t have to worry about trying to verify the algorithm myself.&lt;&#x2F;p&gt;
&lt;p&gt;For &lt;code&gt;f32&lt;&#x2F;code&gt;, their algorithm is refreshingly simple: compute &lt;code&gt;a * b + c&lt;&#x2F;code&gt; in &lt;code&gt;f64&lt;&#x2F;code&gt;, then round it to &lt;code&gt;f32&lt;&#x2F;code&gt;. The only caveat is special handling of rounding errors in the conversion, in case the result falls exactly between two representable values.&lt;&#x2F;p&gt;
&lt;p&gt;Translating the algorithm to SIMD was also straightforward: just do all that basic math per-lane. The special handling of rounding is needed very rarely - you hit it less than once in a million when processing values in the [-1, 1) range, so just put it under an &lt;code&gt;if&lt;&#x2F;code&gt; and it&#x27;ll be fine. &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;linebender&#x2F;fearless_simd&#x2F;pull&#x2F;323&quot;&gt;Done!&lt;&#x2F;a&gt;&lt;&#x2F;p&gt;
&lt;p&gt;Benchmarks look great: it&#x27;s 5x faster than &lt;code&gt;std::simd&lt;&#x2F;code&gt;, even bigger than the expected 4x speedup because Rust standard library has to check if FMA is available on the system in every call, and also has to worry about setting floating-point exception flags, both of which add overhead.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;follow-the-white-rabbit&quot;&gt;Follow the White Rabbit&lt;a class=&quot;zola-anchor&quot; href=&quot;#follow-the-white-rabbit&quot; aria-label=&quot;Anchor link for: follow-the-white-rabbit&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;Just a few hours later, a wild &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;awxkee&quot;&gt;&lt;strong&gt;@awxkee&lt;&#x2F;strong&gt;&lt;&#x2F;a&gt; appeared and &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;linebender&#x2F;fearless_simd&#x2F;pull&#x2F;323#issuecomment-5233682925&quot;&gt;posted some inputs&lt;&#x2F;a&gt; on which my implementation diverged from the hardware results.&lt;&#x2F;p&gt;
&lt;p&gt;Moments like these are why I love open source. I have no idea where he came from or how he even found this PR, he&#x27;s never contributed code to Fearless SIMD before. But there it was, a counter-example that broke my translation of a formally verified algorithm.&lt;&#x2F;p&gt;
&lt;p&gt;Turns out I translated the paper into code incorrectly. I forgot to add special handling for values very close to zero, called &lt;strong&gt;subnormal&lt;&#x2F;strong&gt; values, which have a slightly different representation, and many of the usual floating-point &quot;tricks&quot; don&#x27;t work on them. The check that would apply the rounding fixup handled them incorrectly.&lt;&#x2F;p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Aside:&lt;&#x2F;strong&gt; Subnormals are fascinating! Learn all about them &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;numerical-rust-cpu-81b2c3.pages.in2p3.fr&#x2F;19-subnormal-entertainment.html&quot;&gt;here&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;&#x2F;blockquote&gt;
&lt;p&gt;So I looked at the paper more carefully and fixed my implementation to match it more closely. Then I added tests covering the problematic inputs, random tests that try a million subnormals, and random tests that try a million values that should require fixup, for good measure.&lt;&#x2F;p&gt;
&lt;p&gt;Wait... Why are tests still failing, but &lt;em&gt;only&lt;&#x2F;em&gt; on old systems?&lt;&#x2F;p&gt;
&lt;h2 id=&quot;down-the-rabbit-hole&quot;&gt;Down the rabbit hole&lt;a class=&quot;zola-anchor&quot; href=&quot;#down-the-rabbit-hole&quot; aria-label=&quot;Anchor link for: down-the-rabbit-hole&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;Fearless SIMD has CI configured to run tests in an emulator for every supported SIMD level, on top of running them on the host. This is the only way to test AVX-512 codepaths on CI. It would also fail if we tried to use instructions not available in hardware, although the Rust compiler &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;shnatsel.github.io&#x2F;safe-simd-in-rust-even-on-the-inside&#x2F;&quot;&gt;already verifies this for us&lt;&#x2F;a&gt;. It&#x27;s just good practice to test the codepaths on the hardware where they&#x27;d actually run.&lt;&#x2F;p&gt;
&lt;p&gt;The built-in &lt;code&gt;f32::mul_add&lt;&#x2F;code&gt; in the Rust standard library &lt;strong&gt;has the exact same bug!&lt;&#x2F;strong&gt;&lt;&#x2F;p&gt;
&lt;p&gt;It fails to properly handle rounding for subnormals. As does &lt;code&gt;std::simd&lt;&#x2F;code&gt;. But the software implementation is only invoked for systems that don&#x27;t have FMA in hardware, so without the emulator we wouldn&#x27;t have noticed.&lt;&#x2F;p&gt;
&lt;p&gt;So I &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;compiler-builtins&#x2F;issues&#x2F;1262&quot;&gt;report the bug&lt;&#x2F;a&gt; to the Rust standard library, then transcribe the formally verified algorithm from &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;guillaume.melquiond.fr&#x2F;doc&#x2F;08-tc.pdf&quot;&gt;the paper&lt;&#x2F;a&gt; again (hopefully correctly) to insulate the scalar fallback inside &lt;code&gt;fearless_simd&lt;&#x2F;code&gt; from it. Tests pass.&lt;&#x2F;p&gt;
&lt;p&gt;Here&#x27;s the algorithm, it&#x27;s not that scary:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;rust&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;fn&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; scalar_mul_add_precise_f32&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt;a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt; f32&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt; f32&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; c&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt; f32&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; -&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt; f32&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;    let&lt;&#x2F;span&gt;&lt;span&gt; product&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; (&lt;&#x2F;span&gt;&lt;span&gt;a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt; as f64&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; *&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; (&lt;&#x2F;span&gt;&lt;span&gt;b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt; as f64&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;);&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;    let&lt;&#x2F;span&gt;&lt;span&gt; c&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span&gt; c&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt; as f64&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;    let mut&lt;&#x2F;span&gt;&lt;span&gt; sum&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span&gt; product&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; +&lt;&#x2F;span&gt;&lt;span&gt; c&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;    if&lt;&#x2F;span&gt;&lt;span&gt; sum&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;is_finite&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;() {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;        let&lt;&#x2F;span&gt;&lt;span&gt; virtual_sum&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span&gt; sum&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; -&lt;&#x2F;span&gt;&lt;span&gt; product&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;        let&lt;&#x2F;span&gt;&lt;span&gt; rounding_error&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; (&lt;&#x2F;span&gt;&lt;span&gt;product&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; -&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; (&lt;&#x2F;span&gt;&lt;span&gt;sum&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; -&lt;&#x2F;span&gt;&lt;span&gt; virtual_sum&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;))&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; +&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; (&lt;&#x2F;span&gt;&lt;span&gt;c&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; -&lt;&#x2F;span&gt;&lt;span&gt; virtual_sum&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;);&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;        let&lt;&#x2F;span&gt;&lt;span&gt; sum_bits&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span&gt; sum&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;to_bits&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;();&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;        if&lt;&#x2F;span&gt;&lt;span&gt; rounding_error&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; !=&lt;&#x2F;span&gt;&lt;span style=&quot;color: #FE640B;&quot;&gt; 0&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #FE640B;&quot;&gt;0&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; &amp;amp;&amp;amp;&lt;&#x2F;span&gt;&lt;span&gt; sum_bits&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; &amp;amp;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #FE640B;&quot;&gt; 1&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; ==&lt;&#x2F;span&gt;&lt;span style=&quot;color: #FE640B;&quot;&gt; 0&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;            let&lt;&#x2F;span&gt;&lt;span&gt; corrected_bits&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt; if&lt;&#x2F;span&gt;&lt;span&gt; sum&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;is_sign_negative&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;()&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; ==&lt;&#x2F;span&gt;&lt;span&gt; rounding_error&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;is_sign_negative&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;() {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;                sum_bits&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;wrapping_add&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #FE640B;&quot;&gt;1&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;            }&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt; else&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;                sum_bits&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;wrapping_sub&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #FE640B;&quot;&gt;1&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;            };&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;            sum&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt; f64&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;from_bits&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;corrected_bits&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;);&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;        }&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;    }&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    sum&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt; as f32&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;}&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Just the Rust standard library left to fix.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;how-deep-does-this-go&quot;&gt;How deep does this go?&lt;a class=&quot;zola-anchor&quot; href=&quot;#how-deep-does-this-go&quot; aria-label=&quot;Anchor link for: how-deep-does-this-go&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;Fixing the standard library is trickier because it not only computes the result but also sets the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;FLAGS_register&quot;&gt;floating-point status flags&lt;&#x2F;a&gt;. It also has this codepath not just for &lt;code&gt;f32&lt;&#x2F;code&gt; using &lt;code&gt;f64&lt;&#x2F;code&gt; but also for &lt;code&gt;f64&lt;&#x2F;code&gt; using &lt;code&gt;f128&lt;&#x2F;code&gt; on hardware where that&#x27;s available.&lt;&#x2F;p&gt;
&lt;p&gt;So I write tests, for &lt;code&gt;f32&lt;&#x2F;code&gt; and &lt;code&gt;f64&lt;&#x2F;code&gt; both, then &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;compiler-builtins&#x2F;pull&#x2F;1270&quot;&gt;fix the code&lt;&#x2F;a&gt; to the best of my understanding.&lt;&#x2F;p&gt;
&lt;p&gt;Wait, what&#x27;s this at the top of the file?&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;rust&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;&#x2F;* origin: musl src&#x2F;math&#x2F;fmaf.c Ported to generic Rust algorithm in 2025, TG. *&#x2F;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;I wonder if...&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;c&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;if&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; ((&lt;&#x2F;span&gt;&lt;span&gt;u.i &lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;&amp;amp;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt; 0x&lt;&#x2F;span&gt;&lt;span style=&quot;color: #FE640B;&quot;&gt;1fffffff&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; !=&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt; 0x&lt;&#x2F;span&gt;&lt;span style=&quot;color: #FE640B;&quot;&gt;10000000&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt; &#x2F;* not a halfway case *&#x2F;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Yep, that&#x27;s the same bug: this check ignores subnormals.&lt;&#x2F;p&gt;
&lt;p&gt;So it&#x27;s not just Rust std that&#x27;s buggy, &lt;strong&gt;musl libc is buggy too&lt;&#x2F;strong&gt;. And &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;git.musl-libc.org&#x2F;cgit&#x2F;musl&#x2F;tree&#x2F;src&#x2F;math&#x2F;fmaf.c?id=f21a96538f78fa8e2040831b4209b35f2fb581da&quot;&gt;they credit this implementation to FreeBSD&lt;&#x2F;a&gt;, with copyright years 2005-2011, so who knows where else this buggy code was copy-pasted over the past two decades.&lt;&#x2F;p&gt;
&lt;p&gt;Musl has no bug tracker, so let&#x27;s &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;www.openwall.com&#x2F;lists&#x2F;musl&#x2F;2026&#x2F;08&#x2F;10&#x2F;1&quot;&gt;report it&lt;&#x2F;a&gt; on the mailing list, and pray that it isn&#x27;t an intentional &quot;optimization&quot; (the mailing list has no search function so I can&#x27;t check) and that it won&#x27;t just be forgotten forever once it stops showing up on the recent messages page. Boy do I love development processes from 30 years ago!&lt;&#x2F;p&gt;
&lt;h2 id=&quot;where-is-the-bottom&quot;&gt;Where is the bottom?!&lt;a class=&quot;zola-anchor&quot; href=&quot;#where-is-the-bottom&quot; aria-label=&quot;Anchor link for: where-is-the-bottom&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;Oh hey, CI checks for my standard library PR are complete!&lt;&#x2F;p&gt;
&lt;p&gt;Wait, why are they failing, but only on 32-bit ARM? With completely garbage values?!&lt;&#x2F;p&gt;
&lt;p&gt;Somehow the experimental, nightly-only &lt;code&gt;f128&lt;&#x2F;code&gt; type gets enabled on 32-bit ARM platforms, and... goes completely haywire? I sure am not going to read that much ARM assembly, but an LLM can &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;compiler-builtins&#x2F;pull&#x2F;1270#issuecomment-5246545522&quot;&gt;make quick work of it&lt;&#x2F;a&gt;, and... it&#x27;s a rustc ABI bug. On this type, on this architecture, localized entirely to my pull request.&lt;&#x2F;p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Aside:&lt;&#x2F;strong&gt; In case you&#x27;re wondering, using LLMs for analysis is &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;blog.rust-lang.org&#x2F;inside-rust&#x2F;2026&#x2F;08&#x2F;05&#x2F;rust-langrust-is-adopting-an-llm-policy&#x2F;&quot;&gt;permitted&lt;&#x2F;a&gt; by the Rust LLM policy, under certain conditions.&lt;&#x2F;p&gt;
&lt;&#x2F;blockquote&gt;
&lt;p&gt;So there&#x27;s now &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;compiler-builtins&#x2F;issues&#x2F;1271&quot;&gt;a bug report about this too&lt;&#x2F;a&gt; from a maintainer who could verify the LLM findings and understands this area way better than I do.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;fixes&quot;&gt;Fixes&lt;a class=&quot;zola-anchor&quot; href=&quot;#fixes&quot; aria-label=&quot;Anchor link for: fixes&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;&lt;code&gt;fearless_simd&lt;&#x2F;code&gt; is easy. It&#x27;s fixed. You can use that and get FMA without this bug.&lt;&#x2F;p&gt;
&lt;p&gt;The nightly-only &lt;code&gt;f128&lt;&#x2F;code&gt; bug in the Rust compiler is also fixed.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;compiler-builtins&#x2F;pull&#x2F;1270&quot;&gt;My pull request for the Rust standard library&lt;&#x2F;a&gt; broke setting floating-point exception flags in some cases. It occurred to me to run an LLM to search for bugs I could have introduced, and it found a counter-example. So I&#x27;ve put some more work into making a patch for it that touched as little code as possible. It&#x27;s still awaiting review.&lt;&#x2F;p&gt;
&lt;p&gt;To my surprise, musl developers took my bug report seriously. Both a minimal fix and a substantial rewrite of the &lt;code&gt;fmaf()&lt;&#x2F;code&gt; function were quickly proposed and iterated upon based on feedback from other developers. And after all the feedback was addressed... nothing happened. It&#x27;s still not merged as I&#x27;m writing this.&lt;&#x2F;p&gt;
&lt;p&gt;Looks like code review is a bottleneck regardless of the age of your development tools.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;does-this-bug-matter&quot;&gt;Does this bug matter?&lt;a class=&quot;zola-anchor&quot; href=&quot;#does-this-bug-matter&quot; aria-label=&quot;Anchor link for: does-this-bug-matter&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;I have no idea.&lt;&#x2F;p&gt;
&lt;p&gt;On one hand, it&#x27;s just an incorrect rounding. The result is off by one least significant bit in some rare cases. The original counter-example is off by about 15 parts per million of the correct result.&lt;&#x2F;p&gt;
&lt;p&gt;On the other hand, the error might get amplified dramatically depending on what you do with the result. And those math routines with exact error bounds are used for &lt;em&gt;something&lt;&#x2F;em&gt;, so whatever that is will get the wrong error bounds.&lt;&#x2F;p&gt;
&lt;p&gt;Perhaps most noticeably, deterministic simulations running across different machines will diverge.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;buggy-fma-in-my-computer&quot;&gt;Buggy FMA? In MY computer?&lt;a class=&quot;zola-anchor&quot; href=&quot;#buggy-fma-in-my-computer&quot; aria-label=&quot;Anchor link for: buggy-fma-in-my-computer&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;On x86 this is only a problem on cheap and&#x2F;or very old Intel without AVX2, and on Chinese Hygon x86 chips that don&#x27;t have FMA at all. 64-bit ARM is unaffected.&lt;&#x2F;p&gt;
&lt;p&gt;However, this will probably haunt 32-bit ARM and various embedded systems for years to come. I would not be surprised if the buggy FreeBSD implementation got copy-pasted into lots of different toolchains for various obscure platforms.&lt;&#x2F;p&gt;
&lt;p&gt;If FMA accuracy matters to you, check that &lt;code&gt;fmaf(a, b, c)&lt;&#x2F;code&gt; with these input bit patterns&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;rust&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #FE640B;&quot;&gt; 0x97000800&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #FE640B;&quot;&gt; 0x1cfff001&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;c&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #FE640B;&quot;&gt; 0x00010002&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;evaluates to bit pattern &lt;code&gt;0x00010001&lt;&#x2F;code&gt; (correct) rather than &lt;code&gt;0x00010002&lt;&#x2F;code&gt; (buggy).&lt;&#x2F;p&gt;
</content>
	</entry>
	<entry xml:lang="en">
		<title>Improving std::simd::swizzle_dyn</title>
        <author>
            <name>Sergey &quot;Shnatsel&quot; Davidoff</name>
        </author>
		<published>2026-06-20T00:00:00+00:00</published>
		<updated>2026-06-20T00:00:00+00:00</updated>
		<link href="https://shnatsel.github.io/improving-std-simd-swizzle-dyn/"/>
		<link rel="alternate" href="https://shnatsel.github.io/improving-std-simd-swizzle-dyn/" type="text/html"/>
		<id>https://shnatsel.github.io/improving-std-simd-swizzle-dyn/</id>
        <summary type="html">&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;linebender&#x2F;fearless_simd&quot;&gt;Fearless SIMD&lt;&#x2F;a&gt; recently received a &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;linebender&#x2F;fearless_simd&#x2F;pull&#x2F;265&quot;&gt;pull request&lt;&#x2F;a&gt; that started like this:&lt;&#x2F;p&gt;
&lt;blockquote&gt;
&lt;p&gt;Swizzles are a whole can of worm and I do not intend to figure out how we should implement this generically right now. 😄&lt;&#x2F;p&gt;
&lt;&#x2F;blockquote&gt;</summary>
		<content type="html">&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;linebender&#x2F;fearless_simd&quot;&gt;Fearless SIMD&lt;&#x2F;a&gt; recently received a &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;linebender&#x2F;fearless_simd&#x2F;pull&#x2F;265&quot;&gt;pull request&lt;&#x2F;a&gt; that started like this:&lt;&#x2F;p&gt;
&lt;blockquote&gt;
&lt;p&gt;Swizzles are a whole can of worm and I do not intend to figure out how we should implement this generically right now. 😄&lt;&#x2F;p&gt;
&lt;&#x2F;blockquote&gt;
&lt;span id=&quot;continue-reading&quot;&gt;&lt;&#x2F;span&gt;
&lt;p&gt;The pull request author only wanted to convert RGBA image layouts, e.g. RGBA &amp;lt;-&amp;gt; BGRA, so we solved the immediate need by implementing a &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;linebender&#x2F;fearless_simd&#x2F;pull&#x2F;266&quot;&gt;simpler version&lt;&#x2F;a&gt; that shuffles bytes within 128-bit blocks. That way we can trivially support this operation even on 512-bit vectors on 128-bit hardware, and 512-bit hardware still gets to utilize its full potential with native instructions.&lt;&#x2F;p&gt;
&lt;p&gt;But that made me wonder: how does &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;doc.rust-lang.org&#x2F;std&#x2F;simd&#x2F;index.html&quot;&gt;&lt;code&gt;std::simd&lt;&#x2F;code&gt;&lt;&#x2F;a&gt; implement arbitrary swizzles that aren&#x27;t confined to 128-bit blocks? Turns out the answer is: &lt;strong&gt;poorly.&lt;&#x2F;strong&gt;&lt;&#x2F;p&gt;
&lt;p&gt;Over the course of this article we&#x27;re going to fix some of that.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;what-is-a-swizzle&quot;&gt;What is a swizzle?&lt;a class=&quot;zola-anchor&quot; href=&quot;#what-is-a-swizzle&quot; aria-label=&quot;Anchor link for: what-is-a-swizzle&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;A swizzle, or a shuffle, rearranges elements within an array.&lt;&#x2F;p&gt;
&lt;p&gt;For example, if I have the input array &lt;code&gt;[A,D,F,I,M,R,S,T]&lt;&#x2F;code&gt; and apply the swizzle mask &lt;code&gt;[2,0,6,7,6,3,4,1]&lt;&#x2F;code&gt;, I get &lt;code&gt;[F,A,S,T,S,I,M,D]&lt;&#x2F;code&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;The &lt;code&gt;dyn&lt;&#x2F;code&gt; part indicates that the swizzle mask is dynamic (not hardcoded), so you could supply the swizzle mask &lt;code&gt;[6,4,0,5,7,0,6,6]&lt;&#x2F;code&gt; at runtime to get a different word.&lt;&#x2F;p&gt;
&lt;p&gt;This example uses only 8 bytes, or 64 bits. In practice hardware implements shuffles on vector sizes from 128 to 512 bits, and so does &lt;code&gt;std::simd&lt;&#x2F;code&gt;.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;understanding-the-implementation&quot;&gt;Understanding the implementation&lt;a class=&quot;zola-anchor&quot; href=&quot;#understanding-the-implementation&quot; aria-label=&quot;Anchor link for: understanding-the-implementation&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;On a high level, &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;rust&#x2F;blob&#x2F;d527bc9bfa297ca7fd7f5ae93781eeec42073170&#x2F;library&#x2F;portable-simd&#x2F;crates&#x2F;core_simd&#x2F;src&#x2F;swizzle_dyn.rs&quot;&gt;the current implementation&lt;&#x2F;a&gt; of &lt;code&gt;std::simd::swizzle_dyn&lt;&#x2F;code&gt; is very simple: use the native hardware operation if available, otherwise give up and move bytes one by one.&lt;&#x2F;p&gt;
&lt;p&gt;Even when hardware shuffles for a given size are available, &lt;code&gt;swizzle_dyn&lt;&#x2F;code&gt; doesn&#x27;t always map to the hardware cleanly. It promises that the values for out-of-bounds indices will be set to &lt;code&gt;0&lt;&#x2F;code&gt;, while x86 hardware shuffle just lets them wrap, so &lt;code&gt;swizzle_dyn&lt;&#x2F;code&gt; has to do extra work to uphold this guarantee.&lt;&#x2F;p&gt;
&lt;p&gt;Wait a minute... what is this warning in its documentation?&lt;&#x2F;p&gt;
&lt;blockquote&gt;
&lt;p&gt;Note that the current implementation is selected during build-time of the standard library, so &lt;code&gt;cargo build -Zbuild-std&lt;&#x2F;code&gt; may be necessary to unlock better performance, especially for larger vectors. A planned compiler improvement will enable using &lt;code&gt;#[target_feature]&lt;&#x2F;code&gt; instead.&lt;&#x2F;p&gt;
&lt;&#x2F;blockquote&gt;
&lt;p&gt;Uh oh.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-black-sheep-of-std-simd&quot;&gt;The black sheep of std::simd&lt;a class=&quot;zola-anchor&quot; href=&quot;#the-black-sheep-of-std-simd&quot; aria-label=&quot;Anchor link for: the-black-sheep-of-std-simd&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;Rust uses &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;LLVM&quot;&gt;LLVM&lt;&#x2F;a&gt; for optimizing your code and turning into something CPUs can execute. And LLVM doesn&#x27;t like dealing with platform-specific operations. It&#x27;s much easier to implement an operation like &quot;add two SIMD vectors&quot; generically, run all optimizations on this generic form, and then select the appropriate instructions for a given CPU later.&lt;&#x2F;p&gt;
&lt;p&gt;Most &lt;code&gt;std::simd&lt;&#x2F;code&gt; operations simply emit the generic form such as &quot;add two SIMD vectors&quot;, which makes the implementation remarkably straightforward.&lt;&#x2F;p&gt;
&lt;p&gt;But &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;doc.rust-lang.org&#x2F;std&#x2F;simd&#x2F;struct.Simd.html#method.swizzle_dyn&quot;&gt;swizzle_dyn&lt;&#x2F;a&gt; does not. And &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;rust&#x2F;blob&#x2F;d527bc9bfa297ca7fd7f5ae93781eeec42073170&#x2F;library&#x2F;portable-simd&#x2F;crates&#x2F;core_simd&#x2F;src&#x2F;swizzle_dyn.rs&quot;&gt;the implementation&lt;&#x2F;a&gt; looks something like this:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;rust&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;#&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;[&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;cfg&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;target_feature &lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;=&lt;&#x2F;span&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt; &amp;quot;neon&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;))]&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;unsafe fn&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; armv7_neon_swizzle_u8x16&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt;bytes&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Simd&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;lt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;u8&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; 16&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; idxs&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Simd&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;lt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;u8&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; 16&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; -&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Simd&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;lt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;u8&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; 16&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;    use&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;&quot;&gt; core&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;&quot;&gt;arch&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;&quot;&gt;arm&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;{&lt;&#x2F;span&gt;&lt;span&gt;uint8x8x2_t&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; vcombine_&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;u8&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; vget_high_&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;u8&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; vget_low_&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;u8&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; vtbl2_&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;u8&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;};&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;    unsafe&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;        let&lt;&#x2F;span&gt;&lt;span&gt; bytes&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; uint8x8x2_t&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;vget_low_u8&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;bytes&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;into&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;()),&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; vget_high_u8&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;bytes&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;into&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;()));&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;        let&lt;&#x2F;span&gt;&lt;span&gt; lo&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; vtbl2_u8&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;bytes&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; vget_low_u8&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;idxs&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;into&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;()));&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;        let&lt;&#x2F;span&gt;&lt;span&gt; hi&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; vtbl2_u8&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;bytes&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; vget_high_u8&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;idxs&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;into&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;()));&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;        vcombine_u8&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;lo&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; hi&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;into&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;()&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;    }&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;}&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;I&#x27;ve seen that before! That&#x27;s no LLVM magic, even if it&#x27;s completely inscrutable. Those are plain old SIMD intrinsics!&lt;&#x2F;p&gt;
&lt;p&gt;The problem is those platform-specific intrinsics are under &lt;code&gt;#[cfg(target_feature = ...))]&lt;&#x2F;code&gt;. It&#x27;s that &lt;code&gt;cfg&lt;&#x2F;code&gt; that kills performance, because it gets resolved very early in the build process, before the context from which the function is called is considered. &lt;code&gt;fearless_simd&lt;&#x2F;code&gt; and every other SIMD crate that isn&#x27;t called &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;wide&quot;&gt;&lt;code&gt;wide&lt;&#x2F;code&gt;&lt;&#x2F;a&gt; has a &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;shnatsel.github.io&#x2F;safe-simd-in-rust-even-on-the-inside&#x2F;&quot;&gt;solution&lt;&#x2F;a&gt; for this, but &lt;code&gt;std::simd&lt;&#x2F;code&gt; doesn&#x27;t, and it can&#x27;t just adopt the ecosystem solution without a large change to the API.&lt;&#x2F;p&gt;
&lt;p&gt;The warning in documentation doesn&#x27;t convey the full extent of the problem. Not only is &lt;code&gt;-Z build-std&lt;&#x2F;code&gt; necessary, you also have to pass &lt;code&gt;RUSTFLAGS=-C target-cpu=&lt;&#x2F;code&gt; to get decent performance. But if you do that and you run the program on a CPU that&#x27;s older than the one you selected, the program will crash. The typical solution is selecting the best implementation at runtime via crates such as &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;multiversion&quot;&gt;&lt;code&gt;multiversion&lt;&#x2F;code&gt;&lt;&#x2F;a&gt;, but that does not work for &lt;code&gt;swizzle_dyn&lt;&#x2F;code&gt;, even with &lt;code&gt;-Z build-std&lt;&#x2F;code&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;And let me be clear: I&#x27;m not a compiler engineer. I&#x27;m not the one who&#x27;s going to build whatever compiler improvement is planned to fix this.&lt;&#x2F;p&gt;
&lt;p&gt;But I am a &lt;code&gt;fearless_simd&lt;&#x2F;code&gt; maintainer, and I already know more about SIMD intrinsics than ever wanted to. So there is plenty here that I can improve.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;no-neon&quot;&gt;No NEON?&lt;a class=&quot;zola-anchor&quot; href=&quot;#no-neon&quot; aria-label=&quot;Anchor link for: no-neon&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;The first thing that stuck out to me when reading the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;rust&#x2F;blob&#x2F;d527bc9bfa297ca7fd7f5ae93781eeec42073170&#x2F;library&#x2F;portable-simd&#x2F;crates&#x2F;core_simd&#x2F;src&#x2F;swizzle_dyn.rs&quot;&gt;swizzle_dyn code&lt;&#x2F;a&gt; is that it only implements 128-bit swizzles on ARM NEON, even though NEON does have hardware instructions for larger swizzles.&lt;&#x2F;p&gt;
&lt;p&gt;It turned out that &lt;code&gt;std::simd&lt;&#x2F;code&gt; isn&#x27;t actually developed as part of the Rust standard library - it has &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;portable-simd&#x2F;&quot;&gt;its own repository&lt;&#x2F;a&gt; where development happens, and the result is periodically synced into the standard library. And the missing NEON implementations were &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;portable-simd&#x2F;blob&#x2F;9acfa97d5ce89df77c172b3d6942cafe688c18da&#x2F;crates&#x2F;core_simd&#x2F;src&#x2F;swizzle_dyn.rs#L118-L183&quot;&gt;already implemented there&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;That wasn&#x27;t at all apparent when browsing the code. I nearly wasted a lot of time making pull requests for an outdated version of the code in the wrong repository. Adding headers to synced-in files that specify how to contribute to them would save other people in my position a lot of time.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;doubling-the-vector-width&quot;&gt;Doubling the vector width&lt;a class=&quot;zola-anchor&quot; href=&quot;#doubling-the-vector-width&quot; aria-label=&quot;Anchor link for: doubling-the-vector-width&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;Let&#x27;s say we want to execute &lt;code&gt;[A,D,F,I,M,R,S,T].swizzle_dyn([2,0,6,7,6,3,4,1])&lt;&#x2F;code&gt; but all we have is half-width vectors. &lt;code&gt;swizzle_dyn&lt;&#x2F;code&gt; setting out-of-bounds elements to zero is going to help us here.&lt;&#x2F;p&gt;
&lt;p&gt;To process the first half, &lt;code&gt;[2,0,6,7]&lt;&#x2F;code&gt;, we run &lt;code&gt;[A,D,F,I].swizzle_dyn([2,0,6,7])&lt;&#x2F;code&gt; and &lt;code&gt;[M,R,S,T].swizzle_dyn([-2,-4,2,3])&lt;&#x2F;code&gt;, which gives us &lt;code&gt;[F,A,0,0]&lt;&#x2F;code&gt; and &lt;code&gt;[0,0,S,T]&lt;&#x2F;code&gt;. Then we combine these two intermediate results with a cheap bitwise OR, et voilà! We get &lt;code&gt;[F,A,S,T]&lt;&#x2F;code&gt;! Now repeat the same for the second half, and we&#x27;re done!&lt;&#x2F;p&gt;
&lt;p&gt;It is rare that a portability guarantee actually helps optimize something instead of incurring additional work, but this time we&#x27;re in luck!&lt;&#x2F;p&gt;
&lt;p&gt;This takes four shuffles instead of one, but it&#x27;s still much faster than moving bytes around one at a time. I&#x27;ve measured a 4x performance improvement from implementing this on AVX2 for 512-bit vectors inside &lt;code&gt;fearless_simd&lt;&#x2F;code&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;This is now &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;portable-simd&#x2F;pull&#x2F;540&quot;&gt;merged&lt;&#x2F;a&gt; into &lt;code&gt;std::simd&lt;&#x2F;code&gt;, and will be available in the next release of &lt;code&gt;fearless_simd&lt;&#x2F;code&gt; too.&lt;&#x2F;p&gt;
&lt;p&gt;Using this approach for 4x the vector size would require 16 shuffles, so I don&#x27;t think it&#x27;s worth it, but it might still be worth benchmarking.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;optimizing-avx2&quot;&gt;Optimizing AVX2&lt;a class=&quot;zola-anchor&quot; href=&quot;#optimizing-avx2&quot; aria-label=&quot;Anchor link for: optimizing-avx2&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;AVX2 is weird. It has 256-bit vectors, but most operations process 128-bit halves independently. That&#x27;s fine for most operations, but problematic for a swizzle that may move values across those halves. So the AVX2 implementation needs two separate steps: shuffling within 128-bit blocks, and then using a trick similar to what we&#x27;ve just done.&lt;&#x2F;p&gt;
&lt;p&gt;Another complication is that &lt;code&gt;swizzle_dyn&lt;&#x2F;code&gt; promises that values for out-of-bounds indices will be set to &lt;code&gt;0&lt;&#x2F;code&gt;, which x86 hardware shuffles don&#x27;t do, so we have to implement that in software.&lt;&#x2F;p&gt;
&lt;p&gt;The &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;portable-simd&#x2F;blob&#x2F;9acfa97d5ce89df77c172b3d6942cafe688c18da&#x2F;crates&#x2F;core_simd&#x2F;src&#x2F;swizzle_dyn.rs#L206-L223&quot;&gt;current implementation&lt;&#x2F;a&gt; looks like this:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;rust&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;use&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;&quot;&gt; x86&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span&gt;_mm256_permute2x128_si256 &lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;as&lt;&#x2F;span&gt;&lt;span&gt; avx2_cross_shuffle&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;use&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;&quot;&gt; x86&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span&gt;_mm256_shuffle_epi8 &lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;as&lt;&#x2F;span&gt;&lt;span&gt; avx2_half_pshufb&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;let&lt;&#x2F;span&gt;&lt;span&gt; mid&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Simd&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;splat&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #FE640B;&quot;&gt;16&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;u8&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;);&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;let&lt;&#x2F;span&gt;&lt;span&gt; high&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span&gt; mid&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; +&lt;&#x2F;span&gt;&lt;span&gt; mid&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;&#x2F;&#x2F; This is ordering sensitive, and LLVM will order these how you put them.&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;&#x2F;&#x2F; Most AVX2 impls use ~5 &amp;quot;ports&amp;quot;, and only 1 or 2 are capable of permutes.&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;&#x2F;&#x2F; But the &amp;quot;compose&amp;quot; step will lower to ops that can also use at least 1 other port.&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;&#x2F;&#x2F; So this tries to break up permutes so composition flows through &amp;quot;open&amp;quot; ports.&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;&#x2F;&#x2F; Comparative benches should be done on multiple AVX2 CPUs before reordering this&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;let&lt;&#x2F;span&gt;&lt;span&gt; hihi&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; avx2_cross_shuffle&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;lt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #FE640B;&quot;&gt;0x11&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;bytes&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;into&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(),&lt;&#x2F;span&gt;&lt;span&gt; bytes&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;into&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;());&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;let&lt;&#x2F;span&gt;&lt;span&gt; hi_shuf&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Simd&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;from&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;avx2_half_pshufb&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    hihi&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;        &#x2F;&#x2F; duplicate the vector&amp;#39;s top half&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    idxs&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;into&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(),&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt; &#x2F;&#x2F; so that using only 4 bits of an index still picks bytes 16-31&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;));&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;&#x2F;&#x2F; A zero-fill during the compose step gives the &amp;quot;all-Neon-like&amp;quot; OOB-is-0 semantics&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;let&lt;&#x2F;span&gt;&lt;span&gt; compose&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span&gt; idxs&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;simd_lt&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;high&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;select&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;hi_shuf&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;&quot;&gt; Simd&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;splat&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #FE640B;&quot;&gt;0&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;));&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;let&lt;&#x2F;span&gt;&lt;span&gt; lolo&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; avx2_cross_shuffle&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;lt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #FE640B;&quot;&gt;0x00&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;bytes&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;into&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(),&lt;&#x2F;span&gt;&lt;span&gt; bytes&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;into&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;());&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;let&lt;&#x2F;span&gt;&lt;span&gt; lo_shuf&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Simd&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;from&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;avx2_half_pshufb&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;lolo&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; idxs&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;into&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;()));&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;&#x2F;&#x2F; Repeat, then pick indices &amp;lt; 16, overwriting indices 0-15 from previous compose step&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;let&lt;&#x2F;span&gt;&lt;span&gt; compose&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span&gt; idxs&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;simd_lt&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;mid&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;select&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;lo_shuf&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; compose&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;);&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;compose&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;That&#x27;s pretty clever. The core insight is this:&lt;&#x2F;p&gt;
&lt;p&gt;CPUs have a bunch of different hardware curcuits, and the circuit that handles shuffles is entirely separate from a circuit that handles arithmetic, so you can run them both at the same time. This is known as &quot;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Instruction-level_parallelism&quot;&gt;instruction-level parallelism&lt;&#x2F;a&gt;&quot;.&lt;&#x2F;p&gt;
&lt;p&gt;The above code is carefully tuned to spread the work across different circuits in such a way that it doesn&#x27;t pile too much work on a single circuit (aka &quot;port&quot;) and take advantage of this parallelism to speed things up.&lt;&#x2F;p&gt;
&lt;p&gt;You can measure the effectiveness of this by viewing the resulting assembly with &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;cargo-show-asm&quot;&gt;cargo-show-asm&lt;&#x2F;a&gt; and feeding the result to &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;llvm.org&#x2F;docs&#x2F;CommandGuide&#x2F;llvm-mca.html&quot;&gt;llvm-mca&lt;&#x2F;a&gt; (Machine Code Analyzer), which models the execution of the given assembly on a specific CPU model accounting for all of these effects.&lt;&#x2F;p&gt;
&lt;p&gt;According to llvm-mca, spreading operations across different kinds of ports is important on the early Intel CPUs with AVX2, namely Haswell and Broadwell. On the other end of the spectrum, recent Intel Tiger Lake and all AMD Zen generations are so stuffed with all kinds of circuits that they favor doing less work overall, and you don&#x27;t have to worry about overwhelming one circuit anymore. Intel Skylake is indifferent, dealing with all formulations in an equally mediocre fashion.&lt;&#x2F;p&gt;
&lt;p&gt;It&#x27;s easy to get distracted arguing over which platform should be prioritized and optimized for. On one hand, recent CPUs are much more common these days, so optimizing for them benefits more people. On the other, the older CPUs are underpowered by today&#x27;s standards and need all the help they can get.&lt;&#x2F;p&gt;
&lt;p&gt;We can sidestep that entirely by using a clever algorithm that eliminates the compare-select steps, which improves performance on recent CPUs without regressing it on antiquated ones:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;rust&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;use&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;&quot;&gt; x86&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span&gt;_mm256_permute2x128_si256 &lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;as&lt;&#x2F;span&gt;&lt;span&gt; avx2_cross_shuffle&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;use&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;&quot;&gt; x86&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span&gt;_mm256_shuffle_epi8 &lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;as&lt;&#x2F;span&gt;&lt;span&gt; avx2_half_pshufb&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;let&lt;&#x2F;span&gt;&lt;span&gt; lolo&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; avx2_cross_shuffle&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;lt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #FE640B;&quot;&gt;0x00&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;bytes&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;into&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(),&lt;&#x2F;span&gt;&lt;span&gt; bytes&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;into&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;());&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;let&lt;&#x2F;span&gt;&lt;span&gt; hihi&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; avx2_cross_shuffle&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;lt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #FE640B;&quot;&gt;0x11&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;bytes&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;into&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(),&lt;&#x2F;span&gt;&lt;span&gt; bytes&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;into&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;());&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;&#x2F;&#x2F; Adding 0x60 preserves the low nibble and bit 4 for valid&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;&#x2F;&#x2F; indices 0..=31. Larger indices get their high bit set, so&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;&#x2F;&#x2F; VPSHUFB supplies the required out-of-bounds zeroing.&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;let&lt;&#x2F;span&gt;&lt;span&gt; control&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;&quot;&gt; x86&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;_mm256_adds_epu8&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;idxs&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;into&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(),&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;&quot;&gt; x86&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;_mm256_set1_epi8&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #FE640B;&quot;&gt;0x60&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;));&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;&#x2F;&#x2F; Move index bit 4 into each byte&amp;#39;s sign bit for VPBLENDVB.&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;let&lt;&#x2F;span&gt;&lt;span&gt; select_high&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;&quot;&gt; x86&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;_mm256_slli_epi16&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;lt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #FE640B;&quot;&gt;3&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;control&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;);&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;let&lt;&#x2F;span&gt;&lt;span&gt; from_low&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; avx2_half_pshufb&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;lolo&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; control&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;);&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;let&lt;&#x2F;span&gt;&lt;span&gt; from_high&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; avx2_half_pshufb&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;hihi&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; control&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;);&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #DF8E1D;&quot;&gt;x86&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;_mm256_blendv_epi8&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;from_low&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; from_high&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; select_high&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;into&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;()&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;llvm-mca shows no change on Haswell and Broadwell, but Zen 1, Zen 2 and Tiger Lake show a 23% improvement, and a 7-20% improvement on Zen 3. Skylake is unchanged for the 256-bit shuffle, but the decomposition of a 512-bit shuffle into four 256-bit ones improves by 20%.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;em&gt;I&#x27;m glossing over the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;www.geeksforgeeks.org&#x2F;computer-organization-architecture&#x2F;computer-organization-and-architecture-pipelining-set-1-execution-stages-and-throughput&#x2F;&quot;&gt;throughput vs latency&lt;&#x2F;a&gt; distinction here, since this formulation improves both.&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;p&gt;This change is currently &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;portable-simd&#x2F;pull&#x2F;542&quot;&gt;under review&lt;&#x2F;a&gt; in &lt;code&gt;std::simd&lt;&#x2F;code&gt;, and &lt;code&gt;fearless_simd&lt;&#x2F;code&gt; will ship with this formulation from the start.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;was-it-worth-it&quot;&gt;Was it worth it?&lt;a class=&quot;zola-anchor&quot; href=&quot;#was-it-worth-it&quot; aria-label=&quot;Anchor link for: was-it-worth-it&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;Performance of a single SIMD operation in isolation does not necessarily translate to improvements in algorithms using it.&lt;&#x2F;p&gt;
&lt;p&gt;For example, AVX2 only has 16 SIMD registers to hold data the CPU can directly operate on. If your operation runs faster but uses a lot more registers, the entire algorithm it&#x27;s in may run out of registers and have to spill data onto the stack, which adds costly load&#x2F;store instructions, which negates the gains or even slows down the entire algorithm!&lt;&#x2F;p&gt;
&lt;p&gt;We need to measure performance of some plausible code that uses 512-bit swizzles. 512 bits is 64 bytes, so &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Base64&quot;&gt;base64&lt;&#x2F;a&gt; encoding&#x2F;decoding is a natural fit.&lt;&#x2F;p&gt;
&lt;p&gt;I&#x27;m not aware of any shuffle-based Rust implementations of base64 that I could simply reuse. But it&#x27;s easy enough to have an LLM conjure a &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;Shnatsel&#x2F;intrepid-base64&quot;&gt;prototype&lt;&#x2F;a&gt;, for educational purposes only. Here&#x27;s how it performs with different shuffle implementations:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;plain&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;scalar fallback:   1.6 GiB&#x2F;s encode, 0.9 GiB&#x2F;s decode&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;AVX2 double-width: 8.2 GiB&#x2F;s encode, 5.5 GiB&#x2F;s decode&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;AVX2 dw optimized: 9.4 GiB&#x2F;s encode, 5.9 GiB&#x2F;s decode&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;That&#x27;s a ~6x improvement just from the work described in this article!&lt;&#x2F;p&gt;
&lt;p&gt;Let&#x27;s see how it stacks up against &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;mcy&#x2F;vb64&#x2F;&quot;&gt;Sunny Young&#x27;s base64 using std::simd&lt;&#x2F;a&gt;, which uses a lot of clever tricks and you should totally go read &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;mcyoung.xyz&#x2F;2023&#x2F;11&#x2F;27&#x2F;simd-base64&#x2F;&quot;&gt;their article about it&lt;&#x2F;a&gt;:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;plain&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;AVX2 vb64: 5.9 GiB&#x2F;s encode, 5.7 GiB&#x2F;s decode&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Brute-forcing the work with emulated swizzles and landing in the same ballpark as an implementation that carefully avoids them using &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;mcyoung.xyz&#x2F;2023&#x2F;11&#x2F;27&#x2F;simd-base64&#x2F;#simd-hash-table&quot;&gt;a clever perfect hash&lt;&#x2F;a&gt; is not too shabby! And ours have the added benefit of swapping out the base64 alphabet at will, while &lt;code&gt;vb64&lt;&#x2F;code&gt;&#x27;s perfect hash is tied to the standard base64 alphabet.&lt;&#x2F;p&gt;
&lt;p&gt;A &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;base64&quot;&gt;production implementation without SIMD&lt;&#x2F;a&gt; runs at 2.7 GiB&#x2F;s encode and 2.3 GiB&#x2F;s decode, so SIMD is well worth it. Our implementation on an AVX-512 CPU runs at ~13 GiB&#x2F;s, while vb64 doesn&#x27;t benefit from AVX-512 nearly as much. A &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;1910.05109&quot;&gt;state-of-the art&lt;&#x2F;a&gt; AVX-512 implementation &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;Shnatsel&#x2F;intrepid-base64&#x2F;blob&#x2F;main&#x2F;src&#x2F;avx512.rs&quot;&gt;translated to safe Rust&lt;&#x2F;a&gt; runs at 35 GiB&#x2F;s encode and 20 GiB&#x2F;s decode, so our prototype is not optimal, but still lets us see improvements to individual operations benefit a larger algorithm.&lt;&#x2F;p&gt;
&lt;p&gt;I&#x27;ve run these measurements on my Zen 4 workstation on &lt;code&gt;fearless_simd&lt;&#x2F;code&gt;, because the contributing guide for &lt;code&gt;std::simd&lt;&#x2F;code&gt; doesn&#x27;t say how to benchmark your changes.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;is-fearless-simd-better&quot;&gt;Is fearless_simd better?&lt;a class=&quot;zola-anchor&quot; href=&quot;#is-fearless-simd-better&quot; aria-label=&quot;Anchor link for: is-fearless-simd-better&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;For this one specific operation, &lt;code&gt;swizzle_dyn&lt;&#x2F;code&gt;, yes. The implementations are now identical in both &lt;code&gt;std::simd&lt;&#x2F;code&gt; and &lt;code&gt;fearless_simd&lt;&#x2F;code&gt;, but the latter doesn&#x27;t require &lt;code&gt;-Z build-std&lt;&#x2F;code&gt; and can multiversion this function just fine.&lt;&#x2F;p&gt;
&lt;p&gt;But when &lt;code&gt;std::simd&lt;&#x2F;code&gt; is playing to its strengths, it is really hard to beat or even match.&lt;&#x2F;p&gt;
&lt;p&gt;For example, &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;doc.rust-lang.org&#x2F;stable&#x2F;std&#x2F;simd&#x2F;trait.Swizzle.html#method.swizzle&quot;&gt;&lt;code&gt;std::simd::swizzle&lt;&#x2F;code&gt;&lt;&#x2F;a&gt; that accepts indices known at compile time translates directly into an LLVM &quot;shuffle vector&quot; operation, and LLVM selects the best formulation for your CPU out of a long list of special cases.&lt;&#x2F;p&gt;
&lt;p&gt;Replicating this on stable Rust would require a lot of work, and either a lot of const evaluation or a procedural macro, both of which would negatively impact build time. To the best of my knowledge, nobody has even attempted this yet.&lt;&#x2F;p&gt;
&lt;p&gt;That said, LLVM is not perfect. For example, I&#x27;ve found that it destroys performance of certain AVX-512 shuffles by turning them into a lot of SSE operations instead of just keeping it as a native AVX-512 shuffle. I&#x27;ve &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;rust&#x2F;issues&#x2F;156891&quot;&gt;reported this as an issue in Rust&lt;&#x2F;a&gt;, then a Rust standard library maintainer &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;llvm&#x2F;llvm-project&#x2F;issues&#x2F;199445&quot;&gt;reported it to LLVM&lt;&#x2F;a&gt; and &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;llvm&#x2F;llvm-project&#x2F;pull&#x2F;199454&quot;&gt;attempted to fix it&lt;&#x2F;a&gt;, and that led to a &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;llvm&#x2F;llvm-project&#x2F;pull&#x2F;201618&quot;&gt;whole&lt;&#x2F;a&gt; &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;llvm&#x2F;llvm-project&#x2F;pull&#x2F;201781&quot;&gt;bunch&lt;&#x2F;a&gt; of &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;llvm&#x2F;llvm-project&#x2F;pull&#x2F;201831&quot;&gt;improvements&lt;&#x2F;a&gt; made by LLVM developers. This work will ship in LLVM 23.&lt;&#x2F;p&gt;
&lt;p&gt;I&#x27;ve already uncovered &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;rust&#x2F;issues&#x2F;156946&quot;&gt;another LLVM inefficiency&lt;&#x2F;a&gt;, but &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;llvm&#x2F;llvm-project&#x2F;pull&#x2F;199633&quot;&gt;the PR to fix it&lt;&#x2F;a&gt; just sends LLVM into an infinite loop. So here&#x27;s a fun debugging puzzle if you&#x27;re looking for one!&lt;&#x2F;p&gt;
&lt;p&gt;LLVM issues affect &lt;code&gt;std::simd&lt;&#x2F;code&gt;, all other SIMD crates, and even other languages like C and C++, so fixing them is always a big win.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;takeaways&quot;&gt;Takeaways&lt;a class=&quot;zola-anchor&quot; href=&quot;#takeaways&quot; aria-label=&quot;Anchor link for: takeaways&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;So here&#x27;s what I&#x27;ve learned:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;std::simd::swizzle_dyn&lt;&#x2F;code&gt; is even less useful than it appears, since multiversioning doesn&#x27;t work on it.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;code&gt;std::simd&lt;&#x2F;code&gt; on nightly is outdated compared to the actual development state.&lt;&#x2F;li&gt;
&lt;li&gt;The contributing documentation for &lt;code&gt;std::simd&lt;&#x2F;code&gt; is poor:
&lt;ul&gt;
&lt;li&gt;It&#x27;s easy to open PRs for an outdated version in the wrong repository.&lt;&#x2F;li&gt;
&lt;li&gt;There&#x27;s no documented way to disassemble or benchmark your changes.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;LLVM can mangle perfectly optimal SIMD code, even if you use intrinsics that are supposed to lower to a specific instruction. Not just in Rust, in C and C++ too!&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;And, well, &lt;code&gt;std::simd::swizzle_dyn&lt;&#x2F;code&gt; will be up to 6x faster once this work gets synced into the standard library. You&#x27;re welcome!&lt;&#x2F;p&gt;
&lt;h2 id=&quot;addenum&quot;&gt;Addenum&lt;a class=&quot;zola-anchor&quot; href=&quot;#addenum&quot; aria-label=&quot;Anchor link for: addenum&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;Since this article was originally written, I&#x27;ve also found &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;portable-simd&#x2F;pull&#x2F;545&quot;&gt;a better formulation&lt;&#x2F;a&gt; of &lt;code&gt;swizzle_dyn&lt;&#x2F;code&gt; for AVX-512 with VBMI, and &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;axnsan12&quot;&gt;Cristi Vîjdea&lt;&#x2F;a&gt; suggested &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;portable-simd&#x2F;pull&#x2F;542#issuecomment-5077221668&quot;&gt;a better formulation&lt;&#x2F;a&gt; for ssse3. And together with &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;dzaima&quot;&gt;@dzaima&lt;&#x2F;a&gt; we&#x27;ve found &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;portable-simd&#x2F;pull&#x2F;548&quot;&gt;an even better formulation&lt;&#x2F;a&gt; for AVX2. All of these are now merged into &lt;code&gt;std::simd&lt;&#x2F;code&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;There is also &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;discourse.llvm.org&#x2F;t&#x2F;rfc-ir-ability-to-shuffle-vectors-with-dynamic-mask&#x2F;91282&quot;&gt;an ongoing effort&lt;&#x2F;a&gt; to add a native LLVM operation for &lt;code&gt;swizzle_dyn&lt;&#x2F;code&gt;, the timing of which coincided with my work. This would alleviate the miltiversioning issue in &lt;code&gt;std::simd&lt;&#x2F;code&gt;, and allow natively targeting even more architectures.&lt;&#x2F;p&gt;
&lt;p&gt;There&#x27;s also a separate &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;rust&#x2F;pull&#x2F;158713&quot;&gt;proposal&lt;&#x2F;a&gt; to add a Rust compiler feature that delays target feature checks, which could allow the current intrinsics-based implementation to be multiversioned.&lt;&#x2F;p&gt;
&lt;p&gt;In the meantime &lt;code&gt;fearless_simd&lt;&#x2F;code&gt; picked up the swizzle function matching &lt;code&gt;std::simd&lt;&#x2F;code&gt; behavior, and a cheaper variant that doesn&#x27;t return zero out-of-bounds indices. The cheaper variant produces arbitrary (but memory-safe) values instead. That way you don&#x27;t have to pay for the zeroing if you don&#x27;t use it.&lt;&#x2F;p&gt;
</content>
	</entry>
	<entry xml:lang="en">
		<title>The unreasonable effectiveness of LLMs for auditing Rust code</title>
        <author>
            <name>Sergey &quot;Shnatsel&quot; Davidoff</name>
        </author>
		<published>2026-06-20T00:00:00+00:00</published>
		<updated>2026-06-20T00:00:00+00:00</updated>
		<link href="https://shnatsel.github.io/the-unreasonable-effectiveness-of-llms-for-auditing-rust-code/"/>
		<link rel="alternate" href="https://shnatsel.github.io/the-unreasonable-effectiveness-of-llms-for-auditing-rust-code/" type="text/html"/>
		<id>https://shnatsel.github.io/the-unreasonable-effectiveness-of-llms-for-auditing-rust-code/</id>
        <summary type="html">&lt;p&gt;As a lead of the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;rust-lang.org&#x2F;governance&#x2F;teams&#x2F;#team-wg-secure-code&quot;&gt;Rust Secure Code Working Group&lt;&#x2F;a&gt;, I got free access to GPT-5.5 via the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;openai.com&#x2F;form&#x2F;codex-for-oss&#x2F;&quot;&gt;Codex for Open Source&lt;&#x2F;a&gt;. Since then I’ve found and reported dozens of issues of varying severity in widely used Rust crates.&lt;&#x2F;p&gt;</summary>
		<content type="html">&lt;p&gt;As a lead of the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;rust-lang.org&#x2F;governance&#x2F;teams&#x2F;#team-wg-secure-code&quot;&gt;Rust Secure Code Working Group&lt;&#x2F;a&gt;, I got free access to GPT-5.5 via the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;openai.com&#x2F;form&#x2F;codex-for-oss&#x2F;&quot;&gt;Codex for Open Source&lt;&#x2F;a&gt;. Since then I’ve found and reported dozens of issues of varying severity in widely used Rust crates.&lt;&#x2F;p&gt;
&lt;span id=&quot;continue-reading&quot;&gt;&lt;&#x2F;span&gt;
&lt;p&gt;Separately, the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;rustfoundation.org&#x2F;security-initiative&#x2F;&quot;&gt;Rust Foundation security initiative&lt;&#x2F;a&gt; got access to &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;www.anthropic.com&#x2F;claude&#x2F;mythos&quot;&gt;Mythos&lt;&#x2F;a&gt; via &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;www.anthropic.com&#x2F;glasswing&quot;&gt;Project Glasswing&lt;&#x2F;a&gt;, and their report should also be coming soon. I’ve coordinated with them so that our audit targets would not overlap.&lt;&#x2F;p&gt;
&lt;p&gt;While I haven’t found any truly devastating vulnerabilities, I am very impressed with GPT-5.5 for auditing Rust source code, and I’ll absolutely be adding it to my toolkit alongside fuzzers.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;Note:&lt;&#x2F;strong&gt;&lt;&#x2F;em&gt; &lt;em&gt;All opinions expressed in this article are my own, not that of any organizations I am a part of.&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;h2 id=&quot;methodology&quot;&gt;Methodology&lt;a class=&quot;zola-anchor&quot; href=&quot;#methodology&quot; aria-label=&quot;Anchor link for: methodology&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;Since maintainers could already be dealing with a large amount of vulnerability reports, it is &lt;em&gt;imperative&lt;&#x2F;em&gt; that I do not submit any invalid vulnerability reports and waste maintainers’ already limited time.&lt;&#x2F;p&gt;
&lt;p&gt;So I’ve decided to look for unambiguously problematic class of vulnerability that’s easy to verify: memory safety bugs.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;wait-isn-t-rust-memory-safe&quot;&gt;Wait, isn’t Rust memory-safe?&lt;a class=&quot;zola-anchor&quot; href=&quot;#wait-isn-t-rust-memory-safe&quot; aria-label=&quot;Anchor link for: wait-isn-t-rust-memory-safe&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h3&gt;
&lt;p&gt;Yes, with an asterisk.&lt;&#x2F;p&gt;
&lt;p&gt;Most code you’d write in Rust is memory-safe, but at some point you have to talk to the operating system or a C library or implement things like intrusive data structures, all of which involves raw pointers.&lt;&#x2F;p&gt;
&lt;p&gt;Most languages implement these parts in C (e.g. CPython), provide unsafe interoperability with C, and have you write C for your own unsafe code, while Rust has its own unsafe subset where you can muck about with raw pointers.&lt;&#x2F;p&gt;
&lt;p&gt;This puts Rust’s memory safety properties on par with Python’s, ahead of Go &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;www.ralfj.de&#x2F;blog&#x2F;2025&#x2F;07&#x2F;24&#x2F;memory-safety.html&quot;&gt;which violates safety on data races&lt;&#x2F;a&gt;, and behind browser-sandboxed JavaScript (but you can match that by compiling Rust to WebAssembly).&lt;&#x2F;p&gt;
&lt;p&gt;As for the amount of safe vs unsafe code in the wild, &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;rust&#x2F;comments&#x2F;g0wu9b&#x2F;percentage_of_unsafe_code_per_crate_for&#x2F;&quot;&gt;my own scan from 2020&lt;&#x2F;a&gt; showed that 95% of the code on &lt;a rel=&quot;nofollow external&quot; href=&quot;http:&#x2F;&#x2F;crates.io&quot;&gt;crates.io&lt;&#x2F;a&gt; is memory safe. The authors of the 2020 paper “&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;pm.inf.ethz.ch&#x2F;publications&#x2F;AstrauskasMathejaMuellerPoliSummers20.pdf&quot;&gt;How do programmers use unsafe Rust?&lt;&#x2F;a&gt;” independently arrived to the 95% number, although they didn’t put it into the final paper because they weren’t confident in their methodology for it. My own scan is also rather crude, but two completely different measurements arriving to the same number is encouraging.&lt;&#x2F;p&gt;
&lt;p&gt;In practice memory safety vulnerability rate reduction compared to C++ is &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;blog.google&#x2F;security&#x2F;rust-in-android-move-fast-fix-things&#x2F;&quot;&gt;about 1000x&lt;&#x2F;a&gt;, which is more than you’d expect based on the above figures.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;preventing-false-positives&quot;&gt;Preventing false positives&lt;a class=&quot;zola-anchor&quot; href=&quot;#preventing-false-positives&quot; aria-label=&quot;Anchor link for: preventing-false-positives&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h3&gt;
&lt;p&gt;Rust has a tool called &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;miri&quot;&gt;miri&lt;&#x2F;a&gt; that runs Rust code in an interpreter and tells you precisely whether it committed any crimes against the language rules or not. Safe Rust cannot violate them by construction, but unsafe Rust can, and a validator that immediate tells you whether you messed up or not instead of having to parse dozens of pages of dense prose is indispensable.&lt;&#x2F;p&gt;
&lt;p&gt;It also &lt;strong&gt;completely eliminates false positives&lt;&#x2F;strong&gt; from LLM vulnerability findings.&lt;&#x2F;p&gt;
&lt;p&gt;If the LLM can construct a unit test that causes miri to fail, I can report that to the maintainers and be certain that it’s a bug. I don’t ever have to argue if it’s a real issue or not, either — the proof is right there. And if miri says the execution is completely fine, then the LLM false positive gets discarded before anyone even sees it.&lt;&#x2F;p&gt;
&lt;p&gt;To the best of my knowledge, &lt;strong&gt;no other language&lt;&#x2F;strong&gt; has a practical tool with this level of precision. &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;clang.llvm.org&#x2F;docs&#x2F;AddressSanitizer.html&quot;&gt;Sanitizers&lt;&#x2F;a&gt; are very nice, but can’t catch everything, so verifying against them does not prove absence of issues.&lt;&#x2F;p&gt;
&lt;p&gt;Sadly miri is not without limitations — execution with extra checks is slow, calling into C is not supported, and syscall support is limited. When miri is not applicable, you can fall back on the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;doc.rust-lang.org&#x2F;beta&#x2F;unstable-book&#x2F;compiler-flags&#x2F;sanitizer.html&quot;&gt;sanitizers&lt;&#x2F;a&gt; and get &lt;em&gt;some&lt;&#x2F;em&gt; filtering.&lt;&#x2F;p&gt;
&lt;p&gt;I also had to switch miri to the newer &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;perso.crans.org&#x2F;vanille&#x2F;treebor&#x2F;&quot;&gt;Tree Borrows&lt;&#x2F;a&gt; aliasing model (as opposed to the older &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;plv.mpi-sws.org&#x2F;rustbelt&#x2F;stacked-borrows&#x2F;&quot;&gt;Stacked Borrows&lt;&#x2F;a&gt;) to avoid false positives, but fortunately that’s just one flag, -Zmiri-tree-borrows.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;harness&quot;&gt;Harness&lt;a class=&quot;zola-anchor&quot; href=&quot;#harness&quot; aria-label=&quot;Anchor link for: harness&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h3&gt;
&lt;p&gt;My setup was very basic: just Codex and a &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;gist.github.com&#x2F;Shnatsel&#x2F;e83219d7d6b73255373c2818ee438cda&quot;&gt;prompt&lt;&#x2F;a&gt;, with GPT-5.5 set to xhigh reasoning effort. It is important for the model to be able to write and run unit tests to try and trigger the issue under miri, so I consider something like Codex essential.&lt;&#x2F;p&gt;
&lt;p&gt;It would be interesting to try a more elaborate harness like &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;arm&#x2F;metis&quot;&gt;metis&lt;&#x2F;a&gt;, but even this basic setup was enough to discover interesting bugs.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;findings&quot;&gt;Findings&lt;a class=&quot;zola-anchor&quot; href=&quot;#findings&quot; aria-label=&quot;Anchor link for: findings&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;The most serious issue I’ve found is an &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;tirr-c&#x2F;jxl-oxide&#x2F;security&#x2F;advisories&#x2F;GHSA-5pmv-rx8r-wmv5&quot;&gt;out-of-bounds write in&lt;&#x2F;a&gt; &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;tirr-c&#x2F;jxl-oxide&#x2F;security&#x2F;advisories&#x2F;GHSA-5pmv-rx8r-wmv5&quot;&gt;jxl-grid&lt;&#x2F;a&gt; &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;tirr-c&#x2F;jxl-oxide&#x2F;security&#x2F;advisories&#x2F;GHSA-5pmv-rx8r-wmv5&quot;&gt;crate&lt;&#x2F;a&gt;. It is a part of &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;tirr-c&#x2F;jxl-oxide&quot;&gt;jxl-oxide&lt;&#x2F;a&gt;, a JPEG XL decoder in Rust (not to be confused with &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;libjxl&#x2F;jxl-rs&quot;&gt;jxl-rs&lt;&#x2F;a&gt;, which Firefox and Chromium are adopting for JPEG XL decoding; that one came up clean in my audit).&lt;&#x2F;p&gt;
&lt;p&gt;I’ve already fuzzed this crate earlier, but the fuzzer didn’t catch this issue because it only happens on 32-bit platforms and requires very large image dimensions. I wasn’t running the fuzzer on 32-bit, and the fuzzer is limited to small image dimensions to avoid exhausting my computer’s RAM, so it never had a chance to trigger this condition at runtime.&lt;&#x2F;p&gt;
&lt;p&gt;The initial demonstrator showed this as a jxl-grid issue, but it was not clear if it’s a theoretical problem or if it can be triggered by decoding a crafted image. GPT-5.5 &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;gist.github.com&#x2F;Shnatsel&#x2F;2c4e4f75e5892988d1315aa7ede4e575&quot;&gt;helped analyze that too&lt;&#x2F;a&gt;, and it turned out to be reachable. This was very valuable information to correctly prioritize the bug.&lt;&#x2F;p&gt;
&lt;p&gt;This still isn’t that big a deal in practice because old 32-bit devices and very recent and computationally intensive image formats rarely meet, but it does showcase the capability of the tool quite well.&lt;&#x2F;p&gt;
&lt;p&gt;Here are some more samples of bugs GPT-5.5 has discovered, showing the breadth of the kinds issues it has found:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;Amanieu&#x2F;intrusive-rs&#x2F;pull&#x2F;104&quot;&gt;Use-after-free&lt;&#x2F;a&gt;, &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;Amanieu&#x2F;intrusive-rs&#x2F;pull&#x2F;105&quot;&gt;data races&lt;&#x2F;a&gt; and &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;Amanieu&#x2F;intrusive-rs&#x2F;pull&#x2F;106&quot;&gt;panic safety issues&lt;&#x2F;a&gt; in intrusive-collections&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rkyv&#x2F;rkyv&#x2F;issues&#x2F;670&quot;&gt;Out-of-bounds reads&lt;&#x2F;a&gt; on deserializing crafted archives in rkyv&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;anza-xyz&#x2F;wincode&#x2F;issues&#x2F;306&quot;&gt;Exposing uninitialized memory in serialized data&lt;&#x2F;a&gt; in wincode&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;unicode-org&#x2F;icu4x&#x2F;pull&#x2F;8029&quot;&gt;Incorrect&lt;&#x2F;a&gt; &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;unicode-org&#x2F;icu4x&#x2F;pull&#x2F;8029&quot;&gt;Send&lt;&#x2F;a&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;unicode-org&#x2F;icu4x&#x2F;pull&#x2F;8029&quot;&gt;&#x2F;&lt;&#x2F;a&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;unicode-org&#x2F;icu4x&#x2F;pull&#x2F;8029&quot;&gt;Sync&lt;&#x2F;a&gt; &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;unicode-org&#x2F;icu4x&#x2F;pull&#x2F;8029&quot;&gt;impls&lt;&#x2F;a&gt; in yoke&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;unicode-org&#x2F;icu4x&#x2F;pull&#x2F;7940&quot;&gt;Construction of invalid enum values&lt;&#x2F;a&gt; in zerovec&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;djc&#x2F;hashlink&#x2F;issues&#x2F;42&quot;&gt;Data races for types with interior mutability&lt;&#x2F;a&gt; in hashlink&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;mozilla&#x2F;thin-vec&#x2F;issues&#x2F;86&quot;&gt;Soundness issue in Gecko FFI&lt;&#x2F;a&gt; in thin_vec (crossing languages!)&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;jaredforth&#x2F;webp&#x2F;pull&#x2F;51&quot;&gt;Out-of-bounds reads&lt;&#x2F;a&gt; in webp and &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;imazen&#x2F;webpx&#x2F;commit&#x2F;373015705ec84460ddc8722550805520478a2d57&quot;&gt;another one&lt;&#x2F;a&gt; inwebpx (C library wrappers)&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;zkat&#x2F;miette&#x2F;issues&#x2F;469&quot;&gt;Turning a&lt;&#x2F;a&gt; &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;zkat&#x2F;miette&#x2F;issues&#x2F;469&quot;&gt;&amp;amp;&lt;&#x2F;a&gt; &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;zkat&#x2F;miette&#x2F;issues&#x2F;469&quot;&gt;into a&lt;&#x2F;a&gt; &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;zkat&#x2F;miette&#x2F;issues&#x2F;469&quot;&gt;&amp;amp;mut&lt;&#x2F;a&gt; in miette&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;ruffle-rs&#x2F;nihav-vp6&#x2F;issues&#x2F;2&quot;&gt;Multiple&lt;&#x2F;a&gt; &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;ruffle-rs&#x2F;nihav-vp6&#x2F;issues&#x2F;2&quot;&gt;&amp;amp;mut&lt;&#x2F;a&gt; &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;ruffle-rs&#x2F;nihav-vp6&#x2F;issues&#x2F;2&quot;&gt;pointing to the same memory&lt;&#x2F;a&gt; in nihav-core&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;andylokandy&#x2F;arraydeque&#x2F;issues&#x2F;34&quot;&gt;Aliasing violation&lt;&#x2F;a&gt; in arraydeque&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;rust-av&#x2F;v_frame&#x2F;pull&#x2F;74&quot;&gt;Incorrect alignment handling&lt;&#x2F;a&gt; in v_frame&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;An interesting type confusion issue that’s not yet public&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;This demonstrates not just understanding of generic issues like out-of-bounds accesses, but also the ability to reason about Rust-specific concepts such as panic safety, aliasing, and the Send&#x2F;Sync traits that enforce thread safety.&lt;&#x2F;p&gt;
&lt;p&gt;I also got some results out of Claude before I got GPT-5.5 access:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;servo&#x2F;rust-smallvec&#x2F;pull&#x2F;407&quot;&gt;Use-after-free&lt;&#x2F;a&gt; in one function on a zero-capacity SmallVec (Opus 4.6)&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;rustsec.org&#x2F;packages&#x2F;imageproc.html&quot;&gt;Multiple out-of-bounds reads&lt;&#x2F;a&gt; in imageproc (Opus 4.7)&lt;&#x2F;p&gt;
&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;I haven’t used Claude enough to be able to compare the models. I also cannot compare GPT-5.5 to Mythos, as interesting as that would be, because I deliberately picked different targets to avoid duplicate vulnerability reports putting extra load on maintainers.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;fixes&quot;&gt;Fixes&lt;a class=&quot;zola-anchor&quot; href=&quot;#fixes&quot; aria-label=&quot;Anchor link for: fixes&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;Once the issue is identified and explained, the model usually can also fix it autonomously. I avoid “fix the bugs you found” style prompts and instead discuss possible solutions with the model first, then have the model implement one of them.&lt;&#x2F;p&gt;
&lt;p&gt;Submitting a possible fix alongside the vulnerability report puts less pressure on the maintainers. If you look up my pull requests, the first commit usually adds proof-of-concept tests that cause miri to complain, and the subsequent commit fixes the issues and turns the proof-of-concept snippets into regression tests.&lt;&#x2F;p&gt;
&lt;p&gt;GPT-5.5 has also assisted me in locating the version where the bug was introduced, which is essential for security advisories.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;reflections&quot;&gt;Reflections&lt;a class=&quot;zola-anchor&quot; href=&quot;#reflections&quot; aria-label=&quot;Anchor link for: reflections&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;In my experience Rust code is a lot easier to audit than C code. In C, if I look at a line like data[a + b], I have to trace through the entire codebase and find all the possible values it can be set to just to validate this one line and make sure it doesn&#x27;t have out-of-bounds accesses.&lt;&#x2F;p&gt;
&lt;p&gt;Rust, even unsafe Rust, still relies on local reasoning: if I see unsafe { data.get_unchecked(a + b) }, then the function it&#x27;s in must either validate a and b to make sure their addition is in-bounds, or be itself unsafe to call. Either way, there is a &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;kobzol.github.io&#x2F;rust&#x2F;2026&#x2F;06&#x2F;15&#x2F;how-memory-safety-cves-differ-between-rust-and-c-cpp.html&quot;&gt;clear point where verification must happen&lt;&#x2F;a&gt; - and if it&#x27;s not there, it&#x27;s a bug.&lt;&#x2F;p&gt;
&lt;p&gt;I don’t have to chase the data flow through the entire JPEG XL decoder by hand, and neither does an LLM. This reduces the complexity of auditing the code from combinatorial (all possible combinations of call trees) to linear (each function in isolation).&lt;&#x2F;p&gt;
&lt;p&gt;In that light, it’s not terribly surprising that LLMs are so good at auditing Rust code. And it’s also not terribly surprising that I haven’t found any devastating vulnerabilities after all.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;limitations&quot;&gt;Limitations&lt;a class=&quot;zola-anchor&quot; href=&quot;#limitations&quot; aria-label=&quot;Anchor link for: limitations&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;Just like fuzzers before them, LLMs surface numerous bugs that weren’t economical to discover previously. But at the end of the day, no heuristic tool can prove the absence of vulnerabilities.&lt;&#x2F;p&gt;
&lt;p&gt;For example, a GPT-5.5 alone didn’t discover &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;mariofeter&#x2F;secureloop-findings-public&#x2F;tree&#x2F;master&#x2F;findings&#x2F;rkyv&quot;&gt;several bugs&lt;&#x2F;a&gt; that a combination of a simpler LLM with a fuzzer did.&lt;&#x2F;p&gt;
&lt;p&gt;But we don’t have to rely on heuristics. Rust without unsafe does guarantee the absence of memory safety bugs.&lt;&#x2F;p&gt;
&lt;p&gt;So I’ll &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;shnatsel.medium.com&#x2F;how-to-avoid-bounds-checks-in-rust-without-unsafe-f65e618b4c1e&quot;&gt;keep&lt;&#x2F;a&gt; &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;shnatsel.medium.com&#x2F;safe-simd-in-rust-even-on-the-inside-c6f1ff381828&quot;&gt;shrinking&lt;&#x2F;a&gt; the unsafe surface where I can, and I&#x27;m glad to have these tools for when I can&#x27;t.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;em&gt;This article was originally published &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;shnatsel.medium.com&#x2F;the-unreasonable-effectiveness-of-llms-for-auditing-rust-code-d4df8bf0afd3&quot;&gt;on my Medium&lt;&#x2F;a&gt;&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
</content>
	</entry>
	<entry xml:lang="en">
		<title>Safe SIMD in Rust, even on the inside</title>
        <author>
            <name>Sergey &quot;Shnatsel&quot; Davidoff</name>
        </author>
		<published>2026-06-19T00:00:00+00:00</published>
		<updated>2026-06-19T00:00:00+00:00</updated>
		<link href="https://shnatsel.github.io/safe-simd-in-rust-even-on-the-inside/"/>
		<link rel="alternate" href="https://shnatsel.github.io/safe-simd-in-rust-even-on-the-inside/" type="text/html"/>
		<id>https://shnatsel.github.io/safe-simd-in-rust-even-on-the-inside/</id>
        <summary type="html">&lt;p&gt;&lt;em&gt;Rust’s&lt;&#x2F;em&gt; &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;shnatsel.medium.com&#x2F;the-state-of-simd-in-rust-in-2025-32c263e5f53d&quot;&gt;&lt;em&gt;SIMD abstractions&lt;&#x2F;em&gt;&lt;&#x2F;a&gt; &lt;em&gt;were not as safe as I’d like. Until now.&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;</summary>
		<content type="html">&lt;p&gt;&lt;em&gt;Rust’s&lt;&#x2F;em&gt; &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;shnatsel.medium.com&#x2F;the-state-of-simd-in-rust-in-2025-32c263e5f53d&quot;&gt;&lt;em&gt;SIMD abstractions&lt;&#x2F;em&gt;&lt;&#x2F;a&gt; &lt;em&gt;were not as safe as I’d like. Until now.&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;span id=&quot;continue-reading&quot;&gt;&lt;&#x2F;span&gt;
&lt;p&gt;It’s no secret that raw SIMD intrinsics are unpleasant to use.&lt;&#x2F;p&gt;
&lt;p&gt;You want to write &lt;code&gt;a + b&lt;&#x2F;code&gt;, not this monstrosity:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;rust&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;unsafe&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;    #&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;[&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;cfg&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;all&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;any&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;target_arch &lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;=&lt;&#x2F;span&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt; &amp;quot;x86&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; target_arch &lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;=&lt;&#x2F;span&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt; &amp;quot;x86_64&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;),&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; target_feature &lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;=&lt;&#x2F;span&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt; &amp;quot;avx2&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;))]&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;    _mm256_add_ps&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;    #&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;[&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;cfg&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;all&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;any&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;target_arch &lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;=&lt;&#x2F;span&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt; &amp;quot;x86&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; target_arch &lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;=&lt;&#x2F;span&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt; &amp;quot;x86_64&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;),&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; target_feature &lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;=&lt;&#x2F;span&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt; &amp;quot;sse&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; not&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;target_feature &lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;=&lt;&#x2F;span&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt; &amp;quot;avx2&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)))]&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;    _mm_add_ps&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;    #&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;[&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;cfg&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;all&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;target_arch &lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;=&lt;&#x2F;span&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt; &amp;quot;aarch64&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; target_feature &lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;=&lt;&#x2F;span&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt; &amp;quot;neon&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;))]&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;    vaddq_f32&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;}&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Look at it. It’s hideous. And the whole thing is wrapped in &lt;code&gt;unsafe&lt;&#x2F;code&gt;!&lt;&#x2F;p&gt;
&lt;p&gt;And that’s a simplified example. It still doesn’t handle:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;Other common platforms: AVX-512, 32-bit ARM, WebAssembly&lt;&#x2F;li&gt;
&lt;li&gt;Platforms without SIMD or obscure platforms like RISC-V&lt;&#x2F;li&gt;
&lt;li&gt;Actually loading data like &lt;code&gt;&amp;amp;[f32]&lt;&#x2F;code&gt; into a form that each intrinsic accepts&lt;&#x2F;li&gt;
&lt;li&gt;Selecting the best implementation for the CPU it’s running on&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;Luckily, Rust provides &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;shnatsel.medium.com&#x2F;the-state-of-simd-in-rust-in-2025-32c263e5f53d&quot;&gt;many SIMD abstractions&lt;&#x2F;a&gt; that handle all of that for you and let you simply write &lt;code&gt;a + b&lt;&#x2F;code&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;There is just one wrinkle. Inside, they’re still full of &lt;code&gt;unsafe&lt;&#x2F;code&gt;. It wasn’t gone, just hidden. Vast quantities of it lurking just beneath the surface, getting &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;servo&#x2F;pathfinder&#x2F;issues&#x2F;588&quot;&gt;screwed&lt;&#x2F;a&gt; &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;wingertge&#x2F;macerator&#x2F;issues&#x2F;31&quot;&gt;up&lt;&#x2F;a&gt; &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;sarah-quinones&#x2F;pulp&#x2F;issues&#x2F;28&quot;&gt;occasionally&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;Or rather, they were. Until now.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;why-do-we-even-need-unsafe&quot;&gt;Why do we even need ‘unsafe’?&lt;a class=&quot;zola-anchor&quot; href=&quot;#why-do-we-even-need-unsafe&quot; aria-label=&quot;Anchor link for: why-do-we-even-need-unsafe&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;For the longest time you couldn’t get around wrapping the call to each intrinsic function such as &lt;code&gt;_mm256_add_ps&lt;&#x2F;code&gt; into &lt;code&gt;unsafe&lt;&#x2F;code&gt; because it is illegal to call one when it’s not available on the CPU you’re running on.&lt;&#x2F;p&gt;
&lt;p&gt;So you &lt;em&gt;had&lt;&#x2F;em&gt; to have some mechanism for tracking which instructions are needed for each intrinsic, and which instructions you have access to, and cross-referencing them to decide if it’s safe to call a given function.&lt;&#x2F;p&gt;
&lt;p&gt;It was either tedious if done by hand or complex if done by a code generator, always error-prone, and required &lt;code&gt;unsafe&lt;&#x2F;code&gt; around every intrinsic.&lt;&#x2F;p&gt;
&lt;p&gt;This changed in Rust 1.87 when the compiler started tracking the required instruction sets itself, so you could write this:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;rust&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;#&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;[&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;target_feature&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;enable &lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;=&lt;&#x2F;span&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt; &amp;quot;avx2&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)]&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;fn&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; add_avx2&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt;a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; __m256&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; __m256&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; -&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; __m256&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;    _mm256_add_ps&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt; &#x2F;&#x2F; this is an avx2 intrinsic&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;}&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Look ma, no &lt;code&gt;unsafe&lt;&#x2F;code&gt;!&lt;&#x2F;p&gt;
&lt;p&gt;…yet.&lt;&#x2F;p&gt;
&lt;p&gt;You still cannot write &lt;code&gt;a + b&lt;&#x2F;code&gt; with this. The best you can do is this:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;rust&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;unsafe&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; add_avx2&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;) }&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;This only shifts the &lt;code&gt;unsafe&lt;&#x2F;code&gt; up a layer. You can call intrinsics inside functions annotated with the correct &lt;code&gt;#[target_feature]&lt;&#x2F;code&gt; now, but there still has to be &lt;code&gt;unsafe&lt;&#x2F;code&gt; somewhere in the chain.&lt;&#x2F;p&gt;
&lt;p&gt;The other problem is more fundamental. You cannot put &lt;code&gt;#[target_feature]&lt;&#x2F;code&gt; on the implementation of &lt;code&gt;+&lt;&#x2F;code&gt; for your type, because &lt;code&gt;+&lt;&#x2F;code&gt; must be available always. So no &lt;code&gt;a + b&lt;&#x2F;code&gt; for us using this mechanism.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;lemma-cpu-feature-tokens&quot;&gt;Lemma: CPU feature tokens&lt;a class=&quot;zola-anchor&quot; href=&quot;#lemma-cpu-feature-tokens&quot; aria-label=&quot;Anchor link for: lemma-cpu-feature-tokens&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;To understand how the final solution works, you first need to understand how CPU feature detection works.&lt;&#x2F;p&gt;
&lt;p&gt;Normally, checking for a CPU feature like AVX2 is done at runtime using &lt;code&gt;is_x86_feature_detected!(&quot;avx2&quot;)&lt;&#x2F;code&gt;. But we definitely don’t want to run this check every single time we add two numbers together — that would completely tank performance. We want to check it &lt;em&gt;once&lt;&#x2F;em&gt;, and then prove to the compiler that it’s safe to use AVX2 instructions from that point on.&lt;&#x2F;p&gt;
&lt;p&gt;Instead we can encode this proof into the type system using an &lt;em&gt;unforgeable token&lt;&#x2F;em&gt;: a zero-sized type with a private inner field. The &lt;em&gt;only&lt;&#x2F;em&gt; way to obtain this token is to call a function that performs the CPU feature check. If the check passes, the function hands you the token:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;rust&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;pub struct&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Avx2&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(());&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;fn&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; detect_avx2&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;()&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; -&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Option&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;lt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;Avx2&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;    if&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; is_x86_feature_detected!&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt;&amp;quot;avx2&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;) {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;        Some&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;Avx2&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(()))&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;    }&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt; else&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;        None&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;    }&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;}&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;And because it’s a zero-sized type, passing this token around has no runtime overhead. It exists purely as a compile-time proof.&lt;&#x2F;p&gt;
&lt;p&gt;The upshot is that as long as you have an instance of the &lt;code&gt;Avx2&lt;&#x2F;code&gt; struct, you can be sure that AVX2 instructions are available on the system.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-key-insight&quot;&gt;The key insight&lt;a class=&quot;zola-anchor&quot; href=&quot;#the-key-insight&quot; aria-label=&quot;Anchor link for: the-key-insight&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;The compiler doesn’t know it, but this function is safe to call:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;rust&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;#&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;[&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;target_feature&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;enable &lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;=&lt;&#x2F;span&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt; &amp;quot;avx2&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)]&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;fn&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; add_avx2&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt;token&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Avx2&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; __m256&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; __m256&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; -&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; __m256&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;    _mm256_add_ps&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;}&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;You can &lt;em&gt;only&lt;&#x2F;em&gt; call this function if you have an &lt;code&gt;Avx2&lt;&#x2F;code&gt; token, which you can only get if AVX2 instructions are available on the system.&lt;&#x2F;p&gt;
&lt;p&gt;If we can explain to the compiler that this is valid (using &lt;code&gt;unsafe&lt;&#x2F;code&gt;), we can write that &lt;code&gt;unsafe&lt;&#x2F;code&gt; only once and reuse it everywhere.&lt;&#x2F;p&gt;
&lt;p&gt;What we need is a macro which is safe to invoke:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;rust&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;with_avx2!&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;    fn&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; add_avx2&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt;token&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Avx2&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; __m256&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; __m256&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; -&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; __m256&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;        _mm256_add_ps&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;    }&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;but expands into this behind the scenes:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;rust&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;fn&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; add_avx2&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt;token&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Avx2&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; __m256&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; __m256&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; -&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; __m256&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;    &#x2F;&#x2F; SAFETY: Avx2 is available according to the token,&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;    &#x2F;&#x2F; and we verified that the inner function is not an `unsafe fn`&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;    unsafe&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; inner&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;token&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;) }&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;    #&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;[&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;target_feature&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;enable &lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;=&lt;&#x2F;span&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt; &amp;quot;avx2&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)]&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;    fn&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; inner&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt;token&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Avx2&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; __m256&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; __m256&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; -&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; __m256&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;        _mm256_add_ps&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;    }&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;}&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Now if you use an intrinsic that isn’t in AVX2, &lt;strong&gt;the compiler will reject it!&lt;&#x2F;strong&gt;&lt;&#x2F;p&gt;
&lt;p&gt;We’ve just managed to provide a safe programming interface to SIMD intrinsics without any bespoke tracking of target features!&lt;&#x2F;p&gt;
&lt;p&gt;Even though there is one &lt;code&gt;unsafe&lt;&#x2F;code&gt; block still inside it, it’s encapsulated in a sound API, so you cannot misuse it to cause memory safety bugs. In that sense it is just like &lt;code&gt;println!&lt;&#x2F;code&gt;, safely abstracting unsafe code.&lt;&#x2F;p&gt;
&lt;p&gt;This way you only ever need to review and audit &lt;em&gt;this one macro,&lt;&#x2F;em&gt; not hundreds upon hundreds of bespoke &lt;code&gt;unsafe&lt;&#x2F;code&gt; blocks. And the only things we could &lt;em&gt;possibly&lt;&#x2F;em&gt; screw up in the implementation are:&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;Mapping the token to the wrong &lt;code&gt;#[target_feature]&lt;&#x2F;code&gt;&lt;&#x2F;li&gt;
&lt;li&gt;Allowing an &lt;code&gt;unsafe fn&lt;&#x2F;code&gt; to be called from a safe context&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;p&gt;And both failure modes are quite easy to check for.&lt;&#x2F;p&gt;
&lt;p&gt;So now we can call &lt;code&gt;add_avx2(token, a, b)&lt;&#x2F;code&gt; without &lt;code&gt;unsafe&lt;&#x2F;code&gt;, but that still doesn’t get us to &lt;code&gt;a + b&lt;&#x2F;code&gt;. How do we solve &lt;em&gt;that?&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;h2 id=&quot;generics-to-the-rescue&quot;&gt;Generics to the rescue&lt;a class=&quot;zola-anchor&quot; href=&quot;#generics-to-the-rescue&quot; aria-label=&quot;Anchor link for: generics-to-the-rescue&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;We cannot annotate the implementation of &lt;code&gt;a + b&lt;&#x2F;code&gt; with &lt;code&gt;#[target_feature]&lt;&#x2F;code&gt; because it must be safe to call from anywhere. And we cannot pass a token into the function because it accepts &lt;code&gt;a&lt;&#x2F;code&gt; and &lt;code&gt;b&lt;&#x2F;code&gt; but not &lt;code&gt;token&lt;&#x2F;code&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;But even if we could do that, it would make for a pretty ugly API. We want &lt;code&gt;a + b&lt;&#x2F;code&gt; to always work and automatically use the best SIMD instructions without the user ever fussing with tokens.&lt;&#x2F;p&gt;
&lt;p&gt;We can use generics to solve both problems at once: by defining an &lt;code&gt;f32x8&lt;&#x2F;code&gt; type that’s generic over the available instruction sets, we can implement addition on it that both smuggles a token inside it &lt;em&gt;and&lt;&#x2F;em&gt; creates a separate implementation for each SIMD instruction set!&lt;&#x2F;p&gt;
&lt;p&gt;This is what it looks like:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;rust&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;pub trait&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Level&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {}&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;#&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;[&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;derive&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;Clone&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Copy&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)]&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;pub struct&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Avx2&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(());&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;impl&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Level&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt; for&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Avx2&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {}&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;pub struct&lt;&#x2F;span&gt;&lt;span&gt; f32x8&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;lt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;L&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Level&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;    &#x2F;&#x2F; For simplicity we&amp;#39;ll back this with an array in this example.&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;    &#x2F;&#x2F; In production code we use native SIMD types for the level.&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    data&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; [&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;f32&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #FE640B;&quot;&gt; 8&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;],&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;    &#x2F;&#x2F; The smuggled token!&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    token&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; L&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;}&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;&#x2F;&#x2F;&#x2F; implementation of `a + b` for Avx2&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;impl&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;&quot;&gt; std&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;&quot;&gt;ops&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;Add&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt; for&lt;&#x2F;span&gt;&lt;span&gt; f32x8&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;lt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;Avx2&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;    type&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Output&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #D20F39;&quot;&gt; Self&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;    fn&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; add&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #D20F39;&quot;&gt;self&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; rhs&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #D20F39;&quot;&gt; Self&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; -&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #D20F39;&quot;&gt; Self&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;Output&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;        &#x2F;&#x2F; (type conversions abbreviated)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;        &#x2F;&#x2F; Use the Avx2 token to call our safe wrapper&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;        let&lt;&#x2F;span&gt;&lt;span&gt; result&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; add_avx2&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #D20F39;&quot;&gt;self&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;&quot;&gt;token&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #D20F39;&quot;&gt; self&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; rhs&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;);&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #D20F39;&quot;&gt;        Self&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;            data&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; store_m256&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;result&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;),&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;            token&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #D20F39;&quot;&gt; self&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span&gt;token&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;        }&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;    }&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;}&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;And then we can just as easily make it work for any other instruction set, or when SIMD is not available at all:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;rust&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;#&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;[&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;derive&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;Clone&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Copy&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)]&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;pub struct&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; NoSimd&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(());&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;impl&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Level&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt; for&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; NoSimd&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {}&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;&#x2F;&#x2F;&#x2F; implementation of `a + b` when no SIMD is available&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;impl&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;&quot;&gt; std&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;&quot;&gt;ops&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;Add&lt;&#x2F;span&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt; for&lt;&#x2F;span&gt;&lt;span&gt; f32x8&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;lt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;NoSimd&lt;&#x2F;span&gt;&lt;span style=&quot;color: #04A5E5;&quot;&gt;&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;    type&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Output&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #D20F39;&quot;&gt; Self&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;    fn&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; add&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #D20F39;&quot;&gt;self&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; rhs&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #D20F39;&quot;&gt; Self&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; -&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #D20F39;&quot;&gt; Self&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;Output&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;        let&lt;&#x2F;span&gt;&lt;span&gt; result&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;&quot;&gt; std&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;&quot;&gt;array&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;from_fn&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;|&lt;&#x2F;span&gt;&lt;span&gt;i&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;|&lt;&#x2F;span&gt;&lt;span style=&quot;color: #D20F39;&quot;&gt; self&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;&quot;&gt;data&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;[&lt;&#x2F;span&gt;&lt;span&gt;i&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;]&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; +&lt;&#x2F;span&gt;&lt;span&gt; rhs&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;&quot;&gt;data&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;[&lt;&#x2F;span&gt;&lt;span&gt;i&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;]);&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #D20F39;&quot;&gt;        Self&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;            data&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span&gt; result&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;            token&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #D20F39;&quot;&gt; self&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span&gt;token&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;        }&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;    }&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;}&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;We’ve just solved safety &lt;em&gt;and&lt;&#x2F;em&gt; runtime instruction selection at once!&lt;&#x2F;p&gt;
&lt;p&gt;Add a convenience function that gives you the best &lt;code&gt;Level&lt;&#x2F;code&gt; available on the system, and you get pretty much the perfect API for SIMD!&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-abi-would-like-a-word&quot;&gt;The ABI would like a word&lt;a class=&quot;zola-anchor&quot; href=&quot;#the-abi-would-like-a-word&quot; aria-label=&quot;Anchor link for: the-abi-would-like-a-word&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;There is, unfortunately, a fundamental problem to writing &lt;code&gt;a + b&lt;&#x2F;code&gt; and having it lower to SIMD instructions: &lt;strong&gt;function call overhead.&lt;&#x2F;strong&gt;&lt;&#x2F;p&gt;
&lt;p&gt;Calling a function is not free but pretty cheap — just a handful of CPU instructions. But a &lt;em&gt;handful&lt;&#x2F;em&gt; of instructions is a lot more than &lt;em&gt;one&lt;&#x2F;em&gt; instruction we’ve just used for implementing addition!&lt;&#x2F;p&gt;
&lt;p&gt;So if there’s a function call in the way, addition performance will plummet. And performance is the entire point of using SIMD in the first place!&lt;&#x2F;p&gt;
&lt;p&gt;The compiler is usually pretty good at erasing this overhead via &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;matklad.github.io&#x2F;2021&#x2F;07&#x2F;09&#x2F;inline-in-rust.html&quot;&gt;inlining&lt;&#x2F;a&gt;. It basically copy-pastes the implementation of a function you called into the function calling it, so there’s no more function and no more overhead.&lt;&#x2F;p&gt;
&lt;p&gt;But &lt;code&gt;#[target_feature]&lt;&#x2F;code&gt; annotations throw a wrench in the works. The compiler cannot inline a function that has a &lt;code&gt;#[target_feature]&lt;&#x2F;code&gt; annotation into one that doesn’t, because the required features are not available in it!&lt;&#x2F;p&gt;
&lt;p&gt;And guess what cannot have a &lt;code&gt;#[target_feature]&lt;&#x2F;code&gt; annotation? Yeah.&lt;&#x2F;p&gt;
&lt;p&gt;So how &lt;em&gt;do&lt;&#x2F;em&gt; we make &lt;code&gt;a + b&lt;&#x2F;code&gt; work with SIMD?&lt;&#x2F;p&gt;
&lt;h2 id=&quot;inline-it-inline-it-with-fire&quot;&gt;Inline it, inline it with fire!&lt;a class=&quot;zola-anchor&quot; href=&quot;#inline-it-inline-it-with-fire&quot; aria-label=&quot;Anchor link for: inline-it-inline-it-with-fire&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;We cannot put &lt;code&gt;#[target_feature]&lt;&#x2F;code&gt; on the function that implements &lt;code&gt;a + b&lt;&#x2F;code&gt;, but we &lt;em&gt;can&lt;&#x2F;em&gt; put it on the function that calls &lt;code&gt;a + b&lt;&#x2F;code&gt;!&lt;&#x2F;p&gt;
&lt;p&gt;Then we can use &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;matklad.github.io&#x2F;2021&#x2F;07&#x2F;09&#x2F;inline-in-rust.html&quot;&gt;inlining&lt;&#x2F;a&gt; to get the implementation of &lt;code&gt;a + b&lt;&#x2F;code&gt; copied into the function calling it, and it ends up in a &lt;code&gt;#[target_feature]&lt;&#x2F;code&gt; context.&lt;&#x2F;p&gt;
&lt;p&gt;So the call chain looks like this:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #4C4F69; background-color: #EFF1F5;&quot;&gt;&lt;code data-lang=&quot;rust&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;#&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;[&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;target_feature&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;enable &lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;=&lt;&#x2F;span&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt; &amp;quot;avx2&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)]&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;fn&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; do_stuff&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;() {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;    &#x2F;&#x2F; TODO: some computation&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    c&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; =&lt;&#x2F;span&gt;&lt;span&gt; a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; +&lt;&#x2F;span&gt;&lt;span&gt; b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;    &#x2F;&#x2F; TODO: some more computation&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;}&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;&#x2F;&#x2F; which calls into...&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;#&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;[&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;inline&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;always&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)]&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt; &#x2F;&#x2F; Function body will be copied into the caller&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;fn&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; add&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #D20F39;&quot;&gt;self&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; rhs&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #D20F39;&quot;&gt; Self&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; -&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #D20F39;&quot;&gt; Self&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;::&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;Output&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;    &#x2F;&#x2F; Use the Avx2 token to call our safe wrapper&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;    add_avx2&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #D20F39;&quot;&gt;self&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;.&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;&quot;&gt;token&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #D20F39;&quot;&gt; self&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; rhs&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;);&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;    &#x2F;&#x2F; return statement abbreviated&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;}&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt;&#x2F;&#x2F; which calls into...&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;#&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;[&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;inline&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;]&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;font-style: italic;&quot;&gt; &#x2F;&#x2F; Function body will be copied into the caller if feasible&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;#&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;[&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;target_feature&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt;enable &lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;=&lt;&#x2F;span&gt;&lt;span style=&quot;color: #40A02B;&quot;&gt; &amp;quot;avx2&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)]&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #8839EF;&quot;&gt;fn&lt;&#x2F;span&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt; add_avx2&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt;token&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #DF8E1D;font-style: italic;&quot;&gt; Avx2&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; __m256&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt;:&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; __m256&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: #179299;&quot;&gt; -&amp;gt;&lt;&#x2F;span&gt;&lt;span style=&quot;color: #E64553;&quot;&gt; __m256&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt; {&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #1E66F5;font-style: italic;&quot;&gt;    _mm256_add_ps&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;(&lt;&#x2F;span&gt;&lt;span&gt;a&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;,&lt;&#x2F;span&gt;&lt;span&gt; b&lt;&#x2F;span&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #7C7F93;&quot;&gt;}&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;This works.&lt;&#x2F;p&gt;
&lt;p&gt;You can add these annotations in the right places and abstract over the SIMD levels and you &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;docs.rs&#x2F;fearless_simd&#x2F;latest&#x2F;fearless_simd&#x2F;trait.Simd.html#tymethod.vectorize&quot;&gt;don’t even need macros&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;The problem is that whenever you call &lt;code&gt;a + b&lt;&#x2F;code&gt; on SIMD types, you have to do it from a function with either &lt;code&gt;#[inline(always)]&lt;&#x2F;code&gt; or &lt;code&gt;#[target_feature]&lt;&#x2F;code&gt; on it, otherwise the code still compiles but performance plummets.&lt;&#x2F;p&gt;
&lt;p&gt;Wanna see for yourself? Open &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;rust.godbolt.org&#x2F;z&#x2F;xo1fEfqnT&quot;&gt;this example&lt;&#x2F;a&gt;, remove &lt;code&gt;#[target_feature]&lt;&#x2F;code&gt; from it and watch the generated assembly turn into absolute horror show.&lt;&#x2F;p&gt;
&lt;p&gt;I’m not sure what can be done about this. The limitation seems quite fundamental for any approach that implements &lt;code&gt;a + b&lt;&#x2F;code&gt; with SIMD.&lt;&#x2F;p&gt;
&lt;p&gt;The &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;shnatsel.medium.com&#x2F;github.com&#x2F;rust-lang&#x2F;rfcs&#x2F;pull&#x2F;3525&quot;&gt;Struct Target Features RFC&lt;&#x2F;a&gt; solves this for &lt;code&gt;add_avx2(token, a, b)&lt;&#x2F;code&gt; and &lt;code&gt;add::&amp;lt;S: Simd&amp;gt;(token, a, b)&lt;&#x2F;code&gt; but I don’t see a path to &lt;code&gt;a + b&lt;&#x2F;code&gt; just yet.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;where-does-that-leave-us&quot;&gt;Where does that leave us?&lt;a class=&quot;zola-anchor&quot; href=&quot;#where-does-that-leave-us&quot; aria-label=&quot;Anchor link for: where-does-that-leave-us&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;Despite the inlining wrinkle inherent to all SIMD code, we’ve managed to provide a remarkably pleasant SIMD abstraction at a staggeringly low, never-before-seen level of unsafe code.&lt;&#x2F;p&gt;
&lt;p&gt;You can find the production version of these ideas in &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;fearless_simd&quot;&gt;&lt;code&gt;fearless_simd&lt;&#x2F;code&gt;&lt;&#x2F;a&gt; v0.5, &lt;strong&gt;available now&lt;&#x2F;strong&gt; in a package registry near you! And &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;linebender&#x2F;fearless_simd&#x2F;blob&#x2F;d2411665f0726cd6c09bd6fd98af4af18e3c1778&#x2F;fearless_simd&#x2F;examples&#x2F;sigmoid.rs&quot;&gt;here’s a small example&lt;&#x2F;a&gt; to see how it all fits together in production.&lt;&#x2F;p&gt;
&lt;p&gt;The macro used to implement it is also exposed publicly, so you can easily &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;linebender&#x2F;fearless_simd&#x2F;blob&#x2F;e76b820154dc09f41bca8302670424d8ca079219&#x2F;fearless_simd&#x2F;examples&#x2F;srgb.rs&quot;&gt;mix and match&lt;&#x2F;a&gt; high-level operations like &lt;code&gt;a + b&lt;&#x2F;code&gt; and platform-specific intrinsics to take full advantage of the hardware.&lt;&#x2F;p&gt;
&lt;p&gt;There is more than one &lt;code&gt;unsafe&lt;&#x2F;code&gt; block in &lt;code&gt;fearless_simd&lt;&#x2F;code&gt; because it also provides the functionality of &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;safe_unaligned_simd&quot;&gt;&lt;code&gt;safe_unaligned_simd&lt;&#x2F;code&gt;&lt;&#x2F;a&gt; crate, but that too is done at a &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;okaneco&#x2F;safe_unaligned_simd&#x2F;issues&#x2F;51&quot;&gt;significantly lower&lt;&#x2F;a&gt; amount of unsafe code than the original.&lt;&#x2F;p&gt;
&lt;p&gt;For me the barrier to using high-level SIMD abstractions was always the sheer amount of &lt;code&gt;unsafe&lt;&#x2F;code&gt; they brought. It was scary and hard to justify.&lt;&#x2F;p&gt;
&lt;p&gt;But now SIMD in Rust can be truly fearless.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;acknowledgements&quot;&gt;Acknowledgements&lt;a class=&quot;zola-anchor&quot; href=&quot;#acknowledgements&quot; aria-label=&quot;Anchor link for: acknowledgements&quot;&gt;&lt;i class=&quot;fas fa-link&quot;&gt;&lt;&#x2F;i&gt;&lt;&#x2F;a&gt; 
&lt;&#x2F;h2&gt;
&lt;p&gt;I am shocked that I am the first to put this into production (as far as I can tell), because I’m certainly not the first person to think of this.&lt;&#x2F;p&gt;
&lt;p&gt;CPU feature tokens are an old and common idea. The &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;pulp&quot;&gt;&lt;code&gt;pulp&lt;&#x2F;code&gt;&lt;&#x2F;a&gt; crate has been using them for years, but they relied on handwritten unsafe wrappers around intrinsics, and &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;sarah-quinones&#x2F;pulp&#x2F;issues&#x2F;28&quot;&gt;occasionally got them wrong&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;Generating multiple implementations using generics is also an old idea. It is part of the &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;linebender.org&#x2F;blog&#x2F;towards-fearless-simd&#x2F;&quot;&gt;original &lt;code&gt;fearless_simd&lt;&#x2F;code&gt; concept&lt;&#x2F;a&gt; from 8 years ago. The &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;simdeez&quot;&gt;&lt;code&gt;simdeez&lt;&#x2F;code&gt;&lt;&#x2F;a&gt; crate, which predates it, also &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;docs.rs&#x2F;simdeez&#x2F;1.0.8&#x2F;simdeez&#x2F;&quot;&gt;seems to use something similar&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;The key insight of combining tokens with a single safe wrapper that delegates to &lt;code&gt;rustc&lt;&#x2F;code&gt; is not unique to me either. Just in the context of &lt;code&gt;fearless_simd&lt;&#x2F;code&gt; crate, Raph Levien has &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;linebender&#x2F;fearless_simd&#x2F;commit&#x2F;81b6ab4c8dd6f25064f539c49818c45c2f686815&quot;&gt;experimented with it&lt;&#x2F;a&gt;, and &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;DJMcNab&quot;&gt;Daniel McNab&lt;&#x2F;a&gt; created a &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;linebender&#x2F;fearless_simd&#x2F;pull&#x2F;108&quot;&gt;more elaborate implementation&lt;&#x2F;a&gt; than mine months ago.&lt;&#x2F;p&gt;
&lt;p&gt;Daniel’s approach allows fine-grained tracking of every individual CPU feature, as opposed to a handful of fixed CPU feature levels that &lt;code&gt;fearless_simd&lt;&#x2F;code&gt; uses. It is more expressive, but came at the cost of complexity, and his approach never got merged because no other maintainer stepped up to review it. I still hope it will be published as a standalone crate someday.&lt;&#x2F;p&gt;
&lt;p&gt;Thanks to Daniel and to &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;LaurenzV&quot;&gt;Laurenz Stampfl&lt;&#x2F;a&gt; for reviewing all my PRs to &lt;code&gt;fearless_simd&lt;&#x2F;code&gt;, they were big and the quick reviews are really appreciated!&lt;&#x2F;p&gt;
&lt;p&gt;&lt;em&gt;This article was originally published &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;shnatsel.medium.com&#x2F;safe-simd-in-rust-even-on-the-inside-c6f1ff381828&quot;&gt;on my Medium&lt;&#x2F;a&gt;.&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
</content>
	</entry>
	<entry xml:lang="en">
		<title>My old articles are on Medium</title>
        <author>
            <name>Sergey &quot;Shnatsel&quot; Davidoff</name>
        </author>
		<published>2026-06-18T00:00:00+00:00</published>
		<updated>2026-06-18T00:00:00+00:00</updated>
		<link href="https://shnatsel.github.io/my-old-articles-are-on-medium/"/>
		<link rel="alternate" href="https://shnatsel.github.io/my-old-articles-are-on-medium/" type="text/html"/>
		<id>https://shnatsel.github.io/my-old-articles-are-on-medium/</id>
		<content type="html">&lt;p&gt;Find them at &lt;a rel=&quot;nofollow external&quot; href=&quot;https:&#x2F;&#x2F;shnatsel.medium.com&#x2F;&quot;&gt;https:&#x2F;&#x2F;shnatsel.medium.com&#x2F;&lt;&#x2F;a&gt;&lt;&#x2F;p&gt;
</content>
	</entry>
</feed>
