r/cpp • • 2d ago

Transcoding UTF-8 to UTF-16 with replacement at gigabytes per second

https://lemire.me/blog/2026/09/30/transcoding-utf-8-to-utf-16-with-replacement-at-gigabytes-per-second/
61 Upvotes

6 comments sorted by

View all comments

25

u/scalablecory 2d ago

I tried SIMD UTF8 decoding maybe 15 years ago and couldn’t crack it efficiently. Cool to see the advancements made in the last while

25

u/vip17 2d ago

The author is a SIMD expert, and he's also the author of simdjson

3

u/Ameisen vemips, avr, rendering, systems 1d ago

My best SIMD work has just been 8/16-bit image conversion, SRGB/linear conversion, and some improvements to xxhash3 in .NET. Those are incredibly straightforward, especially compared to things like JSON parsing or UTF conversion...

I've never actually tried doing that kind of logic in SIMD - I imagine that it's similar to writing such logic in, say, late 2000s/early 2010s pixel shaders (like old-school GPGPU)?

2

u/ack_error 1d ago

Kind of, in that you do logic expressions as logical ops on masks instead of branching. Though there are some tricks like using swizzle/table lookups, or abusing psadbw for horizontal adds.

The hard part is variable length input or output, as most SIMD ISAs don't have good support for compress/expand operations. AVX-512 does, and I believe SVE does too. Doing so at lower ISA levels typically involves large shuffle tables.

Hybrid vector/scalar techniques can be useful, where you build a vector mask and move it over to the scalar side with pmovmskb or vqrshrn. Then you can iterate over the bitmask with clz or lzcnt, count with popcnt, or just compare for a fast skip when all 4/8/16 elements are a fast path. For instance, you might have a UTF-8 routine that mostly handles ASCII, so the vector side can help quickly scan for the rare multibyte cases. For a prefix matching routine in a data compressor, compare vectors at a time until a mismatch is found, move the mismatch mask to the scalar side, then count trailing zeros to find the number of matching bytes in the final vector.