Skip to content

Add UTF-16 toWellFormed/isWellFormed (AVX-512, AVX2, SSE, NEON) - #54

Merged
lemire merged 2 commits into
mainfrom
utf16-to-well-formed
Sep 26, 2026
Merged

lemire merged 2 commits into
mainfrom
utf16-to-well-formed

Conversation

@lemire

@lemire lemire commented Sep 26, 2026

Copy link
Copy Markdown
Member

This adds SimdUnicode.UTF16, which makes UTF-16 strings well formed the way JavaScript's String.prototype.toWellFormed() does: every lone surrogate becomes U+FFFD. It also adds validation, like isWellFormed().

string s = UTF16.ToWellFormed(input);          // returns `input` itself (no allocation) when already well formed
bool ok = UTF16.IsWellFormed(span);
UTF16.ToWellFormed(source, destination);       // spans; in place when both are the same buffer
UTF16.ToWellFormed(char* input, int length, char* output);
char* p = UTF16.GetPointerToFirstInvalidChar(char* input, int length);

Kernels

Runtime dispatch picks AVX-512, AVX2, SSE4.1 or ARM64 NEON. Other systems fall back on the runtime's vectorized IndexOfAnyInRange. The kernels are based on simdutf's utf16fix, described in:

Robert Clausecker, Daniel Lemire, Fixing ill-formed UTF-16 strings with SIMD instructions, Software: Practice and Experience, 2026

They differ from simdutf in a few ways:

  • Carried high-surrogate bit. Each block's high-surrogate bitmask is shifted by one code unit, carrying one bit into the next block, instead of loading a second lookback block. On NEON, the shift is an ext in registers.
  • Fast skip. A min-reduction answers "any surrogate in this block?", and blocks without surrogates, which is most text, are skipped.
  • Aligned stores. Out-of-place output stores are aligned, because misaligned vector stores are costly and .NET arrays are only 8-byte aligned.
  • End of input and short strings. AVX-512 uses masked loads. The other kernels use overlapping windows, and inputs shorter than 64 code units first check whether they contain any surrogate at all.
  • Kernels compile on their own. Each kernel is marked [MethodImpl(NoInlining)] so it is compiled on its own. When a kernel was inlined into its caller, the JIT ran out of inlining budget and its small helpers became calls in the hot loop.

Tests

test/UTF16WellFormedTests.cs checks every kernel, in place and out of place, against a rune-based reference. It uses guard code units to detect reads or writes outside the buffers. The inputs are hard-coded edge cases, random strings of length 0 to 10007, a single lone surrogate at every position, and sparse surrogates in lengths 0 to 70 plus a few long ones.

On x64, you can force a lower kernel with DOTNET_EnableAVX512=0 or DOTNET_EnableAVX=0.

Benchmarks

Run dotnet run -c Release --filter "*UTF16WellFormed*" in benchmark/. The README has full tables for the Xeon Gold 6548N (AVX-512, and with AVX2 or SSE forced) and for the Apple M4 Max. Summary for well-formed text, in GB/s of UTF-16 input:

Kernel Validate vs IndexOfAnyInRange Emoji validate Buffer to buffer (plain copy) Emoji buffer
AVX-512 68–115 vs 31–39 53 vs 0.39 42–43 (44–46) 40 vs 0.39
AVX2 58–80 vs 41–50 35 vs 0.50 43 (45–47) 28 vs 0.49
SSE 50–64 vs 28 26 vs 0.56 43–45 (45–47) 22 vs 0.56
NEON (M4 Max) 102–135 vs 62–65 52 vs 2.0 52–80 (63–109) 52 vs 1.6
  • Short strings (1 to 64 code units, one call each): on par with IndexOfAnyInRange, and 6 to 8 times faster on emoji text.
  • Other approaches: string.Concat(s.EnumerateRunes()) runs at 0.2 to 0.4 GB/s.
  • Strings that need fixing: the cost of allocating the new string dominates, so both approaches end up at similar speeds, except on emoji text.

… kernels.

UTF16.ToWellFormed replaces lone surrogates by U+FFFD (JavaScript's
toWellFormed); UTF16.IsWellFormed and GetPointerToFirstInvalidChar
validate. The string overload returns well-formed input as is.

The kernels are based on simdutf's utf16fix (Clausecker and Lemire,
Fixing ill-formed UTF-16 strings with SIMD instructions, Software:
Practice and Experience, 2026). Rather than loading a lookback block,
they carry the high-surrogate bit across blocks, skip surrogate-free
blocks with a min-based pre-check, align output stores, and handle the
end of the input with masked loads (AVX-512) or overlapping windows.
Other systems fall back on IndexOfAnyInRange.

Tests compare every kernel, in place and out of place, against a
rune-based reference. Benchmarks compare against IndexOfAnyInRange,
EnumerateRunes and Encoding.Unicode.
…nRange fallback.

The fallback runs only when no SIMD kernel applies (no ARM64 NEON, AVX-512,
AVX2 or SSE4.1), so the tests now call it directly on every system.
@lemire
lemire merged commit d20b637 into main Sep 26, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant