Add UTF-16 toWellFormed/isWellFormed (AVX-512, AVX2, SSE, NEON) - #54
Merged
Merged
Conversation
… kernels. UTF16.ToWellFormed replaces lone surrogates by U+FFFD (JavaScript's toWellFormed); UTF16.IsWellFormed and GetPointerToFirstInvalidChar validate. The string overload returns well-formed input as is. The kernels are based on simdutf's utf16fix (Clausecker and Lemire, Fixing ill-formed UTF-16 strings with SIMD instructions, Software: Practice and Experience, 2026). Rather than loading a lookback block, they carry the high-surrogate bit across blocks, skip surrogate-free blocks with a min-based pre-check, align output stores, and handle the end of the input with masked loads (AVX-512) or overlapping windows. Other systems fall back on IndexOfAnyInRange. Tests compare every kernel, in place and out of place, against a rune-based reference. Benchmarks compare against IndexOfAnyInRange, EnumerateRunes and Encoding.Unicode.
…nRange fallback. The fallback runs only when no SIMD kernel applies (no ARM64 NEON, AVX-512, AVX2 or SSE4.1), so the tests now call it directly on every system.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This adds
SimdUnicode.UTF16, which makes UTF-16 strings well formed the way JavaScript'sString.prototype.toWellFormed()does: every lone surrogate becomes U+FFFD. It also adds validation, likeisWellFormed().Kernels
Runtime dispatch picks AVX-512, AVX2, SSE4.1 or ARM64 NEON. Other systems fall back on the runtime's vectorized
IndexOfAnyInRange. The kernels are based on simdutf'sutf16fix, described in:They differ from simdutf in a few ways:
extin registers.[MethodImpl(NoInlining)]so it is compiled on its own. When a kernel was inlined into its caller, the JIT ran out of inlining budget and its small helpers became calls in the hot loop.Tests
test/UTF16WellFormedTests.cschecks every kernel, in place and out of place, against a rune-based reference. It uses guard code units to detect reads or writes outside the buffers. The inputs are hard-coded edge cases, random strings of length 0 to 10007, a single lone surrogate at every position, and sparse surrogates in lengths 0 to 70 plus a few long ones.On x64, you can force a lower kernel with
DOTNET_EnableAVX512=0orDOTNET_EnableAVX=0.Benchmarks
Run
dotnet run -c Release --filter "*UTF16WellFormed*"inbenchmark/. The README has full tables for the Xeon Gold 6548N (AVX-512, and with AVX2 or SSE forced) and for the Apple M4 Max. Summary for well-formed text, in GB/s of UTF-16 input:IndexOfAnyInRangeIndexOfAnyInRange, and 6 to 8 times faster on emoji text.string.Concat(s.EnumerateRunes())runs at 0.2 to 0.4 GB/s.