Skip to content

[Vulkan] Stop clamping per-row reductions and propagate NaN - #23244

Draft
mergennachin wants to merge 1 commit into
gh/mergennachin/27/headfrom
gh/mergennachin/28/head
Draft

mergennachin wants to merge 1 commit into
gh/mergennachin/27/headfrom
gh/mergennachin/28/head

Conversation

@mergennachin

@mergennachin mergennachin commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

convert.glslh guarded an fp16 clamp with #if T == float16_t, but neither side is a macro, so the preprocessor compares 0 with 0 and the clamp was always on. Per-row buffer sum, mean, amax and amin therefore saturated fp32 and int results at 65504. The clamp is removed. Single-dimension texture and per-row buffer amax and amin now propagate NaN, argmax and argmin return the index of the first NaN as ATen does, and mean divides in the accumulator type. FP16 texture output uses explicit nearest-even conversion so rounding and overflow agree with ATen across drivers.

Part 5/15 of the Vulkan transformer and operator-conformance stack. Depends on #23243; review against the selected base branch. Integration PR: #23254.

Validation: 3 passed, 2 warnings in 57.33s. Lintrunner and git diff --check pass. Native tests use MoltenVK with portable CPU kernels; hardware and SwiftShader CI are pending.

Recreates #23205 through ghstack. Prior review discussion remains on that PR.

Authored with OpenAI Codex; split planned with Claude Code.

@mergennachin mergennachin added the module: vulkan Issues related to the Vulkan delegate and code under backends/vulkan/ label Sep 29, 2026
@pytorch-bot

pytorch-bot Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23244

Note: Links to docs will display an error until the docs builds have been completed.

❌ 21 New Failures, 1 Pending, 1 Unclassified Failure

As of commit fb905cf with merge base a318382 (image):

NEW FAILURES - The following jobs have failed:

UNCLASSIFIED FAILURE - DrCI could not classify the following job because the workflow did not run on the merge base. The failure may be pre-existing on trunk or introduced by this PR:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 29, 2026
mergennachin added a commit that referenced this pull request Sep 29, 2026
…ormer blocks

The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers.

Fixes #23156.

This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10.

| Part | PR | Change |
| --- | --- | --- |
| 1 | #23240 | Hardware Vulkan CI |
| 2 | #23241 | Scalar cache type and signed-zero keys |
| 3 | #23242 | Vulkan-local signed-zero serialization |
| 4 | #23243 | GELU modes and view kwargs |
| 5 | #23244 | Reduction clamp, NaN, and FP16 rounding |
| 6 | #23245 | Reduction dimension guards |
| 7 | #23246 | Bool staging and logical_not |
| 8 | #23247 | Scalar representability and symbolic guards |
| 9 | #23248 | 64-bit dtype and fusion policy |
| 10 | #23249 | scalar_tensor with exact integer values |
| 11 | #23250 | Typed, resizable full |
| 12 | #23251 | Power special values and logical FP16 dtype |
| 13 | #23252 | mul.Scalar |
| 14 | #23253 | any.dim |
| 15 | #23254 | Dynamic expand and transformer integration |

The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename.

Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons.

Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0.

The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed.

NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads.

Recreates #23162 through ghstack. Prior review discussion remains on that PR.

Authored with OpenAI Codex; split planned with Claude Code.


cc @SS-JIA @manuelcandales @digantdesai @cbilgin

ghstack-source-id: c6a7e94
ghstack-comment-id: 5892899249
Pull-Request: #23254
mergennachin added a commit that referenced this pull request Sep 29, 2026
…ormer blocks

The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers.

Fixes #23156.

This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10.

| Part | PR | Change |
| --- | --- | --- |
| 1 | #23240 | Hardware Vulkan CI |
| 2 | #23241 | Scalar cache type and signed-zero keys |
| 3 | #23242 | Vulkan-local signed-zero serialization |
| 4 | #23243 | GELU modes and view kwargs |
| 5 | #23244 | Reduction clamp, NaN, and FP16 rounding |
| 6 | #23245 | Reduction dimension guards |
| 7 | #23246 | Bool staging and logical_not |
| 8 | #23247 | Scalar representability and symbolic guards |
| 9 | #23248 | 64-bit dtype and fusion policy |
| 10 | #23249 | scalar_tensor with exact integer values |
| 11 | #23250 | Typed, resizable full |
| 12 | #23251 | Power special values and logical FP16 dtype |
| 13 | #23252 | mul.Scalar |
| 14 | #23253 | any.dim |
| 15 | #23254 | Dynamic expand and transformer integration |

The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename.

Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons.

Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0.

The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed.

NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads.

Recreates #23162 through ghstack. Prior review discussion remains on that PR.

Authored with OpenAI Codex; split planned with Claude Code.


cc @SS-JIA @manuelcandales @digantdesai @cbilgin

ghstack-source-id: f8d24d9
ghstack-comment-id: 5892899249
Pull-Request: #23254

This branch was successfully deployed

1 active deployment
cadence — fb905cfe Deployed Sep 29, 2026 by mergennachin via hifi-op-test / hifi4 #30409
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. module: vulkan Issues related to the Vulkan delegate and code under backends/vulkan/

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant