[Vulkan] Keep reductions the reduce shaders cannot handle on CPU - #23245
Draft
mergennachin wants to merge 1 commit into
Draft
mergennachin wants to merge 1 commit into
mergennachin wants to merge 1 commit into
Conversation
…ghstack [ghstack-poisoned]
Contributor
Author
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23245
Note: Links to docs will display an error until the docs builds have been completed. ❌ 2 New FailuresAs of commit 2131a32 with merge base a318382 ( NEW FAILURES - The following jobs have failed:
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This was referenced Sep 29, 2026
mergennachin
added a commit
that referenced
this pull request
Sep 29, 2026
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers. Fixes #23156. This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 | Hardware Vulkan CI | | 2 | #23241 | Scalar cache type and signed-zero keys | | 3 | #23242 | Vulkan-local signed-zero serialization | | 4 | #23243 | GELU modes and view kwargs | | 5 | #23244 | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 | Reduction dimension guards | | 7 | #23246 | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename. Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed. NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: c6a7e94 ghstack-comment-id: 5892899249 Pull-Request: #23254
mergennachin
added a commit
that referenced
this pull request
Sep 29, 2026
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers. Fixes #23156. This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 | Hardware Vulkan CI | | 2 | #23241 | Scalar cache type and signed-zero keys | | 3 | #23242 | Vulkan-local signed-zero serialization | | 4 | #23243 | GELU modes and view kwargs | | 5 | #23244 | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 | Reduction dimension guards | | 7 | #23246 | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename. Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed. NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: f8d24d9 ghstack-comment-id: 5892899249 Pull-Request: #23254
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The general reduce implementation handles one or two reduced dims, but the partitioner accepted empty dim lists and full sum/mean reductions over N-D inputs, which then failed to build or computed the wrong reduction. Texture storage also folds batch into channels, so 4D reductions over dim 0, or over dim 1 with batch above 1, cannot be computed. These cases now stay on CPU.
Part 6/15 of the Vulkan transformer and operator-conformance stack. Depends on #23244; review against the selected base branch. Integration PR: #23254.
Validation: 2 passed, 2 warnings in 70.88s (0:01:10). Lintrunner and git diff --check pass. Native tests use MoltenVK with portable CPU kernels; hardware and SwiftShader CI are pending.
Recreates #23206 through ghstack. Prior review discussion remains on that PR.
Authored with OpenAI Codex; split planned with Claude Code.
cc @SS-JIA @manuelcandales @digantdesai @cbilgin