[Vulkan] Make expand_copy resizable and fully delegate dynamic transformer blocks - #23162
mergennachin wants to merge 1 commit into
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23162
Note: Links to docs will display an error until the docs builds have been completed. ❗ 1 Active SEVsThere are 1 currently active SEVs. If your PR is affected, please view them below: ⏳ No Failures, 1 PendingAs of commit a2543b8 with merge base a318382 ( This comment was automatically generated by Dr. CI and updates every 15 minutes. |
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Unresolved critical scalar-conversion issues and a moderate dynamic SymInt registration issue block approval.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 2
Open (3)
What changed in this PR
This PR enables full dynamic Vulkan delegation for transformer attention blocks, including GELU modes, scalar handling, resizing, reductions, and regression coverage.
Changes:
- Adds operator registrations, runtime implementations, and shader variants.
- Preserves scalar types and supports dynamic tensor operations.
- Adds transformer tests and CI coverage.
Review findings: Two critical issues remain: integer full/full_like values can lose precision through float conversion, and int64 scalar tensors lack a correct runtime/shader path. One moderate issue rejects dynamic SymInt scalar arguments for pow and mul.
| File | Description |
|---|---|
backends/vulkan/test/test_vulkan_transformer.py |
Adds dynamic transformer regression tests. |
backends/vulkan/test/test_vulkan_graph_builder.py |
Tests scalar tensor operator naming. |
backends/vulkan/test/targets.bzl |
Adds the transformer test target. |
backends/vulkan/test/op_tests/cases.py |
Expands scalar, GELU, and multiplication coverage. |
backends/vulkan/serialization/vulkan_graph_builder.py |
Preserves the ATen scalar tensor namespace. |
backends/vulkan/runtime/graph/ops/impl/UnaryOp.cpp |
Selects GELU behavior and registers logical-not. |
backends/vulkan/runtime/graph/ops/impl/ScalarTensor.cpp |
Implements scalar tensor parameter selection. |
backends/vulkan/runtime/graph/ops/impl/Reduce.cpp |
Adds boolean reduction support. |
backends/vulkan/runtime/graph/ops/impl/BinaryScalarOp.cpp |
Adds scalar multiplication support. |
backends/vulkan/runtime/graph/ops/glsl/unary_op.yaml |
Adds the exact GELU shader variant. |
backends/vulkan/runtime/graph/ops/glsl/reduce.yaml |
Adds texture boolean reduction variants. |
backends/vulkan/runtime/graph/ops/glsl/reduce.glsl |
Supports non-floating reduction types. |
backends/vulkan/runtime/graph/ops/glsl/reduce_per_row_buffer.yaml |
Adds buffer boolean reduction variants. |
backends/vulkan/runtime/graph/ops/glsl/reduce_per_row_buffer.glsl |
Handles boolean reduction output. |
backends/vulkan/runtime/graph/ops/glsl/binary_scalar_texture.yaml |
Adds texture scalar multiplication shaders. |
backends/vulkan/runtime/graph/ops/glsl/binary_scalar_buffer.yaml |
Adds buffer scalar multiplication shaders. |
backends/vulkan/runtime/graph/ops/glsl/activations.h |
Implements erf-based GELU. |
backends/vulkan/op_registry.py |
Registers operators and dynamic-shape support. |
.github/workflows/vulkan.yml |
Runs transformer tests in Vulkan CI. |
.github/workflows/pull.yml |
Runs transformer tests in pull-request CI. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Three unresolved findings remain, including one critical scalar-range issue.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 3
Open (5)
Reject out-of-range float scalars before int32 conversion · New Long scalar_tensor outputs incorrectly use float extraction Dynamic integer full values lose precision through float conversion Reject constant bool tensors requiring unsupported buffer storage · New Rejects SymInt scalar arguments for Vulkan binary operations
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Unresolved int64 fill/scalar-tensor handling and incomplete bool-buffer capability checks can cause shader failures or incorrect execution.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 3
Open (5)
Unsupported int64 output falls through to float shader path · New Long scalar_tensor outputs incorrectly use float extraction Dynamic integer full values lose precision through float conversion Reject constant bool tensors requiring unsupported buffer storage Rejects SymInt scalar arguments for Vulkan binary operations
Resolved since last review (1)
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
The broad Vulkan runtime, shader, partitioning, and CI changes require final human review.
Review effort: Lite
Findings: None
Resolved since last review (6)
Unsupported int64 output falls through to float shader path Reject out-of-range float scalars before int32 conversion Long scalar_tensor outputs incorrectly use float extraction Dynamic integer full values lose precision through float conversion Reject constant bool tensors requiring unsupported buffer storage Rejects SymInt scalar arguments for Vulkan binary operations
…ormer blocks expand_copy already had a resize function but was marked supports_resize=False, so under require_dynamic_shapes it split masked attention blocks. With it marked resizable, a transformer block with a length-derived attention mask lowers to a single Vulkan delegate for both the SDPA and explicit-softmax formulations and runs correctly across sequence lengths. Fixes #23156 Authored with OpenAI Codex; split planned with Claude Code.
1634b9e to
a2543b8
Compare
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers. Fixes #23156. This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 | Hardware Vulkan CI | | 2 | #23241 | Scalar cache type and signed-zero keys | | 3 | #23242 | Vulkan-local signed-zero serialization | | 4 | #23243 | GELU modes and view kwargs | | 5 | #23244 | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 | Reduction dimension guards | | 7 | #23246 | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename. Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed. NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: c6a7e94 ghstack-comment-id: 5892899249 Pull-Request: #23254
…ormer blocks The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers. Fixes #23156. This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10. | Part | PR | Change | | --- | --- | --- | | 1 | #23240 | Hardware Vulkan CI | | 2 | #23241 | Scalar cache type and signed-zero keys | | 3 | #23242 | Vulkan-local signed-zero serialization | | 4 | #23243 | GELU modes and view kwargs | | 5 | #23244 | Reduction clamp, NaN, and FP16 rounding | | 6 | #23245 | Reduction dimension guards | | 7 | #23246 | Bool staging and logical_not | | 8 | #23247 | Scalar representability and symbolic guards | | 9 | #23248 | 64-bit dtype and fusion policy | | 10 | #23249 | scalar_tensor with exact integer values | | 11 | #23250 | Typed, resizable full | | 12 | #23251 | Power special values and logical FP16 dtype | | 13 | #23252 | mul.Scalar | | 14 | #23253 | any.dim | | 15 | #23254 | Dynamic expand and transformer integration | The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename. Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons. Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0. The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed. NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads. Recreates #23162 through ghstack. Prior review discussion remains on that PR. Authored with OpenAI Codex; split planned with Claude Code. cc @SS-JIA @manuelcandales @digantdesai @cbilgin ghstack-source-id: f8d24d9 ghstack-comment-id: 5892899249 Pull-Request: #23254


Superseded by #23254, part 15/15 of the ghstack replacement series. The current integration PR is #23254. This PR is closed in favor of the replacement; its previous description and review history are retained.
The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes
expand_copyresizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers.Fixes #23156.
This is part 15 of a linear stack. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10.
The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now
test_vulkan_dynamic.py; CI and Buck references follow the rename.Each part was linted and tested before publication. The combined changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at
a31838280f9309f04af1b375147375b702d29339found no new regressions or unresolved comparisons.Validation uses the shared
backends/test/harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference usestexture_limits=(1,1,1), with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision:3b8c778c99766a8b4d0d04563ae0b16cbb276829, seed 0.The complete native regression run on final head
a2543b8a82passed 50 tests with one expected SwiftShader-only skip on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. Shared FlatBuffers and serialization checks also passed (24 tests). Every stack commit passed lintrunner andgit diff --check; GitHub lint and mypy checks passed for the final head.The NVIDIA and SwiftShader jobs are queued for the final head; the Vulkan Windows build is in progress.
Authored with OpenAI Codex; split planned with Claude Code.