Skip to content

[Vulkan] Make expand_copy resizable and fully delegate dynamic transformer blocks - #23162

Closed
mergennachin wants to merge 1 commit into
mergennachin/vulkan-23156-14-anyfrom
mergennachin/vulkan-dynamic-transformer-23156
Closed

mergennachin wants to merge 1 commit into
mergennachin/vulkan-23156-14-anyfrom
mergennachin/vulkan-dynamic-transformer-23156

Conversation

@mergennachin

@mergennachin mergennachin commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Superseded by #23254, part 15/15 of the ghstack replacement series. The current integration PR is #23254. This PR is closed in favor of the replacement; its previous description and review history are retained.


The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes expand_copy resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers.

Fixes #23156.

This is part 15 of a linear stack. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10.

Part PR Change
1 #23201 Hardware Vulkan CI
2 #23202 Scalar cache type and signed-zero keys
3 #23203 FlatBuffers signed-zero serialization
4 #23204 GELU modes and view kwargs
5 #23205 Reduction clamp, NaN, and FP16 rounding
6 #23206 Reduction dimension guards
7 #23207 Bool staging and logical_not
8 #23208 Scalar representability and symbolic guards
9 #23209 64-bit dtype and fusion policy
10 #23210 scalar_tensor with exact integer values
11 #23211 Typed, resizable full
12 #23212 Power special values and logical FP16 dtype
13 #23213 mul.Scalar
14 #23214 any.dim
15 #23162 Dynamic expand and transformer integration

The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now test_vulkan_dynamic.py; CI and Buck references follow the rename.

Each part was linted and tested before publication. The combined changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at a31838280f9309f04af1b375147375b702d29339 found no new regressions or unresolved comparisons.

Validation uses the shared backends/test/ harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses texture_limits=(1,1,1), with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: 3b8c778c99766a8b4d0d04563ae0b16cbb276829, seed 0.

The complete native regression run on final head a2543b8a82 passed 50 tests with one expected SwiftShader-only skip on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. Shared FlatBuffers and serialization checks also passed (24 tests). Every stack commit passed lintrunner and git diff --check; GitHub lint and mypy checks passed for the final head.

The NVIDIA and SwiftShader jobs are queued for the final head; the Vulkan Windows build is in progress.

Authored with OpenAI Codex; split planned with Claude Code.

Copilot AI lite review requested due to automatic review settings September 25, 2026 17:54
@mergennachin mergennachin added module: vulkan Issues related to the Vulkan delegate and code under backends/vulkan/ release notes: vulkan Changes to the Vulkan backend delegate labels Sep 25, 2026
@pytorch-bot

pytorch-bot Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23162

Note: Links to docs will display an error until the docs builds have been completed.

❗ 1 Active SEVs

There are 1 currently active SEVs. If your PR is affected, please view them below:

⏳ No Failures, 1 Pending

As of commit a2543b8 with merge base a318382 (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 25, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unresolved critical scalar-conversion issues and a moderate dynamic SymInt registration issue block approval.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 2 High severity · 1 Medium severity

Open (3)
What changed in this PR

This PR enables full dynamic Vulkan delegation for transformer attention blocks, including GELU modes, scalar handling, resizing, reductions, and regression coverage.

Changes:

  • Adds operator registrations, runtime implementations, and shader variants.
  • Preserves scalar types and supports dynamic tensor operations.
  • Adds transformer tests and CI coverage.

Review findings: Two critical issues remain: integer full/full_like values can lose precision through float conversion, and int64 scalar tensors lack a correct runtime/shader path. One moderate issue rejects dynamic SymInt scalar arguments for pow and mul.

File Description
backends/​vulkan/​test/​test_vulkan_transformer.py Adds dynamic transformer regression tests.
backends/​vulkan/​test/​test_vulkan_graph_builder.py Tests scalar tensor operator naming.
backends/​vulkan/​test/​targets.bzl Adds the transformer test target.
backends/​vulkan/​test/​op_tests/​cases.py Expands scalar, GELU, and multiplication coverage.
backends/​vulkan/​serialization/​vulkan_graph_builder.py Preserves the ATen scalar tensor namespace.
backends/​vulkan/​runtime/​graph/​ops/​impl/​UnaryOp.cpp Selects GELU behavior and registers logical-not.
backends/​vulkan/​runtime/​graph/​ops/​impl/​ScalarTensor.cpp Implements scalar tensor parameter selection.
backends/​vulkan/​runtime/​graph/​ops/​impl/​Reduce.cpp Adds boolean reduction support.
backends/​vulkan/​runtime/​graph/​ops/​impl/​BinaryScalarOp.cpp Adds scalar multiplication support.
backends/​vulkan/​runtime/​graph/​ops/​glsl/​unary_op.yaml Adds the exact GELU shader variant.
backends/​vulkan/​runtime/​graph/​ops/​glsl/​reduce.yaml Adds texture boolean reduction variants.
backends/​vulkan/​runtime/​graph/​ops/​glsl/​reduce.glsl Supports non-floating reduction types.
backends/​vulkan/​runtime/​graph/​ops/​glsl/​reduce_per_row_buffer.yaml Adds buffer boolean reduction variants.
backends/​vulkan/​runtime/​graph/​ops/​glsl/​reduce_per_row_buffer.glsl Handles boolean reduction output.
backends/​vulkan/​runtime/​graph/​ops/​glsl/​binary_scalar_texture.yaml Adds texture scalar multiplication shaders.
backends/​vulkan/​runtime/​graph/​ops/​glsl/​binary_scalar_buffer.yaml Adds buffer scalar multiplication shaders.
backends/​vulkan/​runtime/​graph/​ops/​glsl/​activations.h Implements erf-based GELU.
backends/​vulkan/​op_registry.py Registers operators and dynamic-shape support.
.github/​workflows/​vulkan.yml Runs transformer tests in Vulkan CI.
.github/​workflows/​pull.yml Runs transformer tests in pull-request CI.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread backends/vulkan/op_registry.py
Comment thread backends/vulkan/runtime/graph/ops/impl/ScalarTensor.cpp
Comment thread backends/vulkan/op_registry.py Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Three unresolved findings remain, including one critical scalar-range issue.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 3 High severity · 2 Medium severity

Open (5)

Comment thread backends/vulkan/op_registry.py Outdated
Comment thread backends/vulkan/runtime/VulkanBackend.cpp

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unresolved int64 fill/scalar-tensor handling and incomplete bool-buffer capability checks can cause shader failures or incorrect execution.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 3 High severity · 2 Medium severity

Open (5)
Resolved since last review (1)

Comment thread backends/vulkan/runtime/graph/ops/impl/Full.cpp

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot AI review requested due to automatic review settings September 28, 2026 17:44

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

…ormer blocks

expand_copy already had a resize function but was marked supports_resize=False, so under require_dynamic_shapes it split masked attention blocks. With it marked resizable, a transformer block with a length-derived attention mask lowers to a single Vulkan delegate for both the SDPA and explicit-softmax formulations and runs correctly across sequence lengths.

Fixes #23156

Authored with OpenAI Codex; split planned with Claude Code.
@mergennachin mergennachin changed the title [Vulkan] Fully delegate dynamic transformer blocks [Vulkan] Make expand_copy resizable and fully delegate dynamic transformer blocks Sep 28, 2026
@mergennachin
mergennachin force-pushed the mergennachin/vulkan-dynamic-transformer-23156 branch from 1634b9e to a2543b8 Compare September 28, 2026 22:30
@mergennachin
mergennachin changed the base branch from main to mergennachin/vulkan-23156-14-any September 28, 2026 22:30
mergennachin added a commit that referenced this pull request Sep 29, 2026
…ormer blocks

The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers.

Fixes #23156.

This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10.

| Part | PR | Change |
| --- | --- | --- |
| 1 | #23240 | Hardware Vulkan CI |
| 2 | #23241 | Scalar cache type and signed-zero keys |
| 3 | #23242 | Vulkan-local signed-zero serialization |
| 4 | #23243 | GELU modes and view kwargs |
| 5 | #23244 | Reduction clamp, NaN, and FP16 rounding |
| 6 | #23245 | Reduction dimension guards |
| 7 | #23246 | Bool staging and logical_not |
| 8 | #23247 | Scalar representability and symbolic guards |
| 9 | #23248 | 64-bit dtype and fusion policy |
| 10 | #23249 | scalar_tensor with exact integer values |
| 11 | #23250 | Typed, resizable full |
| 12 | #23251 | Power special values and logical FP16 dtype |
| 13 | #23252 | mul.Scalar |
| 14 | #23253 | any.dim |
| 15 | #23254 | Dynamic expand and transformer integration |

The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename.

Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons.

Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0.

The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed.

NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads.

Recreates #23162 through ghstack. Prior review discussion remains on that PR.

Authored with OpenAI Codex; split planned with Claude Code.


cc @SS-JIA @manuelcandales @digantdesai @cbilgin

ghstack-source-id: c6a7e94
ghstack-comment-id: 5892899249
Pull-Request: #23254
mergennachin added a commit that referenced this pull request Sep 29, 2026
…ormer blocks

The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers.

Fixes #23156.

This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10.

| Part | PR | Change |
| --- | --- | --- |
| 1 | #23240 | Hardware Vulkan CI |
| 2 | #23241 | Scalar cache type and signed-zero keys |
| 3 | #23242 | Vulkan-local signed-zero serialization |
| 4 | #23243 | GELU modes and view kwargs |
| 5 | #23244 | Reduction clamp, NaN, and FP16 rounding |
| 6 | #23245 | Reduction dimension guards |
| 7 | #23246 | Bool staging and logical_not |
| 8 | #23247 | Scalar representability and symbolic guards |
| 9 | #23248 | 64-bit dtype and fusion policy |
| 10 | #23249 | scalar_tensor with exact integer values |
| 11 | #23250 | Typed, resizable full |
| 12 | #23251 | Power special values and logical FP16 dtype |
| 13 | #23252 | mul.Scalar |
| 14 | #23253 | any.dim |
| 15 | #23254 | Dynamic expand and transformer integration |

The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename.

Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons.

Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0.

The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed.

NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads.

Recreates #23162 through ghstack. Prior review discussion remains on that PR.

Authored with OpenAI Codex; split planned with Claude Code.


cc @SS-JIA @manuelcandales @digantdesai @cbilgin

ghstack-source-id: f8d24d9
ghstack-comment-id: 5892899249
Pull-Request: #23254

This branch was successfully deployed

1 active deployment
cadence — a2543b8a Deployed Sep 28, 2026 by mergennachin via hifi-op-test / hifi4 #30222
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. module: vulkan Issues related to the Vulkan delegate and code under backends/vulkan/ release notes: vulkan Changes to the Vulkan backend delegate

Projects

None yet

2 participants