Skip to content

[Vulkan] Make expand_copy resizable and fully delegate dynamic transformer blocks - #23254

Draft
mergennachin wants to merge 2 commits into
gh/mergennachin/37/headfrom
gh/mergennachin/38/head
Draft

mergennachin wants to merge 2 commits into
gh/mergennachin/37/headfrom
gh/mergennachin/38/head

Conversation

@mergennachin

@mergennachin mergennachin commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes expand_copy resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers.

Fixes #23156.

This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10.

Part PR Change
1 #23240 Hardware Vulkan CI
2 #23241 Scalar cache type and signed-zero keys
3 #23242 Vulkan-local signed-zero serialization
4 #23243 GELU modes and view kwargs
5 #23244 Reduction clamp, NaN, and FP16 rounding
6 #23245 Reduction dimension guards
7 #23246 Bool staging and logical_not
8 #23247 Scalar representability and symbolic guards
9 #23248 64-bit dtype and fusion policy
10 #23249 scalar_tensor with exact integer values
11 #23250 Typed, resizable full
12 #23251 Power special values and logical FP16 dtype
13 #23252 mul.Scalar
14 #23253 any.dim
15 #23254 Dynamic expand and transformer integration

The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. On native FP16 devices, the scalar-tensor, full texture, single-dimension texture reduction, and binary scalar shaders round their FP16 outputs to nearest-even. The binary scalar shaders also preserve the requested FP16 dtype when storage is emulated with FP32. Other operators retain their existing FP32 intermediate behavior on those devices. PR4 introduces test_vulkan_dynamic.py and its CI/Buck references.

Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at a31838280f9309f04af1b375147375b702d29339 found no new regressions or unresolved comparisons.

Validation uses the shared backends/test/ harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses texture_limits=(1,1,1), with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: 3b8c778c99766a8b4d0d04563ae0b16cbb276829, seed 0.

The complete native regression run on the previously published head a2543b8a82 passed 50 tests with one expected SwiftShader-only skip on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines --force-defaults to Vulkan and restores the shared exir FlatBuffers API. That revised source tree passed 26 focused tests, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and git diff --check passed.

The prior Vulkan NVIDIA, SwiftShader, and Windows checks have completed successfully. CI validation of the new exponent-conversion and chained-multiplication tests is pending on the updated ghstack heads.

Recreates #23162 through ghstack. Prior review discussion remains on that PR.

Follow-up regression coverage adds scalar exponents 2.0001 and 2049 and a two-node FP16 mul.Scalar chain. The power test and chained-multiplication/signed-zero tests pass individually on MoltenVK; the full dynamic suite on the restacked tip passed 35 tests with one expected SwiftShader-only skip on MoltenVK. Lintrunner and git diff --check passed. These changes preserve the existing runtime and shader implementation.

Authored with OpenAI Codex; split planned with Claude Code.

cc @SS-JIA @manuelcandales @digantdesai @cbilgin

@mergennachin mergennachin added the module: vulkan Issues related to the Vulkan delegate and code under backends/vulkan/ label Sep 29, 2026
@pytorch-bot

pytorch-bot Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23254

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 Pending, 1 Unclassified Failure

As of commit b82ffd0 with merge base a318382 (image):

UNCLASSIFIED FAILURE - DrCI could not classify the following job because the workflow did not run on the merge base. The failure may be pre-existing on trunk or introduced by this PR:

  • Test CoreML Backend (gh) (this job did not run on the merge base, so DrCI cannot tell whether the failure is pre-existing)

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@mergennachin mergennachin added the release notes: vulkan Changes to the Vulkan backend delegate label Sep 29, 2026
mergennachin added a commit that referenced this pull request Sep 29, 2026
…ormer blocks

The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers.

Fixes #23156.

This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10.

| Part | PR | Change |
| --- | --- | --- |
| 1 | #23240 | Hardware Vulkan CI |
| 2 | #23241 | Scalar cache type and signed-zero keys |
| 3 | #23242 | Vulkan-local signed-zero serialization |
| 4 | #23243 | GELU modes and view kwargs |
| 5 | #23244 | Reduction clamp, NaN, and FP16 rounding |
| 6 | #23245 | Reduction dimension guards |
| 7 | #23246 | Bool staging and logical_not |
| 8 | #23247 | Scalar representability and symbolic guards |
| 9 | #23248 | 64-bit dtype and fusion policy |
| 10 | #23249 | scalar_tensor with exact integer values |
| 11 | #23250 | Typed, resizable full |
| 12 | #23251 | Power special values and logical FP16 dtype |
| 13 | #23252 | mul.Scalar |
| 14 | #23253 | any.dim |
| 15 | #23254 | Dynamic expand and transformer integration |

The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename.

Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons.

Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0.

The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed.

NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads.

Recreates #23162 through ghstack. Prior review discussion remains on that PR.

Authored with OpenAI Codex; split planned with Claude Code.


cc @SS-JIA @manuelcandales @digantdesai @cbilgin

ghstack-source-id: c6a7e94
ghstack-comment-id: 5892899249
Pull-Request: #23254
@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 29, 2026
mergennachin added a commit that referenced this pull request Sep 29, 2026
…ormer blocks

The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers.

Fixes #23156.

This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10.

| Part | PR | Change |
| --- | --- | --- |
| 1 | #23240 | Hardware Vulkan CI |
| 2 | #23241 | Scalar cache type and signed-zero keys |
| 3 | #23242 | Vulkan-local signed-zero serialization |
| 4 | #23243 | GELU modes and view kwargs |
| 5 | #23244 | Reduction clamp, NaN, and FP16 rounding |
| 6 | #23245 | Reduction dimension guards |
| 7 | #23246 | Bool staging and logical_not |
| 8 | #23247 | Scalar representability and symbolic guards |
| 9 | #23248 | 64-bit dtype and fusion policy |
| 10 | #23249 | scalar_tensor with exact integer values |
| 11 | #23250 | Typed, resizable full |
| 12 | #23251 | Power special values and logical FP16 dtype |
| 13 | #23252 | mul.Scalar |
| 14 | #23253 | any.dim |
| 15 | #23254 | Dynamic expand and transformer integration |

The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename.

Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons.

Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0.

The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed.

NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads.

Recreates #23162 through ghstack. Prior review discussion remains on that PR.

Authored with OpenAI Codex; split planned with Claude Code.


cc @SS-JIA @manuelcandales @digantdesai @cbilgin

ghstack-source-id: f8d24d9
ghstack-comment-id: 5892899249
Pull-Request: #23254

This branch was successfully deployed

1 active deployment
cadence — b82ffd01 Deployed Sep 29, 2026 by mergennachin via hifi-op-test / hifi4 #30492
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. module: vulkan Issues related to the Vulkan delegate and code under backends/vulkan/ release notes: vulkan Changes to the Vulkan backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant