Skip to content

[Vulkan] Pass typed fill values to full and make it resizable - #23250

Draft
mergennachin wants to merge 1 commit into
gh/mergennachin/33/headfrom
gh/mergennachin/34/head
Draft

mergennachin wants to merge 1 commit into
gh/mergennachin/33/headfrom
gh/mergennachin/34/head

Conversation

@mergennachin

@mergennachin mergennachin commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

The full shaders always received the fill value as a float, which loses int32 fills above 2^24 and is ill-defined for bools. Full.cpp now passes the fill value in the output's accumulator type. full, full_like, zeros, ones and their _like variants are marked resizable so they can live inside dynamic-shape delegates, and fill values the output dtype cannot represent stay on CPU.

Part 11/15 of the Vulkan transformer and operator-conformance stack. Depends on #23249; review against the selected base branch. Integration PR: #23254.

Validation: 6 passed, 2 warnings in 113.26s (0:01:53). Lintrunner and git diff --check pass. Native tests use MoltenVK with portable CPU kernels; hardware and SwiftShader CI are pending.

Recreates #23211 through ghstack. Prior review discussion remains on that PR.

Authored with OpenAI Codex; split planned with Claude Code.

cc @SS-JIA @manuelcandales @digantdesai @cbilgin

@mergennachin mergennachin added the module: vulkan Issues related to the Vulkan delegate and code under backends/vulkan/ label Sep 29, 2026
@pytorch-bot

pytorch-bot Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23250

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures, 3 Pending

As of commit bebd775 with merge base a318382 (image):

NEW FAILURES - The following jobs have failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

mergennachin added a commit that referenced this pull request Sep 29, 2026
…ormer blocks

The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers.

Fixes #23156.

This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10.

| Part | PR | Change |
| --- | --- | --- |
| 1 | #23240 | Hardware Vulkan CI |
| 2 | #23241 | Scalar cache type and signed-zero keys |
| 3 | #23242 | Vulkan-local signed-zero serialization |
| 4 | #23243 | GELU modes and view kwargs |
| 5 | #23244 | Reduction clamp, NaN, and FP16 rounding |
| 6 | #23245 | Reduction dimension guards |
| 7 | #23246 | Bool staging and logical_not |
| 8 | #23247 | Scalar representability and symbolic guards |
| 9 | #23248 | 64-bit dtype and fusion policy |
| 10 | #23249 | scalar_tensor with exact integer values |
| 11 | #23250 | Typed, resizable full |
| 12 | #23251 | Power special values and logical FP16 dtype |
| 13 | #23252 | mul.Scalar |
| 14 | #23253 | any.dim |
| 15 | #23254 | Dynamic expand and transformer integration |

The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename.

Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons.

Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0.

The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed.

NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads.

Recreates #23162 through ghstack. Prior review discussion remains on that PR.

Authored with OpenAI Codex; split planned with Claude Code.


cc @SS-JIA @manuelcandales @digantdesai @cbilgin

ghstack-source-id: c6a7e94
ghstack-comment-id: 5892899249
Pull-Request: #23254
@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 29, 2026
mergennachin added a commit that referenced this pull request Sep 29, 2026
…ormer blocks

The dynamic transformer reproductions now lower to one Vulkan delegate across changing sequence lengths. This final PR makes `expand_copy` resizable and adds the eager-attention and SDPA integration tests, including fully masked rows on devices with 8-bit storage buffers.

Fixes #23156.

This is part 15 of a linear ghstack series. Review each PR against its selected base branch; this PR contains only the final expansion and transformer changes. The scalar-tensor issue #23158 is owned by part 10.

| Part | PR | Change |
| --- | --- | --- |
| 1 | #23240 | Hardware Vulkan CI |
| 2 | #23241 | Scalar cache type and signed-zero keys |
| 3 | #23242 | Vulkan-local signed-zero serialization |
| 4 | #23243 | GELU modes and view kwargs |
| 5 | #23244 | Reduction clamp, NaN, and FP16 rounding |
| 6 | #23245 | Reduction dimension guards |
| 7 | #23246 | Bool staging and logical_not |
| 8 | #23247 | Scalar representability and symbolic guards |
| 9 | #23248 | 64-bit dtype and fusion policy |
| 10 | #23249 | scalar_tensor with exact integer values |
| 11 | #23250 | Typed, resizable full |
| 12 | #23251 | Power special values and logical FP16 dtype |
| 13 | #23252 | mul.Scalar |
| 14 | #23253 | any.dim |
| 15 | #23254 | Dynamic expand and transformer integration |

The stack retains the FACTO and ATen conformance fixes and their regression tests: typed integer fills, NaN and signed-zero behavior, reduction range and accumulation fixes, safe scalar fallbacks, and the FP16 power contract. FP16 writes use explicit nearest-even conversion, and power keeps the requested tensor dtype when a device emulates FP16 storage with FP32. The test module is now `test_vulkan_dynamic.py`; CI and Buck references follow the rename.

Each code patch was linted and tested before its original publication. Prior validation of the operator changes passed 116 FACTO cases against both ATen and portable kernels, plus 144 boundary cases against ATen, with texture and buffer preferences: all 520 Vulkan-configuration executions matched ATen, including CPU fallback where unsupported. The additional 239-case sweep covers all 91 available FACTO specifications for registered ATen overloads and 57 targeted examples. Exact-input replays of every candidate Vulkan failure on a separately built main at `a31838280f9309f04af1b375147375b702d29339` found no new regressions or unresolved comparisons.

Validation uses the shared `backends/test/` harness on Apple M1 Pro / MoltenVK with Release portable CPU kernels and Vulkan; optimized kernels and XNNPACK are disabled. Buffer preference uses `texture_limits=(1,1,1)`, with actual storage and fallback recorded separately. Single-texel tensors can still select textures. FACTO revision: `3b8c778c99766a8b4d0d04563ae0b16cbb276829`, seed 0.

The complete native regression run on the previously published head `a2543b8a82` passed **50 tests with one expected SwiftShader-only skip** on MoltenVK, covering all 35 dynamic tests, graph-builder and serialization tests, and PT2E quantized linear with downcasting disabled. The subsequent serializer refinement confines `--force-defaults` to Vulkan and restores the shared `exir` FlatBuffers API. That revised source tree passed **26 focused tests**, including shared/Vulkan serialization, signed-zero execution, and power special values. All other code patches are unchanged, and every replacement PR was checked against the tested source trees. Lintrunner and `git diff --check` passed.

NVIDIA, SwiftShader, and Windows CI are pending verification on the new ghstack heads.

Recreates #23162 through ghstack. Prior review discussion remains on that PR.

Authored with OpenAI Codex; split planned with Claude Code.


cc @SS-JIA @manuelcandales @digantdesai @cbilgin

ghstack-source-id: f8d24d9
ghstack-comment-id: 5892899249
Pull-Request: #23254

This branch was successfully deployed

1 active deployment
cadence — bebd7753 Deployed Sep 29, 2026 by mergennachin via hifi-op-test / hifi4 #30422
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. module: vulkan Issues related to the Vulkan delegate and code under backends/vulkan/

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant