Skip to content

[Vulkan] Stop clamping per-row reductions and propagate NaN - #23205

Closed
mergennachin wants to merge 1 commit into
mergennachin/vulkan-23156-04-gelufrom
mergennachin/vulkan-23156-05-reduction-contract
Closed

mergennachin wants to merge 1 commit into
mergennachin/vulkan-23156-04-gelufrom
mergennachin/vulkan-23156-05-reduction-contract

Conversation

@mergennachin

@mergennachin mergennachin commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Superseded by #23244, part 5/15 of the ghstack replacement series. The current integration PR is #23254. This PR is closed in favor of the replacement; its previous description and review history are retained.


convert.glslh guarded an fp16 clamp with #if T == float16_t, but neither side is a macro, so the preprocessor compares 0 with 0 and the clamp was always on. Per-row buffer sum, mean, amax and amin therefore saturated fp32 and int results at 65504. The clamp is removed. Single-dimension texture and per-row buffer amax and amin now propagate NaN, argmax and argmin return the index of the first NaN as ATen does, and mean divides in the accumulator type. FP16 texture output uses explicit nearest-even conversion so rounding and overflow agree with ATen across drivers.

Part 5/15 of the Vulkan transformer and operator-conformance stack. Depends on #23204; review against the selected base branch. Integration PR: #23162.

Validation: 3 passed, 2 warnings in 57.33s. Lintrunner and git diff --check pass. Native tests use MoltenVK with portable CPU kernels; hardware and SwiftShader CI are pending.

Authored with OpenAI Codex; split planned with Claude Code.

cc @SS-JIA @manuelcandales @digantdesai @cbilgin

convert.glslh guarded an fp16 clamp with `#if T == float16_t`, but neither side is a macro, so the preprocessor compares 0 with 0 and the clamp was always on. Per-row buffer sum, mean, amax and amin therefore saturated fp32 and int results at 65504. The clamp is removed. Single-dimension texture and per-row buffer amax and amin now propagate NaN, argmax and argmin return the index of the first NaN as ATen does, and mean divides in the accumulator type. FP16 texture output uses explicit nearest-even conversion so rounding and overflow agree with ATen across drivers.

Authored with OpenAI Codex; split planned with Claude Code.
@pytorch-bot pytorch-bot Bot added the module: vulkan Issues related to the Vulkan delegate and code under backends/vulkan/ label Sep 28, 2026
@pytorch-bot

pytorch-bot Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23205

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit 000b4ae with merge base a318382 (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

This branch was successfully deployed

1 active deployment
cadence — 000b4aea Deployed Sep 28, 2026 by mergennachin via hifi-op-test / hifi4 #30191
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. module: vulkan Issues related to the Vulkan delegate and code under backends/vulkan/

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant