You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Enable ConvTranspose + BatchNorm fusion in the Qualcomm/QNN pass pipeline after #23170 lands.
FuseBatchNormWithConv.can_fuse currently rejects all transposed convolutions. This workaround avoids the incorrect output-channel scaling in the shared pass discussed in #22994. PR #23170 fixes that shared pass, including grouped transposed weights, but leaves the QNN guard in place. Consequently, QNN will continue to retain standalone BatchNorm operations after that fix.
Add Qualcomm regression cases for ConvTranspose + BatchNorm, covering equal and unequal input/output channel counts, grouped convolutions, and convolution bias enabled/disabled where supported by QNN.
Use nontrivial BatchNorm running statistics and affine parameters. Assert that BatchNorm is folded and the transformed output matches eager execution.
Run QNN delegation/execution coverage for the supported configurations and retain coverage for ordinary convolution fusion.
Additional context
Local CPU review validation of #23170 with PyTorch 2.13 passed all six new tests; four reproduce the original bug with the pre-fix pass. Eighteen additional grouped/depthwise ConvTranspose1d/2d/3d cases passed through the exported-program transform path. These checks validate the shared pass; QNN integration still needs the coverage above.
Feature and motivation
Enable
ConvTranspose + BatchNormfusion in the Qualcomm/QNN pass pipeline after #23170 lands.FuseBatchNormWithConv.can_fusecurrently rejects all transposed convolutions. This workaround avoids the incorrect output-channel scaling in the shared pass discussed in #22994. PR #23170 fixes that shared pass, including grouped transposed weights, but leaves the QNN guard in place. Consequently, QNN will continue to retain standalone BatchNorm operations after that fix.Proposed work
ConvTranspose + BatchNorm, covering equal and unequal input/output channel counts, grouped convolutions, and convolution bias enabled/disabled where supported by QNN.Additional context
Local CPU review validation of #23170 with PyTorch 2.13 passed all six new tests; four reproduce the original bug with the pre-fix pass. Eighteen additional grouped/depthwise ConvTranspose1d/2d/3d cases passed through the exported-program transform path. These checks validate the shared pass; QNN integration still needs the coverage above.
cc @cccclai @winskuo-quic @shewu-quic @haowhsu-quic @DannyYuyang-quic @cbilgin @abhinaykukkadapu @psiddh