Expected behavior
op-support.csv lists slice_copy for fp32 with constraints 'no zero-dim tensors, no dynamic shapes'. A static-shape float32 slice with step 2 and no zero dims satisfies every stated constraint, so with a default XnnpackPartitioner (documented to run as much as possible on XNNPACK) it should delegate exactly like the step-1 slice.
Actual behavior
Static float32 slice with step 2 (shape (1,28,28,96), explicit end, no zero dims) yields 0 delegate calls; the leftover aten.slice_copy.Tensor falls back to portable kernels. The identical slice with step 1 (open or explicit end) yields 1 delegate call. Upstream swin_t shows the same pattern: all 24 non-delegated slices are PatchMerging step-2 slices; the partitioner debug log gives no reason for them (the stride check returns False silently), while the 20 leftover int64 adds are correctly rejected as invalid dtype. Numerics are unaffected: XNNPACK .pte matches eager with allclose rtol=1e-3 atol=1e-5 True on 3 seeded inputs (max abs <= 7.7e-06, top-1 agrees), same as the portable control. Perf at 1 thread (warmup 20, 3 alternating x10 batches, pinned CPU, threadpool verified 1): XNNPACK medians 129.27/128.88/129.85 ms vs portable 7062.49/7072.42/7068.93 ms, every XNNPACK sample below every control sample, max load/CPU 0.045; expected direction, no anomaly.
Model and versions
Model: https://github.com/pytorch/vision/blob/9769007f09193b307a90d2b833aac31d364f03e3/torchvision/models/swin_transformer.py
ExecuTorch 1.6.0+6c8bc89 from pinned public source (pybind preset, Release, XNNPACK ON; KERNELS_OPTIMIZED/LLM/LLM-AOT/EXT-LLM/LLM-RUNNER/OPENVINO OFF; -j4), torch 2.14.0+cpu, torchvision 0.29.0+cpu (swin_transformer.py sha256 9ec73f23 byte-identical to pinned vision commit), torchao 0.18.0+cpu, submodules XNNPACK 92a7ad50 cpuinfo f9a03241 pthreadpool a56dcd79 FP16 4dfe081c FXdiv b408327a flatbuffers 595bf000 flatcc 896db547 gflags a738fdf9 json ac0133ea pybind11 d03662f0, gcc 11.5, cmake 3.31.8, AMD EPYC-Genoa x86_64
Documentation checked
https://raw.githubusercontent.com/pytorch/executorch/6c8bc89be41b3fe3fb42ce401556ecad65df7919/docs/source/backends/xnnpack/op-support.csv
slice_copy,"fp16, fp32",,"no zero-dim tensors, no dynamic shapes"
https://raw.githubusercontent.com/pytorch/executorch/6c8bc89be41b3fe3fb42ce401556ecad65df7919/docs/source/backends/xnnpack/xnnpack-partitioner.rst
Passing an XnnpackPartitioner instance with no additional parameters will run as much of the model as possible on the XNNPACK backend.
Reproduction
Save repro.sh, repro.py, requirements.txt into one fresh directory on a Linux x86_64 CPU host with python3 (3.10-3.14), git, gcc, cmake; run bash repro.sh (creates repro_env, installs pinned public wheels, clones pinned public sources, builds ExecuTorch from source, prints GAP_REPRODUCED). Repeat without redownloading/rebuilding via REPRO_VENV=
/repro_env bash repro.sh.
repro.sh
#!/bin/bash
# Reproducer: XNNPACK slice_copy with stride != 1 never delegates, although
# op-support.csv only excludes zero-dim tensors and dynamic shapes.
#
# Fresh mode (default): creates ./repro_env, installs pinned public deps,
# clones pinned public sources, builds ExecuTorch from source, runs repro.py.
# bash repro.sh
# Reuse mode: skip download/build with an existing venv that already has the
# pinned torch + a built executorch (e.g. a previous fresh run):
# REPRO_VENV=/path/to/repro_env bash repro.sh
#
# Expects a Linux x86_64 CPU host with python3 (3.10-3.14), git, gcc, cmake.
set -euo pipefail
ET_COMMIT=6c8bc89be41b3fe3fb42ce401556ecad65df7919
HERE="$(cd "$(dirname "$0")" && pwd)"
cd "$HERE"
if [ -n "${REPRO_VENV:-}" ]; then
echo "reuse mode: using $REPRO_VENV"
"$REPRO_VENV/bin/python" repro.py
exit 0
fi
VENV="$HERE/repro_env"
python3 -m venv "$VENV"
"$VENV/bin/pip" install --upgrade pip setuptools wheel
# Pinned torch/torchao CPU wheels + executorch runtime deps (see requirements.txt).
"$VENV/bin/pip" install -r requirements.txt
# Deterministic build tools: prefer a working cmake; neutralize host ccache
# (passthrough shim) so the build never depends on ccache behavior.
mkdir -p build_bin
printf '#!/bin/sh\nexec "$@"\n' > build_bin/ccache
chmod +x build_bin/ccache
if cmake --version >/dev/null 2>&1; then
CMAKE_BIN="$(command -v cmake)"
else
CMAKE_BIN=/usr/bin/cmake
fi
ln -sf "$CMAKE_BIN" build_bin/cmake
export PATH="$HERE/build_bin:$PATH"
cmake --version | head -1
# Pinned public sources, minimal submodules for a pybind+XNNPACK build.
if [ ! -d executorch ]; then
git clone https://github.com/pytorch/executorch.git executorch
fi
cd executorch
git checkout -q "$ET_COMMIT"
git submodule update --init third-party/json third-party/gflags \
third-party/pybind11 third-party/flatbuffers third-party/flatcc \
backends/xnnpack/third-party/FP16 backends/xnnpack/third-party/FXdiv \
backends/xnnpack/third-party/XNNPACK backends/xnnpack/third-party/cpuinfo \
backends/xnnpack/third-party/pthreadpool
git submodule status | grep -E "XNNPACK$|pybind11|flatbuffers" || true
cd "$HERE"
# Build from source (Release). Optimized kernels need an unreachable Eigen
# host, LLM/tokenizer parts are unnecessary for this CV repro, and the Linux
# pip default of OpenVINO is disabled to keep the build minimal.
CMAKE_ARGS="-DEXECUTORCH_BUILD_KERNELS_OPTIMIZED=OFF \
-DEXECUTORCH_BUILD_KERNELS_LLM=OFF \
-DEXECUTORCH_BUILD_KERNELS_LLM_AOT=OFF \
-DEXECUTORCH_BUILD_EXTENSION_LLM=OFF \
-DEXECUTORCH_BUILD_EXTENSION_LLM_RUNNER=OFF \
-DEXECUTORCH_BUILD_OPENVINO=OFF" \
CMAKE_BUILD_PARALLEL_LEVEL="${CMAKE_BUILD_PARALLEL_LEVEL:-4}" \
"$VENV/bin/pip" install --no-deps --no-build-isolation ./executorch
"$VENV/bin/python" repro.py
repro.py
"""Minimal reproducer: slice_copy with stride != 1 never delegates to XNNPACK.
op-support.csv lists slice_copy (fp32) with constraints "no zero-dim tensors,
no dynamic shapes". Both modules below are static-shape float32 slices with no
zero dims; they differ ONLY in stride. Expected per the table: both delegate.
Observed: stride==1 delegates (1 call), stride==2 does not (0 calls).
Prints GAP_REPRODUCED when the gap is present, GAP_NOT_REPRODUCED otherwise.
"""
import torch
from torch import nn
from executorch.backends.xnnpack.partition.xnnpack_partitioner import XnnpackPartitioner
from executorch.backends.xnnpack.utils.configs import get_transform_passes
from executorch.exir import to_edge_transform_and_lower
def delegate_calls(mod, shape):
torch.manual_seed(0)
edge = to_edge_transform_and_lower(
torch.export.export(mod.eval(), (torch.randn(*shape),)),
partitioner=[XnnpackPartitioner()],
transform_passes=get_transform_passes(),
)
gm = edge.exported_program().graph_module
n = 0
leftover = []
for node in gm.graph.nodes:
if node.op == "call_function" and "executorch_call_delegate" in str(node.target):
n += 1
elif node.op == "call_function":
leftover.append(str(node.target).split(":")[0][:80])
return n, leftover
class SliceStep1(nn.Module): # control: satisfies every stated table constraint
def forward(self, x):
return x[:, 2:, :, :]
class SliceStep2(nn.Module): # repro: same, but stride 2 (as in Swin PatchMerging)
def forward(self, x):
return x[:, 0:28:2, :, :]
def main():
print("torch:", torch.__version__)
from importlib.metadata import version
print("executorch:", version("executorch"))
shape = (1, 28, 28, 96) # Swin-T stage-1 PatchMerging feature shape
c_calls, c_left = delegate_calls(SliceStep1(), shape)
r_calls, r_left = delegate_calls(SliceStep2(), shape)
print(f"control stride=1: delegate_calls={c_calls} leftover={c_left}")
print(f"repro stride=2: delegate_calls={r_calls} leftover={r_left}")
if c_calls == 1 and r_calls == 0:
print("GAP_REPRODUCED")
else:
print("GAP_NOT_REPRODUCED")
if __name__ == "__main__":
main()
requirements.txt
# Pinned public dependencies for the slice_copy stride reproducer.
# Install with: pip install -r requirements.txt (see repro.sh; torch/torchao
# come from the CPU index, the rest from PyPI). ExecuTorch itself is built
# from pinned public source by repro.sh (no released wheel is assumed).
--index-url https://download.pytorch.org/whl/cpu
--extra-index-url https://pypi.org/simple
torch==2.14.0
torchao==0.18.0
flatbuffers
pyyaml
ruamel.yaml
tabulate
expecttest
parameterized
py-cpuinfo
requests
pytorch-tokenizers
pandas
hydra-core
omegaconf
observed.log
# Observed evidence: slice_copy stride != 1 never delegates (docs gap)
# Environment: ExecuTorch 1.6.0+6c8bc89 built from pinned source (pybind
# preset, Release, XNNPACK ON; KERNELS_OPTIMIZED/LLM/LLM-AOT/EXT-LLM/
# LLM-RUNNER/OPENVINO OFF; -j4), torch 2.14.0+cpu, torchvision 0.29.0+cpu,
# torchao 0.18.0+cpu, XNNPACK 92a7ad50, gcc 11.5, cmake 3.31.8,
# AMD EPYC-Genoa x86_64. Weights: public swin_t-704ceda3.pth.
## 1. Upstream model: torchvision swin_t FP32, documented flow
# torch.export (1014 nodes) -> to_edge_transform_and_lower(XnnpackPartitioner,
# get_transform_passes()) -> .pte (113408416 bytes, contains XnnpackBackend)
delegate calls: 66, lowered modules: 66
# All 24 leftover float32 slice_copy nodes are strided (PatchMerging x[::2]):
# aten_slice_copy_tensor_27 -> size=(1, 28, 56, 96) args: [.., dim=1, 0, MAXINT, step=2]
# aten_slice_copy_tensor_28 -> size=(1, 28, 28, 96) args: [.., dim=2, 0, MAXINT, step=2]
# (... 22 more, same step=2 pattern at every stage transition ...)
# Partitioner debug log gives NO reason for these slices (silent reject);
# the 20 leftover add.Tensor are int64 index math ("invalid dtype", consistent
# with the table's fp16/fp32 compute dtypes).
## 2. Correctness (documented tolerance rtol=1e-3, atol=1e-5), 3 seeded inputs
seed=0 xnn_allclose=true xnn_max_abs=5.72e-06 top1_match=true
seed=1 xnn_allclose=true xnn_max_abs=7.63e-06 top1_match=true
seed=2 xnn_allclose=true xnn_max_abs=4.29e-06 top1_match=true
# Portable-kernel control .pte also allclose=true on all 3 inputs.
## 3. Minimal reproducer (repro.py): only stride differs, both static fp32
control stride=1: delegate_calls=1 leftover=['<built-in function getitem>']
repro stride=2: delegate_calls=0 leftover=['<EdgeOpOverload aten.slice_copy>']
GAP_REPRODUCED
# Open-ended step-1 partial slice also delegates (1 call): stride is the gate.
# Source: backends/xnnpack/partition/config/generic_node_configs.py,
# SliceCopyConfig.check_constraints: "Support slicing with stride = 1 ..."
# returns False for stride != 1 with no diagnostic; the published table lists
# only "no zero-dim tensors, no dynamic shapes".
## 4. Performance (1 thread, pinned CPU, warmup 20, 3 alternating x10 batches)
# XNNPACK medians: 129.27 / 128.88 / 129.85 ms (CV<=0.0032)
# Portable medians: 7062.49 / 7072.42 / 7068.93 ms (CV<=0.0041)
# Every XNNPACK sample below every control sample (~55x faster, expected
# direction); max load/CPU 0.045. No performance anomaly.
verify-1.log
# Verification run 1: fresh clean-environment execution of submitted repro
# Date (UTC): 2026-09-29. Host: Linux x86_64, 1 logical CPU, gcc 11.5.
# NOTE: host background load was elevated (loadavg ~3.7 at start on 1 CPU).
# This finding is about deterministic partitioner behavior (delegate counts),
# not timing, so load does not affect the outcome; both runs agree exactly.
## Command (fresh directory <workdir>/verify1, copies of submitted files)
# cd verify1 && bash repro.sh
# Exit: 0. Wall: 6m22s (includes venv + pip install + git clone + source build).
# Build jobs: default from repro.sh (CMAKE_BUILD_PARALLEL_LEVEL=4 max). No ccache.
## Clean-setup steps observed in full log (private paths redacted here)
- python3 -m venv repro_env; pip install -r requirements.txt (pinned wheels)
- cmake version 3.31.8
- git clone https://github.com/pytorch/executorch.git executorch
- checked out commit 6c8bc89be41b3fe3fb42ce401556ecad65df7919 (matches sources.json pin)
- submodule checkouts: XNNPACK 92a7ad50, cpuinfo f9a03241, pthreadpool a56dcd79,
FP16 4dfe081, FXdiv b408327, flatbuffers 595bf000, flatcc 896db547,
gflags a738fdf, json ac0133ea, pybind11 d03662f0
(all match the candidate's stated environment string)
- pip install --no-deps --no-build-isolation ./executorch -> wheel
executorch-1.6.0+6c8bc89 built and installed successfully
## Experiment output (repro.py, run by repro.sh in the fresh env)
torch: 2.14.0+cpu
executorch: 1.6.0+6c8bc89
control stride=1: delegate_calls=1 leftover=['<built-in function getitem>']
repro stride=2: delegate_calls=0 leftover=['<EdgeOpOverload']
GAP_REPRODUCED
## Verifier cross-checks (this run)
- Fresh checkout HEAD verified: 6c8bc89be41b3fe3fb42ce401556ecad65df7919
- Fresh venv torch verified: 2.14.0+cpu
- Result matches candidate's claimed 1-vs-0 delegation and GAP_REPRODUCED.
verify-2.log
# Verification run 2: re-execution of the experiment in a separate process
# Date (UTC): 2026-09-29. Reused run-1 env only; experiment re-run from scratch.
# Host load at start: ~6.1 on 1 CPU (background contention; outcome is
# deterministic partitioning, identical to run 1).
## Command
# cd verify1 && REPRO_VENV=<workdir>/verify1/repro_env bash repro.sh
# Exit: 0. Wall: 9.5s (experiment only, no rebuild). Separate process from run 1.
## Experiment output
reuse mode: using <workdir>/verify1/repro_env
torch: 2.14.0+cpu
executorch: 1.6.0+6c8bc89
control stride=1: delegate_calls=1 leftover=['<built-in function getitem>']
repro stride=2: delegate_calls=0 leftover=['<EdgeOpOverload']
GAP_REPRODUCED
## Independent verifier probes (separate processes, torch threads=1)
Probe A (fairness + backend identity + generality), same pinned toolchain:
- pre-partition aten op identical for ALL variants: aten.slice.Tensor
step1-implicit args (dim=1,start=2,end=MAX) -> 1 delegate call
step1-explicit args (dim=1,start=2,end=28) -> 1 delegate call
step2 args (dim=1,start=0,end=28,step=2) -> 0 delegate calls
step3 args (dim=1,start=0,end=28,step=3) -> 0 delegate calls
step2-on-dim2 args (dim=2,start=0,end=28,step=2) -> 0 delegate calls
- delegate call backend_id = XnnpackBackend (real XNNPACK delegation, not just a count)
- inner op of the step-1 delegate call: aten.slice_copy.Tensor
- leftover in step-2 case: the identical aten.slice_copy.Tensor edge op
- conclusion: control is fair (same op pre-partition, real XNNPACK backend);
gap generalizes to step 3 and to other dims, not just the reported step-2 case.
(An initial full-range 0:28:1 probe variant folded to aten.alias at export and
was discarded as a no-op slice; replaced with partial 2:28:1 which delegates.)
## Document/source cross-checks (pinned sources)
- op-support.csv line 42 verbatim: slice_copy,"fp16, fp32",,"no zero-dim tensors, no dynamic shapes"
- xnnpack-partitioner.rst verbatim: default XnnpackPartitioner "will run as much
of the model as possible on the XNNPACK backend"
- No stride constraint for slice documented anywhere under
docs/executorch/docs/source/backends/xnnpack/; the table DOES document a stride
constraint for max_pool2d ("stride <= kernel_size"), showing such limits are
normally recorded there.
- executorch/backends/xnnpack/partition/config/generic_node_configs.py,
SliceCopyConfig.check_constraints (pinned commit 6c8bc89): returns False for
stride != 1 with NO why() diagnostic, while every other rejection in the same
function logs a reason -> candidate's "silent reject" claim holds.
- Model provenance: installed torchvision 0.29.0+cpu swin_transformer.py is
byte-identical (sha256 9ec73f23...) to vision/ at pinned commit 9769007.
## Scope note
- The runnable reproduction covers the documentation gap (partitioning); it was
executed twice above. The swin_t numerics/perf figures in observed.log are
supporting context, not the finding (kind=documentation, performance=null, no
perf anomaly claimed), and were corroborated at source/doc level rather than
re-benchmarked. No timing thresholds apply to this verdict.
Independent reproduction
Tried to disprove the docs-gap claim and could not. Ran the submitted repro twice in separate processes: (1) fresh clean-env run in a new dir (venv, pinned wheels, clone + checkout 6c8bc89, submodule pins all matching, source build of executorch-1.6.0+6c8bc89, 6m22s, exit 0) printed control 1 vs repro 0 delegate calls and GAP_REPRODUCED; (2) reuse-mode re-run (9.5s, exit 0) printed identical counts and GAP_REPRODUCED. Independent probe (torch threads=1) showed the control is fair: pre-partition aten op is aten.slice.Tensor for all variants, the step-1 delegate call has backend_id=XnnpackBackend containing aten.slice_copy.Tensor while the identical edge slice_copy is left over for step 2, and the gap generalizes to step 3 and other dims. Source check at pinned commit: SliceCopyConfig.check_constraints returns False for stride!=1 with no why() diagnostic (silent, unlike sibling paths). Doc check: op-support.csv line 42 lists only 'no zero-dim tensors, no dynamic shapes', partitioner doc promises default delegation of 'as much as possible', no stride limit for slice appears anywhere in the xnnpack docs while max_pool2d's stride limit IS tabulated. Model provenance holds: installed torchvision 0.29.0 swin_transformer.py byte-identical to pinned vision commit 9769007. Repro artifacts use only public URLs/relative paths; full build logs with private paths kept out of out/. Perf/numerics in observed.log are supporting context (kind=documentation, no anomaly claimed), not re-benchmarked.
cc @GregoryComer @digantdesai @cbilgin
Expected behavior
op-support.csv lists slice_copy for fp32 with constraints 'no zero-dim tensors, no dynamic shapes'. A static-shape float32 slice with step 2 and no zero dims satisfies every stated constraint, so with a default XnnpackPartitioner (documented to run as much as possible on XNNPACK) it should delegate exactly like the step-1 slice.
Actual behavior
Static float32 slice with step 2 (shape (1,28,28,96), explicit end, no zero dims) yields 0 delegate calls; the leftover aten.slice_copy.Tensor falls back to portable kernels. The identical slice with step 1 (open or explicit end) yields 1 delegate call. Upstream swin_t shows the same pattern: all 24 non-delegated slices are PatchMerging step-2 slices; the partitioner debug log gives no reason for them (the stride check returns False silently), while the 20 leftover int64 adds are correctly rejected as invalid dtype. Numerics are unaffected: XNNPACK .pte matches eager with allclose rtol=1e-3 atol=1e-5 True on 3 seeded inputs (max abs <= 7.7e-06, top-1 agrees), same as the portable control. Perf at 1 thread (warmup 20, 3 alternating x10 batches, pinned CPU, threadpool verified 1): XNNPACK medians 129.27/128.88/129.85 ms vs portable 7062.49/7072.42/7068.93 ms, every XNNPACK sample below every control sample, max load/CPU 0.045; expected direction, no anomaly.
Model and versions
Model: https://github.com/pytorch/vision/blob/9769007f09193b307a90d2b833aac31d364f03e3/torchvision/models/swin_transformer.py
ExecuTorch 1.6.0+6c8bc89 from pinned public source (pybind preset, Release, XNNPACK ON; KERNELS_OPTIMIZED/LLM/LLM-AOT/EXT-LLM/LLM-RUNNER/OPENVINO OFF; -j4), torch 2.14.0+cpu, torchvision 0.29.0+cpu (swin_transformer.py sha256 9ec73f23 byte-identical to pinned vision commit), torchao 0.18.0+cpu, submodules XNNPACK 92a7ad50 cpuinfo f9a03241 pthreadpool a56dcd79 FP16 4dfe081c FXdiv b408327a flatbuffers 595bf000 flatcc 896db547 gflags a738fdf9 json ac0133ea pybind11 d03662f0, gcc 11.5, cmake 3.31.8, AMD EPYC-Genoa x86_64
Documentation checked
https://raw.githubusercontent.com/pytorch/executorch/6c8bc89be41b3fe3fb42ce401556ecad65df7919/docs/source/backends/xnnpack/op-support.csv
https://raw.githubusercontent.com/pytorch/executorch/6c8bc89be41b3fe3fb42ce401556ecad65df7919/docs/source/backends/xnnpack/xnnpack-partitioner.rst
Reproduction
Save repro.sh, repro.py, requirements.txt into one fresh directory on a Linux x86_64 CPU host with python3 (3.10-3.14), git, gcc, cmake; run bash repro.sh (creates repro_env, installs pinned public wheels, clones pinned public sources, builds ExecuTorch from source, prints GAP_REPRODUCED). Repeat without redownloading/rebuilding via REPRO_VENV=
/repro_env bash repro.sh.repro.sh
repro.py
requirements.txt
observed.log
verify-1.log
verify-2.log
Independent reproduction
Tried to disprove the docs-gap claim and could not. Ran the submitted repro twice in separate processes: (1) fresh clean-env run in a new dir (venv, pinned wheels, clone + checkout 6c8bc89, submodule pins all matching, source build of executorch-1.6.0+6c8bc89, 6m22s, exit 0) printed control 1 vs repro 0 delegate calls and GAP_REPRODUCED; (2) reuse-mode re-run (9.5s, exit 0) printed identical counts and GAP_REPRODUCED. Independent probe (torch threads=1) showed the control is fair: pre-partition aten op is aten.slice.Tensor for all variants, the step-1 delegate call has backend_id=XnnpackBackend containing aten.slice_copy.Tensor while the identical edge slice_copy is left over for step 2, and the gap generalizes to step 3 and other dims. Source check at pinned commit: SliceCopyConfig.check_constraints returns False for stride!=1 with no why() diagnostic (silent, unlike sibling paths). Doc check: op-support.csv line 42 lists only 'no zero-dim tensors, no dynamic shapes', partitioner doc promises default delegation of 'as much as possible', no stride limit for slice appears anywhere in the xnnpack docs while max_pool2d's stride limit IS tabulated. Model provenance holds: installed torchvision 0.29.0 swin_transformer.py byte-identical to pinned vision commit 9769007. Repro artifacts use only public URLs/relative paths; full build logs with private paths kept out of out/. Perf/numerics in observed.log are supporting context (kind=documentation, no anomaly claimed), not re-benchmarked.
cc @GregoryComer @digantdesai @cbilgin