Repository navigation
lowering to Ethos-U85 with input of HW1 #18047
Replies: 1 comment
|
Hi Shmulik, the 16× increase is consistent with NHCWB16 channel padding: it stores channels in groups of 16, so a single-channel, 8-bit tensor can occupy One experiment is to use channels-last memory format throughout export. For a standard PyTorch model = model.eval().to(memory_format=torch.channels_last)
example_input = torch.randn(1, 1, 300, 400).to(
memory_format=torch.channels_last
)For your ExecuTorch documents channels-last as a way to avoid boundary transposes, but currently describes this path as experimental. Keep calibration/export inputs and the runtime tensor layout consistent, then check whether the initial transpose and its allocation disappear, and validate the outputs. This is a diagnostic step, not a guaranteed fix for Vela’s internal format choices. ExecuTorch guidance. Could you share your PyTorch, ExecuTorch and Vela versions, plus a minimal model/export script and compile options? That would help identify why this full-resolution intermediate is being materialized. |
Uh oh!
There was an error while loading. Please reload this page.
Hello all,
I'm trying to lower a tiny model so it can run in a very memory-tight environment. input image is [300,400,1] of size. when looking in the tensor allocation table vela reports, I see the following lines right the the table top:
0 - 1: 0x1d4c00 - 0x1f20c0: 120000: 2040512: NHWC : quantized_decomposed_quantize_per_tensor_default
0 - 3: 0x0 - 0x1d4c00: 1920000: 2400944: NHCWB16 : tosa_transpose_default
as 300x400x16 = 1920000, I'm obviously not handling the NCHW->NHWC correctly, and the explicit transpose creates that huge (in my world of scarce memory) tensor.
the peak memory for that model, if not that ill behavior, is lower than 600K.
I'll be very grateful if someone who already solved this will share.
Shmulik
All reactions