Skip to content

WebAssembly / Web runtime (both for wasm-simd and WebGPU) #3497

Description

@vadimkantorov

I'm wondering if ExecuTorch can be compiled for WebAssembly target? As far as I understand, XNNPACK exists for wasm-simd, so theoretically at least for CPU it can be done? (e.g. to be compared with tflite+tfjs, ort-web and tvm-wasm at least for some popular models like MobileNets)

(This is especially interesting if strong fusion/codegen can be done to produce fused wasm-simd code/fused WebGPU programs - although maybe this is an ask for Inductor)

Activity

  1. SS-JIA commented on May 3, 2024

    @SS-JIA
    Contributor

    cc: @mcr229 or @digantdesai regarding running XNNPACK via wasm

  2. SS-JIA commented on May 3, 2024

    @SS-JIA
    Contributor

    Also cc: @mergennachin

  3. JacobSzwejbka commented on May 9, 2024

    @JacobSzwejbka
    Contributor

    I've talked with @digantdesai about this before. I think for xnnpack he mentioned it should just be plug and play. Ive been wanting to try out wasm for sometime now just havent had the bandwidth.

  4. added
    enhancementNot as big of a feature, but technically not a bug. Should be easy to fix
    on May 9, 2024
  5. vadimkantorov commented on May 9, 2024

    @vadimkantorov
    Author

    I also wonder about the fusion capabilities of executorch :) Does it allow Inductor codegen'd fused kernels (e.g. think quant/dequant fused into the flash attn kernel directly, with positional embedding computation also fused into this kernel)?

    Another interesting backend is webgpu/wgpu: https://github.com/huggingface/ratchet or even directly wgpu/wgsl shaders could in theory be a compilation target for fused kernels

    But even if executorch does not support wild codegen/fusions - it's still be good to have it as a baseline with comparisons against ort-web and tflate-tfjs and tvm-wasm and ggml compiled to wasm. This should show roughly where all these frameworks stand (especially if compiling is relatively doable)

  6. vadimkantorov commented on May 9, 2024

    @vadimkantorov
    Author

    And given that currently PyTorch does not have its own inference wasm/WebGPU story, having executorch compiled to wasm-simd might be a nice baseline to have (especially if it's minimalistic and relatively simple to compile)

  7. kimishpatel commented on Nov 21, 2024

    @kimishpatel
    Contributor

    I suspect much of the core should compilable with emscripten cpp compiler. Probably not optimized operators though and not too sure about backends/xnnpack

  8. vadimkantorov commented on Nov 21, 2024

    @vadimkantorov
    Author

    Maybe best would be adding some sort of GitHub Actions CI test compiling it with emscripten... (even if no tests using it exist so far)

  9. digantdesai commented on Nov 21, 2024

    @digantdesai
    Contributor

    not too sure about backends/xnnpack

    It should be, given a bunch of WASM[SIMD] kernels. I haven't tried it myself though. IIRC there aren't any CI for that on github/xnnpack either.

  10. vadimkantorov commented on Nov 21, 2024

    @vadimkantorov
    Author

    xnnpack is also known to compile (and maybe even tested) for wasm/simd, so somehow this should be achievable... don't know if any compact backend library/project exists for webgpu kernels

  11. locked and limited conversation to collaborators on Feb 5, 2025
  12. converted this issue into a discussion #8216 on Feb 5, 2025
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

enhancementNot as big of a feature, but technically not a bug. Should be easy to fix

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions