From 80108f5101f90aaef6578eadb41a040ce4d4fdfa Mon Sep 17 00:00:00 2001 From: Matt Cossins Date: Mon, 6 Jul 2026 10:49:45 +0100 Subject: [PATCH 1/4] First draft MLIA LP --- .../1-overview.md | 83 +++++++ .../2-install-and-discover.md | 187 +++++++++++++++ .../3-tflite-comparison.md | 212 ++++++++++++++++++ .../4-analyze-tosa-with-vela.md | 69 ++++++ .../5-executorch-pt2-pte.md | 179 +++++++++++++++ .../6-python-api.md | 95 ++++++++ .../_index.md | 74 ++++++ .../_next-steps.md | 9 + 8 files changed, 908 insertions(+) create mode 100644 content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/1-overview.md create mode 100644 content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/2-install-and-discover.md create mode 100644 content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/3-tflite-comparison.md create mode 100644 content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/4-analyze-tosa-with-vela.md create mode 100644 content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/5-executorch-pt2-pte.md create mode 100644 content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/6-python-api.md create mode 100644 content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/_index.md create mode 100644 content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/_next-steps.md diff --git a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/1-overview.md b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/1-overview.md new file mode 100644 index 0000000000..91a884b241 --- /dev/null +++ b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/1-overview.md @@ -0,0 +1,83 @@ +--- +title: What is the ML Inference Advisor? + +weight: 2 + +### FIXED, DO NOT MODIFY +layout: "learningpathall" +--- + +## Understand what MLIA does + +Arm ML Inference Advisor (MLIA) helps you evaluate whether a machine learning model is suitable for a target inference platform. + +In this Learning Path, you use MLIA from the command line to check model compatibility, estimate performance, and read advice that points toward useful model changes. These examples use Arm Ethos-U as an example target. + +MLIA is most useful before full deployment or runtime profiling, when you are asking questions such as: + +- Will this model map cleanly to my target? +- Which target profile should I use for early analysis? +- Which operators or layers are likely to matter most for performance? +- Is the model compute-bound, memory-bound, or affected by low MAC utilization? +- What should I investigate before building firmware or running on a board? + +MLIA does not make the final optimization decision for you. It gives target-aware evidence so you can decide what to change, what to measure next, and which workflow stage deserves attention. + +## Use the CLI first + +The MLIA CLI is the primary workflow in this Learning Path. You will use it to: + +- discover installed targets, target profiles, and backends +- run compatibility checks +- run performance analysis +- request JSON output +- inspect advice and metrics + +After you understand the CLI workflow, you will briefly use the Python API. The API is useful when you want to embed MLIA results in another product, dashboard, CI job, or tool. + +## Using MLIA alongside other tools + +MLIA is not a replacement for graph visualization or runtime profiling. It is an advisory layer that helps earlier in the model preparation workflow. + +| Tool | Use it to answer | +| --- | --- | +| MLIA | Is this model suitable for my target, and what should I change? | +| Model Explorer | What does the generated model artifact graph look like? | +| Vela | How does the Ethos-U compiler map supported work onto the NPU? | +| Runtime-specific profiling tools | What happened when the model actually ran? For example, use ETRecord, ETDump, and ExecuTorch Inspector for ExecuTorch deployments, or TensorFlow Lite benchmark and profiling tools for TFLite deployments. | + +For example, MLIA can tell you which layers dominate estimated cycles or have low MAC utilization. Model Explorer can show how an ExecuTorch `.pte` artifact is partitioned into delegate regions. Runtime-specific profiling tools can show behavior after you have a runnable deployment. + +Use these tools together: + +- Use MLIA before or during model preparation. +- Use Model Explorer to inspect generated artifacts and delegation structure. +- Use runtime profiling tools after you can execute the model. + +## Understand the plugin model + +MLIA uses a plugin model. The `mlia` core package provides the shared command-line interface, output structure, and Python API. Target, backend, and converter support is added through plugins. + +The important repositories are: + +- `arm/mlia`: core MLIA package +- `arm/mlia-ethos-u`: Ethos-U target plugin and Vela/Corstone backend plugins +- `arm/mlia-converters-pytorch`: PyTorch `.pt2` converter plugin for TOSA and PTE routes +- `arm/mlia-legacy`: legacy support for older MLIA flows + +## Understand the model formats + +MLIA can analyze different kinds of model artifacts depending on what workflow you are using and the stage you want to analyze. + +| Format | Where it fits | +| --- | --- | +| `.pt2` | PyTorch exported program input for PyTorch and ExecuTorch-oriented workflows. | +| `.tosa` | Intermediate representation consumed by compiler/backend flows such as Ethos-U Vela. | +| `.pte` | Serialized ExecuTorch program. Ethos-U `.pte` performance analysis uses Corstone backends. | +| `.tflite` | TensorFlow Lite model format used in many Ethos-U and embedded ML workflows. | + +## What you have learned + +You have learned what MLIA does, what you use it for, and how MLIA fits alongside other tools. + +Next, you will install MLIA and inspect the capabilities available in your environment. diff --git a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/2-install-and-discover.md b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/2-install-and-discover.md new file mode 100644 index 0000000000..631d83bf42 --- /dev/null +++ b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/2-install-and-discover.md @@ -0,0 +1,187 @@ +--- +title: Install MLIA and discover capabilities + +weight: 3 + +### FIXED, DO NOT MODIFY +layout: "learningpathall" +--- + +## Check your environment + +Use Ubuntu 22.04 LTS or another compatible Linux environment with Python 3.10 or later. + +Some Python environments also require the Python development package, such as `libpython3.10-dev`, before installing MLIA packages. + +Check that Git LFS is installed: + +```bash +git lfs version +``` + +If this command fails, install Git LFS: + +```bash +sudo apt update +sudo apt install -y git-lfs +``` + +## Create a Python environment + +Create a virtual environment so the MLIA packages do not conflict with any existing PyTorch, ExecuTorch, or TensorFlow environment. + +```bash +python3 -m venv mlia_env +source mlia_env/bin/activate +python -m pip install --upgrade pip +``` + +## Install MLIA + +{{% notice TODO %}} +Confirm installation +{{% /notice %}} + +MLIA uses plugins. The examples in this Learning Path use Ethos-U as the target, so install the Ethos-U plugin package: + +```bash +pip install mlia-ethos-u +``` + +The Ethos-U plugin package depends on a compatible MLIA core package. Installing the target plugin is the recommended starting point because it brings in the matching MLIA core dependency. + +## Confirm the CLI works + +Display top-level help: + +```bash +mlia --help +``` + +You should see commands similar to: + +```output + check Generate compatibility/performance advice for a model + backend Manage MLIA backends + target Manage MLIA targets +``` + +The `mlia check` command is the main command you will use to ask MLIA compatibility and performance questions about model artifacts. + +## Discover target profiles + +MLIA target profiles describe the target configuration used for analysis. List the target profiles available in your environment: + +```bash +mlia target list +``` + +For Ethos-U, typical bundled profiles include: + +```output +ethos-u55-128 +ethos-u55-256 +ethos-u65-256 +ethos-u65-512 +ethos-u85-128 +ethos-u85-256 +ethos-u85-512 +ethos-u85-1024 +ethos-u85-2048 +``` + +In this Learning Path, the examples use one Ethos-U85 profile: + +```output +ethos-u85-256 +``` + +Use a different profile if you want MLIA to evaluate the same model for a different Ethos-U configuration. + +## Discover backends + +Backends perform the work behind an MLIA analysis flow. List available and installed backends: + +```bash +mlia backend list +``` + +For this Ethos-U demonstration, you should expect Vela and Corstone backend options. Vela is used for compiler-oriented compatibility and performance analysis. Corstone backends are used for simulation-oriented performance flows, including supported ExecuTorch `.pte` workloads. + +```output +Name Installed Installable +corstone-300 no yes +corstone-310 no yes +corstone-320 no yes +vela no yes +``` + +Install Vela: + +```bash +mlia backend install vela +``` + +Check the backend list again: + +```bash +mlia backend list +``` + +You should now see `vela` in the installed backend list. + +## Clone model artifacts + +This Learning Path uses prebuilt artifacts from the Arm ML model artifacts repository: + +```bash +git lfs install +git clone --filter=blob:none --sparse https://github.com/arm-education/ml-model-artifacts.git +cd ml-model-artifacts +git sparse-checkout set pt2 pte tflite tosa +git lfs pull \ + --include="pt2/toy_conditional_select_fp32.pt2,pte/toy_conditional_select_int8_ethos_u55_256.pte,pte/toy_conditional_select_int8_ethos_u85_256.pte,tflite/mv2_fp32.tflite,tflite/mv2_int8.tflite,tosa/mv2_fp32.tosa,tosa/mv2_int8.tosa" \ + --exclude="" +git lfs checkout +``` + +This downloads only the artifacts used in this Learning Path. It avoids larger unrelated files, such as transformer `.pte`, `.etdp`, and `.etrecord` artifacts. + +Confirm that the artifacts are real model files, not Git LFS pointer files: + +```bash +wc -c tflite/mv2_int8.tflite +``` + +You should see a size of several megabytes, similar to: + +```output +3942808 tflite/mv2_int8.tflite +``` + +If the file is about 100 to 200 bytes, it is still a Git LFS pointer file. Run the `git lfs pull` command again from the `ml-model-artifacts` directory, then rerun the size check. + +The model artifacts are provided for learning and analysis exercises. Use them to explore MLIA workflows, model formats, and target-aware advice, not as accuracy reference models. + +The repository contains model artifacts such as: + +```output +ml-model-artifacts/ +├── pte/ +│ ├── toy_conditional_select_int8_ethos_u55_256.pte +│ └── toy_conditional_select_int8_ethos_u85_256.pte +├── pt2/ +│ └── toy_conditional_select_fp32.pt2 +├── tflite/ +│ ├── mv2_fp32.tflite +│ └── mv2_int8.tflite +└── tosa/ + ├── mv2_fp32.tosa + └── mv2_int8.tosa +``` + +## What you have learned + +You have installed MLIA, along with the Ethos-U plugin, discovered available target profiles and backends from the CLI, installed Vela, and cloned model artifacts for analysis. + +Next, you will run your first MLIA compatibility and performance checks. diff --git a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/3-tflite-comparison.md b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/3-tflite-comparison.md new file mode 100644 index 0000000000..82c7b1a1d8 --- /dev/null +++ b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/3-tflite-comparison.md @@ -0,0 +1,212 @@ +--- +title: Analyze TensorFlow Lite artifacts with Vela + +weight: 4 + +### FIXED, DO NOT MODIFY +layout: "learningpathall" +--- + +## Run a compatibility check + +Start by asking MLIA whether a TensorFlow Lite model can map to the selected target profile. + +TensorFlow Lite is a compact model format and runtime stack for deploying machine learning models on mobile, embedded, and edge devices. In embedded ML workflows, a `.tflite` file is often the artifact handed to backend tools for target-specific compatibility checks, compilation, or runtime deployment. + +The model artifacts repository includes floating-point and quantized MobileNetV2 TensorFlow Lite variants: + +```output +tflite/mv2_fp32.tflite +tflite/mv2_int8.tflite +``` + +From the `ml-model-artifacts` directory, run MLIA on the FP32 TFLite artifact: + +```bash +mlia check tflite/mv2_fp32.tflite \ + --target-profile ethos-u85-256 \ + --compatibility \ + --backend vela \ + --json +``` + +The `--json` option makes the output easier to inspect, compare, and automate. + +You can also use the shorter `-t` and `-b` options instead of `--target-profile` and `--backend`. The `-t` option still means target profile: + +```bash +mlia check tflite/mv2_fp32.tflite -t ethos-u85-256 --compatibility -b vela --json +``` + +This report should show the FP32 TFLite model as incompatible for Ethos-U acceleration, with `accelerator_operator_percentage` set to `0`. This is expected because Ethos-U acceleration requires supported quantized integer workloads. The failed checks explain that the input, output, and weight tensors are missing quantization parameters. + +This does not mean TensorFlow Lite is the problem. It means this particular TFLite artifact is not in the numeric form the selected Ethos-U target needs. Use a quantized model instead. + +Now run the same compatibility check on the quantized INT8 model: + +```bash +mlia check tflite/mv2_int8.tflite \ + --target-profile ethos-u85-256 \ + --compatibility \ + --backend vela \ + --json +``` + +You will generate a lengthy report. A small snippet is included below: + +```output +{ + "target": { + "configuration": { + "target": "ethos-u85", + "mac": 256 + } + }, + "model": { + "name": "mv2_int8.tflite", + "format": "tflite" + }, + "backends": [ + { + "id": "vela", + "name": "Vela Compiler", + "version": "5.0.0" + } + ], + "results": [ + { + "kind": "compatibility", + "status": "ok", + "metrics": [ + { + "name": "accelerator_operator_percentage", + "value": 100.0 + } + ] + } + ] +} +``` + +## What does this report tell us? + +| Field | What it tells you | +| --- | --- | +| `schema_version`, `run_id`, `timestamp` | Which output schema was used, and how to identify this specific run later. | +| `tool` | The MLIA version that generated the report. | +| `target` | The selected target profile and its configuration, such as Ethos-U85 with 256 MACs. | +| `model` | The artifact name, format, and hash. | +| `context` | The CLI command that produced the report. | +| `backends` | The backend MLIA used, including the Vela version and compiler configuration. | +| `results` | The answer MLIA produced for the requested check. | +| `checks` | The individual operator support checks. | +| `entities` | The operators MLIA analyzed, including placement and operator type. | + +For the INT8 TFLite file, the important result is that `status` is `ok` and `accelerator_operator_percentage` is `100.0`. That means Vela found the operators in this quantized MobileNetV2 TFLite artifact compatible with the selected `ethos-u85-256` target profile, and MLIA expects the operator work to map to the NPU path for this compatibility check. It does not prove runtime latency or application accuracy. It tells you the model is a good candidate for the next step: performance estimation and deeper deployment testing. + +## Read operator placement + +The summary result tells you that the model is compatible overall. To see how MLIA reached that result, look at the `entities` list. Each operator entity describes one analyzed operator and includes a `placement` field: + +```json +{ + "scope": "operator", + "name": "...", + "placement": "npu", + "attributes": { + "op_type": "Conv2D" + } +} +``` + +Names and attributes vary by input format and MLIA version. The important part is the placement: it tells you where MLIA and the backend analysis expect the operator to land for this target profile. If a future model has unsupported operators, this is where you start narrowing down the problem: find the operator whose placement or check status differs from the expected NPU path, then inspect that part of the model graph or change the model before deployment. + +## Run a performance check + +Next, ask MLIA for a target-aware performance estimate using the INT8 TFLite model: + +```bash +mlia check tflite/mv2_int8.tflite \ + --target-profile ethos-u85-256 \ + --performance \ + --backend vela \ + --json +``` + +The performance report uses the same top-level structure as the compatibility report, but the `results` object now contains estimated performance metrics, operator-level breakdowns, and advice: + +```output +{ + "results": [ + { + "kind": "performance", + "status": "ok", + "warnings": [ + "The performance figures above refer to NPU only" + ], + "metrics": [ + { + "name": "npu_cycles", + "value": 3623146 + }, + { + "name": "total_cycles", + "value": 5006357 + }, + { + "name": "inference_time", + "unit": "ms", + "value": 5.006357 + }, + { + "name": "inferences_per_second", + "unit": "inferences/s", + "value": 199.74604288108097 + }, + { + "name": "target_utilization", + "unit": "%", + "value": 72.3709076280417 + } + ], + "advice": [ + { + "category": "performance", + "severity": "warning", + "message": "The following layers make up the majority of operator cycles..." + }, + { + "category": "performance", + "severity": "warning", + "message": "Among the layers with the highest impact, 5 layers have been identified with low MAC utilization..." + }, + { + "category": "performance", + "severity": "warning", + "message": "Among the layers with the highest impact, 5 layers have been identified as possibly memory bound..." + } + ] + } + ] +} +``` + +## What does this report tell us? + +| Field | What it tells you | +| --- | --- | +| `warnings` | Important scope limits for the result, such as the estimate referring to NPU work only. | +| `metrics` | Summary estimates for cycles, inference time, throughput, utilization, model size, and memory use. | +| `breakdowns` | Per-operator metrics, including operator cycles, memory access cycles, MAC count, and MAC utilization. | +| `advice` | MLIA's interpretation of the metrics, including which layers dominate cycles or may be inefficient. | +| `availability` and `reason` | Why a metric is not available from the selected backend, if MLIA cannot report it. | + +For this INT8 TFLite file, the Vela-backed estimate reports about `5.01M` total cycles, about `5.01 ms` inference time for batch size 1, about `199.7` inferences per second, and about `72.4%` target utilization. It also reports about `3.62M` NPU cycles, plus SRAM and DRAM access cycles. Treat these as target-aware estimates for the NPU portion of the model, not as final runtime measurements from hardware. + +MLIA is now advising on where to investigate to improve target performance. In this report, the advice identifies the ten layers that make up most operator cycles, flags five high-impact layers with low MAC utilization, and flags five high-impact layers as possibly memory-bound. Low MAC utilization can be expected for layers with small channel counts, small spatial dimensions, or heavy memory movement, so these are the layers to consider adjusting. + +## What you have learned + +You have used MLIA to check compatibility and estimate performance with TensorFlow Lite and Vela. You have also learned how to read target metadata, backend metadata, metrics, operator placement, and advice. + +Next, you will inspect the TOSA intermediate representation using MLIA. diff --git a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/4-analyze-tosa-with-vela.md b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/4-analyze-tosa-with-vela.md new file mode 100644 index 0000000000..f57756cfc6 --- /dev/null +++ b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/4-analyze-tosa-with-vela.md @@ -0,0 +1,69 @@ +--- +title: Analyze TOSA IR artifacts with Vela + +weight: 5 + +### FIXED, DO NOT MODIFY +layout: "learningpathall" +--- + +## What is TOSA? + +TOSA stands for Tensor Operator Set Architecture. It is an intermediate representation for machine learning graphs: a stable set of tensor operators that can sit between a model framework and a target backend. + +Instead of asking every backend to understand every framework operator directly, a conversion flow can lower supported parts of a model into TOSA. Backend tools can then analyze or compile that TOSA graph for a target. + +TOSA can appear in more than one kind of ML workflow. Some compiler flows from TensorFlow or TensorFlow Lite models can use TOSA as an intermediate representation, while other TensorFlow Lite flows hand a `.tflite` file directly to a backend tool such as Vela. In the ExecuTorch Arm Ethos-U flow, supported PyTorch graph regions are lowered to TOSA before Vela compiles them for Ethos-U. + +## Compare FP32 and INT8 TOSA artifacts + +The model artifacts repository includes floating-point and quantized TOSA variants: + +```output +tosa/mv2_fp32.tosa +tosa/mv2_int8.tosa +``` + +Start with the FP32 TOSA model: + +```bash +mlia check tosa/mv2_fp32.tosa \ + --target-profile ethos-u85-256 \ + --compatibility \ + --backend vela \ + --json +``` + +This result should tell a similar story to the FP32 TFLite check: the model can be expressed as an artifact, but it is not in the supported quantized integer form required for Ethos-U acceleration with this target profile. The important distinction is that TOSA describes an intermediate graph form, not a complete runtime deployment. + +Run the same compatibility check on the quantized INT8 TOSA model: + +```bash +mlia check tosa/mv2_int8.tosa \ + --target-profile ethos-u85-256 \ + --compatibility \ + --backend vela \ + --json +``` + +The INT8 TOSA report should show `status` as `ok` and `accelerator_operator_percentage` as `100.0`. The model format is now TOSA, but the target profile and Vela backend configuration are the same as the TFLite run. The difference from the FP32 TOSA artifact is that the INT8 artifact has the quantized representation required for Ethos-U acceleration, so the operator support checks pass and MLIA expects the operator work to map to the NPU path. + +You can also run a performance check using the INT8 `.tosa` model: + +```bash +mlia check tosa/mv2_int8.tosa \ + --target-profile ethos-u85-256 \ + --performance \ + --backend vela \ + --json +``` + +The report should tell the same broad story as the INT8 TFLite performance result: the model maps to the NPU path, MLIA reports NPU-scoped estimated metrics, and the advice points you toward operators that dominate estimated cycles or have low utilization. + +Whether it is useful to you to use TOSA with MLIA, will depend on your workflow. Many developers will likely be using ExecuTorch or TFLite artifacts directly. + +## What you have learned + +You have seen how the same MLIA CLI pattern applies to TOSA, and also how TOSA bridges model formats and backend compilation. + +Next, you will look at how the ExecuTorch `.pt2` and `.pte` routes fit into the same MLIA workflow. diff --git a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/5-executorch-pt2-pte.md b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/5-executorch-pt2-pte.md new file mode 100644 index 0000000000..7f5bb17fe9 --- /dev/null +++ b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/5-executorch-pt2-pte.md @@ -0,0 +1,179 @@ +--- +title: Analyze ExecuTorch artifacts with Corstone + +weight: 6 + +### FIXED, DO NOT MODIFY +layout: "learningpathall" +--- + +## Understand the ExecuTorch routes + +The previous sections used TensorFlow Lite and TOSA artifacts with Vela. You can also use MLIA if you are working with a PyTorch / ExecuTorch flow. + +A `.pte` file is a portable ExecuTorch executable: the packaged artifact that ExecuTorch can load and run. A `.pt2` file is a PyTorch exported program. It is useful earlier in the workflow, before the model has been packaged for a specific ExecuTorch runtime path. + +In this section, you first compare two packaged `.pte` artifacts with Corstone backends. Then you briefly look at the `.pt2` converter route that MLIA can use earlier in an ExecuTorch-oriented workflow. + +## Compare prebuilt .pte artifacts + +The model artifacts repository includes prebuilt Ethos-U `.pte` files: + +```output +pte/toy_conditional_select_int8_ethos_u55_256.pte +pte/toy_conditional_select_int8_ethos_u85_256.pte +``` + +These are synthetic learning artifacts, not benchmark models. They use a small convolution plus a conditional selection pattern to make target-dependent delegation easier to see. A model can be packaged for different Ethos-U targets and show different delegation and runtime-counter behavior. + +For supported `.pte` workloads, MLIA performance analysis uses a Corstone backend. In this context, Corstone refers to an Arm reference subsystem platform, and the backend uses an FVP, or Fixed Virtual Platform. An FVP is a software model of a hardware platform. It lets you run a packaged artifact in a target-like environment and collect performance counters without needing a physical board on your desk. + +For `.pte` artifacts, MLIA currently supports performance advice only. The artifact has already been packaged for an ExecuTorch deployment path, so Corstone is used to run the packaged program on an FVP and collect NPU counters. Use compatibility checks earlier in the flow, before packaging, when you are still asking whether operators and tensors can map to the target. + +Install the Corstone backends used in this section: + +```bash +mlia backend install corstone-320 +mlia backend install corstone-300 +``` + +{{% notice Note %}} +Corstone backend installation requires accepting a license. You will be prompted in the terminal to agree. +{{% /notice %}} + +If you do not install a Corstone backend before running the check command, MLIA can prompt you to install it during the check. + +{{% notice Note %}} +The Corstone FVP used by this backend may require the Python 3.9 shared library on the host. Install it before running the `.pte` checks: + +```bash +sudo apt install -y libpython3.9 +``` + +If your Ubuntu package repositories do not include `libpython3.9`, add the deadsnakes PPA and try again: + +```bash +sudo apt install -y software-properties-common +sudo add-apt-repository -y ppa:deadsnakes/ppa +sudo apt update +sudo apt install -y libpython3.9 +``` +{{% /notice %}} + +Run the Ethos-U55 artifact with the Corstone-300 backend: + +```bash +mlia check pte/toy_conditional_select_int8_ethos_u55_256.pte \ + --target-profile ethos-u55-256 \ + --performance \ + --backend corstone-300 +``` + +Then run the Ethos-U85 artifact with the Corstone-320 backend: + +```bash +mlia check pte/toy_conditional_select_int8_ethos_u85_256.pte \ + --target-profile ethos-u85-256 \ + --performance \ + --backend corstone-320 +``` + +The reports should show Corstone running each `.pte` artifact and collecting NPU performance counters. Read the Corstone report as runtime counter evidence: + +| Field | What it tells you | +| --- | --- | +| `NPU active cycles` | Cycles where the NPU was doing work. | +| `NPU idle cycles` | Cycles where the NPU was present but not active. A very small value means this run kept the NPU busy once work was issued. | +| `NPU total cycles` | Active plus idle cycles for the NPU portion of the run. | +| `NPU AXI0 RD/WR data beat` | Memory traffic on the AXI0 port, configured as SRAM for this target profile. | +| `NPU AXI1 RD/WR data beat` | Memory traffic on the AXI1 port, configured as DRAM for this target profile. | + +Use the reports to compare the NPU counters for each packaged artifact: + +| Artifact | Target profile | Backend | NPU total cycles | +| --- | --- | --- | --- | +| `toy_conditional_select_int8_ethos_u55_256.pte` | `ethos-u55-256` | `corstone-300` | `273,088` | +| `toy_conditional_select_int8_ethos_u85_256.pte` | `ethos-u85-256` | `corstone-320` | `19,056` | + +The comparison uses `ethos-u55-256` and `ethos-u85-256`, so both target profiles use 256 MACs per cycle. However, Ethos-U85 is a newer and higher performance NPU than Ethos-U55, and the target information in the reports also shows other platform differences, such as accelerator clock and memory configuration. + +The U55 report shows `271,725` NPU active cycles and `1,363` NPU idle cycles. The U85 report shows `18,183` NPU active cycles and `873` NPU idle cycles, for `19,056` total NPU cycles. This difference reflects the combined effect of improvements in the U85 over the U55. + +The Corstone report focuses on NPU counters. If part of the graph runs outside the NPU delegate, the CPU-side cost is not fully represented by the NPU cycle table. Corstone output is different from the earlier Vela estimates. Vela gave compiler-estimated layer-level advice before deployment. Corstone runs the packaged `.pte` artifact on an FVP and reports runtime-oriented NPU counters. For `.pte` inputs, the advice can be less detailed at the layer level, but the counter values are useful for comparing packaged artifacts under the same target profile and backend. + +## Use Model Explorer for graph structure + +Across this Learning Path, you used MLIA with TensorFlow Lite, TOSA, ExecuTorch `.pte`, and PyTorch exported program artifacts. MLIA answers target-aware questions about compatibility, estimated performance, runtime counters, and advice. To see how these artifacts can be visualized as graphs in Model Explorer, continue with the [Explore model artifacts with Model Explorer](/learning-paths/cross-platform/explore-model-artifacts-with-model-explorer/) Learning Path, which uses many of the same files from the model artifacts repository. + +Model Explorer can open some formats directly, while other formats use adapters. Use it alongside MLIA when you want to connect target advice with the graph structure that produced it. + +In the example above, not only is the U85 expected to perform better due to its platform differences, there is also a difference in the way the model delegates to the U85 vs U55. + +The model pattern is: + +```output +convolution / activation +conditional select +convolution / activation / convolution +``` + +On Ethos-U55, the conditional select part is not kept inside the Ethos-U delegate path. The partitioning therefore looks like this: + +```output +EthosUBackend region +aten::gt / aten::where outside delegate +EthosUBackend region +``` + +On Ethos-U85, that pattern is handled more cleanly by this target/backend flow, so the packaged artifact has one `EthosUBackend` region. This will provide further performance benefit, as there is reduced fragmentation between CPU and NPU. To visualize the delegation of models to different backends, we can use the Model Explorer tool, with Arm adapters. + +{{% notice Note %}} +An ExecuTorch `.pte` file can contain work outside an accelerator delegate, but there is no guarantee that every non-delegated operator can run on the CPU runtime you deploy. CPU execution depends on the kernel libraries linked into that runtime and the operators, dtypes, layouts, and shapes they support. Cortex-M bare-metal runtimes are usually built with a smaller, more selective kernel set than Cortex-A runtimes, because Cortex-M systems have tighter memory and storage constraints. In both cases, actual CPU fallback support depends on which kernels are included in the runtime build. +{{% /notice %}} + +## Use the .pt2 converter route + +To analyze `.pt2` inputs, you need the PyTorch converter plugin: + +```bash +pip install mlia-converters-pytorch +``` + +This plugin registers transformer names used by MLIA, including: + +```output +pt2_to_tosa +pt2_to_pte +pte_to_delegate +``` + +These transformer names describe how MLIA can prepare PyTorch or ExecuTorch artifacts for downstream analysis. For example, when you give MLIA a `.pt2` file, the converter plugin can prepare the exported program for a target-specific analysis route instead of treating the `.pt2` file as a final deployment artifact. + +The model artifacts repository includes a PyTorch exported program: + +```output +pt2/toy_conditional_select_fp32.pt2 +``` + +This artifact contains the same small model pattern used by the packaged `.pte` examples. + +Run MLIA on the `.pt2` file: + +```bash +mlia check pt2/toy_conditional_select_fp32.pt2 \ + --target-profile ethos-u85-256 \ + --performance \ + --backend vela +``` + +{{% notice TODO %}} +Confirm installation and what the above does +{{% /notice %}} + +With `mlia-converters-pytorch` installed, MLIA can use the converter route to prepare the PyTorch exported program for the requested target analysis. + +## What you have learned + +You have learned how `.tflite`, `.tosa`, `.pte`, and `.pt2` fit into MLIA workflows. You have also seen why PTE is useful for ExecuTorch artifact analysis, where TOSA can fit as an intermediate handoff, and why Model Explorer remains useful for graph structure. + +Next, you will use the Python API to integrate MLIA into another workflow. diff --git a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/6-python-api.md b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/6-python-api.md new file mode 100644 index 0000000000..149a922c84 --- /dev/null +++ b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/6-python-api.md @@ -0,0 +1,95 @@ +--- +title: Use the Python API + +weight: 7 + +### FIXED, DO NOT MODIFY +layout: "learningpathall" +--- + +## Use the API after the CLI + +The MLIA CLI is the primary workflow for this Learning Path. It is the best way to learn the tool, inspect output, and debug your environment. + +Use the Python API when you want another product, dashboard, workflow runner, or CI system to integrate MLIA. + +The `mlia` Python package exposes the same advisor functionality used by the CLI. In this section, you use `run_advisor()` as the main API entry point, and helper functions such as `list_targets()`, `list_target_profiles()`, and `list_backends()` to discover what the installed environment supports. + +## Run MLIA from Python and compare two models + +This example shows using the API to analyze two TFLite model variants, and then printing results and advice: + +```bash +cat > compare_mlia_models.py <<'PY' +from pathlib import Path + +from mlia import run_advisor + + +models = [ + Path("tflite/mv2_fp32.tflite"), + Path("tflite/mv2_int8.tflite"), +] + +for model in models: + result = run_advisor( + advice_category="compatibility", + target_profile="ethos-u85-256", + model=model, + backends=["vela"], + ) + + print() + print(model) + + for item in result["results"]: + print(item["kind"], item["status"]) + + for advice in item.get("advice", []): + print("advice:", advice["severity"], advice["message"].splitlines()[0]) +PY +``` + +Run the comparison script: + +```bash +python compare_mlia_models.py +``` + +The expected result is that `mv2_fp32.tflite` reports `compatibility incompatible`, while `mv2_int8.tflite` reports `compatibility ok`. You might still see warning advice for both models. For example, MLIA can report that `SOFTMAX` is a suboptimal activation even when the quantized model is otherwise compatible with the NPU. Compatibility tells you whether the model can map to the target; advice can still point out ways to improve it. + +You could take this further to: + +- extract selected metrics into a dashboard +- compare performance between model revisions +- fail a CI job if a key metric regresses +- surface advice messages in an internal model review tool + +## Discover capabilities from Python + +MLIA also exposes helper functions for discovery. Depending on the installed MLIA version, useful helpers can include: + +```bash +cat > discover_mlia.py <<'PY' +from mlia import list_backends, list_target_profiles, list_targets + + +print(list_targets()) +print(list_target_profiles()) +print(list_backends()) +PY +``` + +Run the discovery script: + +```bash +python discover_mlia.py +``` + +Use discovery in integrations so your product can report what the current environment supports. + +## What you have learned + +You have used the Python API to run the same kind of analysis you performed from the CLI. You have also seen how to compare model variants programmatically and why the API is useful for product integration or automation. + +Next, review where to go from here. diff --git a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/_index.md b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/_index.md new file mode 100644 index 0000000000..e93dac5fdf --- /dev/null +++ b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/_index.md @@ -0,0 +1,74 @@ +--- +title: Analyze ML models with Arm ML Inference Advisor + +description: Learn how to use Arm ML Inference Advisor from the command line to check model compatibility, estimate performance, and identify target-aware model improvement opportunities using Ethos-U as the example target. + +minutes_to_complete: 45 + +who_is_this_for: This Learning Path is for ML developers who want to use Arm's ML Inference Advisor (MLIA) to evaluate whether a model is suitable for a target before moving into deployment, graph inspection, or runtime profiling. + +learning_objectives: + - Explain what MLIA does and where it fits in model preparation + - Use the MLIA CLI to discover targets, target profiles, and backends + - Run compatibility and performance analysis on model artifacts + - Interpret MLIA JSON output, metrics, unavailable fields, and advice + - Compare how TOSA, TensorFlow Lite, and ExecuTorch artifacts enter MLIA workflows + - Understand how to use the MLIA Python API to integrate MLIA with other tools and workflows. + +prerequisites: + - Ubuntu 22.04 LTS or another compatible Linux environment + - Python 3.10 or later + - Git and Git LFS to download the model artifacts + - Basic familiarity with machine learning model deployment concepts + - Basic familiarity with command-line tools + +author: + - Matt Cossins + +### Tags +skilllevels: Introductory +subjects: ML +armips: + - Cortex-M + - Ethos-U + +operatingsystems: + - Linux + +tools_software_languages: + - MLIA + - Vela + - ExecuTorch + - PyTorch + - Python + - TOSA + - TensorFlow Lite + +further_reading: + - resource: + title: Arm MLIA + link: https://github.com/arm/mlia + type: repository + - resource: + title: MLIA Ethos-U Plugin + link: https://github.com/arm/mlia-ethos-u + type: repository + - resource: + title: MLIA PyTorch Converter Plugin + link: https://github.com/arm/mlia-converters-pytorch + type: repository + - resource: + title: Arm ML model artifacts + link: https://github.com/arm-education/ml-model-artifacts + type: repository + - resource: + title: Ethos-U Vela compiler + link: https://gitlab.arm.com/artificial-intelligence/ethos-u/ethos-u-vela + type: repository + +### FIXED, DO NOT MODIFY +# ================================================================================ +weight: 1 +layout: "learningpathall" +learning_path_main_page: "yes" +--- diff --git a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/_next-steps.md b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/_next-steps.md new file mode 100644 index 0000000000..cce01e5164 --- /dev/null +++ b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/_next-steps.md @@ -0,0 +1,9 @@ +--- +# ================================================================================ +# FIXED, DO NOT MODIFY THIS FILE +# ================================================================================ +weight: 8 +title: "Next Steps" +layout: "learningpathall" +--- + From b72ac8b2a8fddba50a68ca51654ccf19a3506490 Mon Sep 17 00:00:00 2001 From: Matt Cossins Date: Sat, 12 Sep 2026 10:56:13 +0100 Subject: [PATCH 2/4] Updates --- .../1-overview.md | 24 ++--- .../2-install-and-discover.md | 32 ++----- .../3-tflite-comparison.md | 32 +++---- .../4-analyze-tosa-with-vela.md | 14 +-- ...cutorch-pt2-pte.md => 5-executorch-pte.md} | 89 +++---------------- .../6-python-api.md | 12 +-- .../_index.md | 23 +++-- 7 files changed, 69 insertions(+), 157 deletions(-) rename content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/{5-executorch-pt2-pte.md => 5-executorch-pte.md} (53%) diff --git a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/1-overview.md b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/1-overview.md index 91a884b241..57d0adb36b 100644 --- a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/1-overview.md +++ b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/1-overview.md @@ -1,6 +1,8 @@ --- title: What is the ML Inference Advisor? +description: Understand where MLIA fits in model preparation and how it works with Vela, Model Explorer, and runtime profiling tools. + weight: 2 ### FIXED, DO NOT MODIFY @@ -11,7 +13,7 @@ layout: "learningpathall" Arm ML Inference Advisor (MLIA) helps you evaluate whether a machine learning model is suitable for a target inference platform. -In this Learning Path, you use MLIA from the command line to check model compatibility, estimate performance, and read advice that points toward useful model changes. These examples use Arm Ethos-U as an example target. +In this Learning Path, you use MLIA from the command line to check model compatibility, estimate performance, and read advice that points toward useful model changes. This Learning Path uses Arm Ethos-U as an example target. MLIA is most useful before full deployment or runtime profiling, when you are asking questions such as: @@ -33,7 +35,7 @@ The MLIA CLI is the primary workflow in this Learning Path. You will use it to: - request JSON output - inspect advice and metrics -After you understand the CLI workflow, you will briefly use the Python API. The API is useful when you want to embed MLIA results in another product, dashboard, CI job, or tool. +If you want to automate the same checks, there is an optional Python API section at the end of this Learning Path. The API is useful when you want to embed MLIA results in another product, dashboard, CI job, or tool. ## Using MLIA alongside other tools @@ -44,9 +46,9 @@ MLIA is not a replacement for graph visualization or runtime profiling. It is an | MLIA | Is this model suitable for my target, and what should I change? | | Model Explorer | What does the generated model artifact graph look like? | | Vela | How does the Ethos-U compiler map supported work onto the NPU? | -| Runtime-specific profiling tools | What happened when the model actually ran? For example, use ETRecord, ETDump, and ExecuTorch Inspector for ExecuTorch deployments, or TensorFlow Lite benchmark and profiling tools for TFLite deployments. | +| Runtime-specific profiling tools | What happened when the model actually ran? For example, use ETRecord, ETDump, and ExecuTorch Inspector for ExecuTorch deployments, or LiteRT benchmark and profiling tools for LiteRT deployments. | -For example, MLIA can tell you which layers dominate estimated cycles or have low MAC utilization. Model Explorer can show how an ExecuTorch `.pte` artifact is partitioned into delegate regions. Runtime-specific profiling tools can show behavior after you have a runnable deployment. +For example, Vela-backed MLIA reports can tell you which layers dominate estimated cycles or have low MAC utilization. Model Explorer can show how an ExecuTorch `.pte` artifact is partitioned into delegate regions. Runtime-specific profiling tools can show behavior after you have a runnable deployment. Use these tools together: @@ -54,27 +56,15 @@ Use these tools together: - Use Model Explorer to inspect generated artifacts and delegation structure. - Use runtime profiling tools after you can execute the model. -## Understand the plugin model - -MLIA uses a plugin model. The `mlia` core package provides the shared command-line interface, output structure, and Python API. Target, backend, and converter support is added through plugins. - -The important repositories are: - -- `arm/mlia`: core MLIA package -- `arm/mlia-ethos-u`: Ethos-U target plugin and Vela/Corstone backend plugins -- `arm/mlia-converters-pytorch`: PyTorch `.pt2` converter plugin for TOSA and PTE routes -- `arm/mlia-legacy`: legacy support for older MLIA flows - ## Understand the model formats MLIA can analyze different kinds of model artifacts depending on what workflow you are using and the stage you want to analyze. | Format | Where it fits | | --- | --- | -| `.pt2` | PyTorch exported program input for PyTorch and ExecuTorch-oriented workflows. | | `.tosa` | Intermediate representation consumed by compiler/backend flows such as Ethos-U Vela. | | `.pte` | Serialized ExecuTorch program. Ethos-U `.pte` performance analysis uses Corstone backends. | -| `.tflite` | TensorFlow Lite model format used in many Ethos-U and embedded ML workflows. | +| `.tflite` | LiteRT model format used in many Ethos-U and embedded ML workflows. | ## What you have learned diff --git a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/2-install-and-discover.md b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/2-install-and-discover.md index 631d83bf42..3c06913b55 100644 --- a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/2-install-and-discover.md +++ b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/2-install-and-discover.md @@ -1,6 +1,8 @@ --- title: Install MLIA and discover capabilities +description: Install the MLIA Ethos-U plugin, inspect target profiles and backends, and download the model artifacts used in the analysis examples. + weight: 3 ### FIXED, DO NOT MODIFY @@ -28,7 +30,7 @@ sudo apt install -y git-lfs ## Create a Python environment -Create a virtual environment so the MLIA packages do not conflict with any existing PyTorch, ExecuTorch, or TensorFlow environment. +Create a virtual environment so the MLIA packages do not conflict with any existing ML framework environment. ```bash python3 -m venv mlia_env @@ -38,10 +40,6 @@ python -m pip install --upgrade pip ## Install MLIA -{{% notice TODO %}} -Confirm installation -{{% /notice %}} - MLIA uses plugins. The examples in this Learning Path use Ethos-U as the target, so install the Ethos-U plugin package: ```bash @@ -106,7 +104,7 @@ Backends perform the work behind an MLIA analysis flow. List available and insta mlia backend list ``` -For this Ethos-U demonstration, you should expect Vela and Corstone backend options. Vela is used for compiler-oriented compatibility and performance analysis. Corstone backends are used for simulation-oriented performance flows, including supported ExecuTorch `.pte` workloads. +For this Ethos-U demonstration, you should expect Vela and Corstone backend options. Vela is used for compiler-oriented compatibility and performance analysis, including per-operator performance estimates for supported formats. Corstone backends are used for simulation-oriented performance flows, including supported ExecuTorch `.pte` workloads, and report model-wide NPU counters. ```output Name Installed Installable @@ -116,19 +114,7 @@ corstone-320 no yes vela no yes ``` -Install Vela: - -```bash -mlia backend install vela -``` - -Check the backend list again: - -```bash -mlia backend list -``` - -You should now see `vela` in the installed backend list. +When we later use `mlia check`, any missing backends required by your target will be installed. ## Clone model artifacts @@ -138,9 +124,9 @@ This Learning Path uses prebuilt artifacts from the Arm ML model artifacts repos git lfs install git clone --filter=blob:none --sparse https://github.com/arm-education/ml-model-artifacts.git cd ml-model-artifacts -git sparse-checkout set pt2 pte tflite tosa +git sparse-checkout set pte tflite tosa git lfs pull \ - --include="pt2/toy_conditional_select_fp32.pt2,pte/toy_conditional_select_int8_ethos_u55_256.pte,pte/toy_conditional_select_int8_ethos_u85_256.pte,tflite/mv2_fp32.tflite,tflite/mv2_int8.tflite,tosa/mv2_fp32.tosa,tosa/mv2_int8.tosa" \ + --include="pte/toy_conditional_select_int8_ethos_u55_256.pte,pte/toy_conditional_select_int8_ethos_u85_256.pte,tflite/mv2_fp32.tflite,tflite/mv2_int8.tflite,tosa/mv2_fp32.tosa,tosa/mv2_int8.tosa" \ --exclude="" git lfs checkout ``` @@ -170,8 +156,6 @@ ml-model-artifacts/ ├── pte/ │ ├── toy_conditional_select_int8_ethos_u55_256.pte │ └── toy_conditional_select_int8_ethos_u85_256.pte -├── pt2/ -│ └── toy_conditional_select_fp32.pt2 ├── tflite/ │ ├── mv2_fp32.tflite │ └── mv2_int8.tflite @@ -182,6 +166,6 @@ ml-model-artifacts/ ## What you have learned -You have installed MLIA, along with the Ethos-U plugin, discovered available target profiles and backends from the CLI, installed Vela, and cloned model artifacts for analysis. +You have installed MLIA, along with the Ethos-U plugin, discovered available target profiles and backends from the CLI, and cloned model artifacts for analysis. Next, you will run your first MLIA compatibility and performance checks. diff --git a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/3-tflite-comparison.md b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/3-tflite-comparison.md index 82c7b1a1d8..f0a21a1ac9 100644 --- a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/3-tflite-comparison.md +++ b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/3-tflite-comparison.md @@ -1,5 +1,7 @@ --- -title: Analyze TensorFlow Lite artifacts with Vela +title: Analyze LiteRT artifacts with Vela + +description: Run MLIA compatibility and Vela performance checks on LiteRT models and inspect JSON metrics, operator placement, and advice. weight: 4 @@ -9,18 +11,18 @@ layout: "learningpathall" ## Run a compatibility check -Start by asking MLIA whether a TensorFlow Lite model can map to the selected target profile. +Start by asking MLIA whether a LiteRT model can map to the selected target profile. -TensorFlow Lite is a compact model format and runtime stack for deploying machine learning models on mobile, embedded, and edge devices. In embedded ML workflows, a `.tflite` file is often the artifact handed to backend tools for target-specific compatibility checks, compilation, or runtime deployment. +LiteRT is a compact model format and runtime stack for deploying machine learning models on mobile, embedded, and edge devices. In embedded ML workflows, a `.tflite` file is often the artifact handed to backend tools for target-specific compatibility checks, compilation, or runtime deployment. -The model artifacts repository includes floating-point and quantized MobileNetV2 TensorFlow Lite variants: +The model artifacts repository includes floating-point and quantized MobileNetV2 LiteRT variants: ```output tflite/mv2_fp32.tflite tflite/mv2_int8.tflite ``` -From the `ml-model-artifacts` directory, run MLIA on the FP32 TFLite artifact: +From the `ml-model-artifacts` directory, run MLIA on the FP32 LiteRT artifact: ```bash mlia check tflite/mv2_fp32.tflite \ @@ -32,15 +34,15 @@ mlia check tflite/mv2_fp32.tflite \ The `--json` option makes the output easier to inspect, compare, and automate. -You can also use the shorter `-t` and `-b` options instead of `--target-profile` and `--backend`. The `-t` option still means target profile: +You can also use the shorter `-t` and `-b` options instead of `--target-profile` and `--backend`. Note `--compatibility` and `vela` are both defaults that could be omitted: ```bash -mlia check tflite/mv2_fp32.tflite -t ethos-u85-256 --compatibility -b vela --json +mlia check tflite/mv2_fp32.tflite -t ethos-u85-256 --json ``` -This report should show the FP32 TFLite model as incompatible for Ethos-U acceleration, with `accelerator_operator_percentage` set to `0`. This is expected because Ethos-U acceleration requires supported quantized integer workloads. The failed checks explain that the input, output, and weight tensors are missing quantization parameters. +This report should show the FP32 LiteRT model as incompatible for Ethos-U acceleration, with `accelerator_operator_percentage` set to `0`. This is expected because Ethos-U acceleration requires supported quantized integer workloads. The failed checks explain that the input, output, and weight tensors are missing quantization parameters. -This does not mean TensorFlow Lite is the problem. It means this particular TFLite artifact is not in the numeric form the selected Ethos-U target needs. Use a quantized model instead. +This does not mean LiteRT is the problem. It means this particular LiteRT artifact is not in the numeric form the selected Ethos-U target needs. Use a quantized model instead. Now run the same compatibility check on the quantized INT8 model: @@ -52,7 +54,7 @@ mlia check tflite/mv2_int8.tflite \ --json ``` -You will generate a lengthy report. A small snippet is included below: +You will generate a lengthy report, saved in the `mlia-output` directory in the root of the `mlia` execution directory. A small snippet is included below: ```output { @@ -102,7 +104,7 @@ You will generate a lengthy report. A small snippet is included below: | `checks` | The individual operator support checks. | | `entities` | The operators MLIA analyzed, including placement and operator type. | -For the INT8 TFLite file, the important result is that `status` is `ok` and `accelerator_operator_percentage` is `100.0`. That means Vela found the operators in this quantized MobileNetV2 TFLite artifact compatible with the selected `ethos-u85-256` target profile, and MLIA expects the operator work to map to the NPU path for this compatibility check. It does not prove runtime latency or application accuracy. It tells you the model is a good candidate for the next step: performance estimation and deeper deployment testing. +For the INT8 LiteRT file, the important result is that `status` is `ok` and `accelerator_operator_percentage` is `100.0`. That means Vela found the operators in this quantized MobileNetV2 LiteRT artifact compatible with the selected `ethos-u85-256` target profile, and MLIA expects the operator work to map to the NPU path for this compatibility check. It does not prove runtime latency or application accuracy. It tells you the model is a good candidate for the next step: performance estimation and deeper deployment testing. ## Read operator placement @@ -123,7 +125,7 @@ Names and attributes vary by input format and MLIA version. The important part i ## Run a performance check -Next, ask MLIA for a target-aware performance estimate using the INT8 TFLite model: +Next, ask MLIA for a target-aware performance estimate using the INT8 LiteRT model: ```bash mlia check tflite/mv2_int8.tflite \ @@ -133,7 +135,7 @@ mlia check tflite/mv2_int8.tflite \ --json ``` -The performance report uses the same top-level structure as the compatibility report, but the `results` object now contains estimated performance metrics, operator-level breakdowns, and advice: +Because this run uses Vela, the performance report uses the same top-level structure as the compatibility report, but the `results` object now contains estimated performance metrics, operator-level breakdowns, and advice: ```output { @@ -201,12 +203,12 @@ The performance report uses the same top-level structure as the compatibility re | `advice` | MLIA's interpretation of the metrics, including which layers dominate cycles or may be inefficient. | | `availability` and `reason` | Why a metric is not available from the selected backend, if MLIA cannot report it. | -For this INT8 TFLite file, the Vela-backed estimate reports about `5.01M` total cycles, about `5.01 ms` inference time for batch size 1, about `199.7` inferences per second, and about `72.4%` target utilization. It also reports about `3.62M` NPU cycles, plus SRAM and DRAM access cycles. Treat these as target-aware estimates for the NPU portion of the model, not as final runtime measurements from hardware. +For this INT8 LiteRT file, the Vela-backed estimate reports about `5.01M` total cycles, about `5.01 ms` inference time for batch size 1, about `199.7` inferences per second, and about `72.4%` target utilization. It also reports about `3.62M` NPU cycles, plus SRAM and DRAM access cycles. Treat these as target-aware estimates for the NPU portion of the model, not as final runtime measurements from hardware. MLIA is now advising on where to investigate to improve target performance. In this report, the advice identifies the ten layers that make up most operator cycles, flags five high-impact layers with low MAC utilization, and flags five high-impact layers as possibly memory-bound. Low MAC utilization can be expected for layers with small channel counts, small spatial dimensions, or heavy memory movement, so these are the layers to consider adjusting. ## What you have learned -You have used MLIA to check compatibility and estimate performance with TensorFlow Lite and Vela. You have also learned how to read target metadata, backend metadata, metrics, operator placement, and advice. +You have used MLIA to check compatibility and estimate performance with LiteRT and Vela. You have also learned how to read target metadata, backend metadata, metrics, operator placement, and advice. Next, you will inspect the TOSA intermediate representation using MLIA. diff --git a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/4-analyze-tosa-with-vela.md b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/4-analyze-tosa-with-vela.md index f57756cfc6..de032969a8 100644 --- a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/4-analyze-tosa-with-vela.md +++ b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/4-analyze-tosa-with-vela.md @@ -1,6 +1,8 @@ --- title: Analyze TOSA IR artifacts with Vela +description: Analyze TOSA artifacts with MLIA and Vela to compare FP32 and INT8 compatibility, performance estimates, and NPU mapping. + weight: 5 ### FIXED, DO NOT MODIFY @@ -13,7 +15,7 @@ TOSA stands for Tensor Operator Set Architecture. It is an intermediate represen Instead of asking every backend to understand every framework operator directly, a conversion flow can lower supported parts of a model into TOSA. Backend tools can then analyze or compile that TOSA graph for a target. -TOSA can appear in more than one kind of ML workflow. Some compiler flows from TensorFlow or TensorFlow Lite models can use TOSA as an intermediate representation, while other TensorFlow Lite flows hand a `.tflite` file directly to a backend tool such as Vela. In the ExecuTorch Arm Ethos-U flow, supported PyTorch graph regions are lowered to TOSA before Vela compiles them for Ethos-U. +TOSA can appear in more than one kind of ML workflow. Some compiler flows for LiteRT models can use TOSA as an intermediate representation, while other LiteRT flows hand a `.tflite` file directly to a backend tool such as Vela. In the ExecuTorch Arm Ethos-U flow, supported PyTorch graph regions are lowered to TOSA before Vela compiles them for Ethos-U. ## Compare FP32 and INT8 TOSA artifacts @@ -34,7 +36,7 @@ mlia check tosa/mv2_fp32.tosa \ --json ``` -This result should tell a similar story to the FP32 TFLite check: the model can be expressed as an artifact, but it is not in the supported quantized integer form required for Ethos-U acceleration with this target profile. The important distinction is that TOSA describes an intermediate graph form, not a complete runtime deployment. +This result should tell a similar story to the FP32 LiteRT check: the model can be expressed as an artifact, but it is not in the supported quantized integer form required for Ethos-U acceleration with this target profile. The important distinction is that TOSA describes an intermediate graph form, not a complete runtime deployment. Run the same compatibility check on the quantized INT8 TOSA model: @@ -46,7 +48,7 @@ mlia check tosa/mv2_int8.tosa \ --json ``` -The INT8 TOSA report should show `status` as `ok` and `accelerator_operator_percentage` as `100.0`. The model format is now TOSA, but the target profile and Vela backend configuration are the same as the TFLite run. The difference from the FP32 TOSA artifact is that the INT8 artifact has the quantized representation required for Ethos-U acceleration, so the operator support checks pass and MLIA expects the operator work to map to the NPU path. +The INT8 TOSA report should show `status` as `ok` and `accelerator_operator_percentage` as `100.0`. The model format is now TOSA, but the target profile and Vela backend configuration are the same as the LiteRT run. The difference from the FP32 TOSA artifact is that the INT8 artifact has the quantized representation required for Ethos-U acceleration, so the operator support checks pass and MLIA expects the operator work to map to the NPU path. You can also run a performance check using the INT8 `.tosa` model: @@ -58,12 +60,12 @@ mlia check tosa/mv2_int8.tosa \ --json ``` -The report should tell the same broad story as the INT8 TFLite performance result: the model maps to the NPU path, MLIA reports NPU-scoped estimated metrics, and the advice points you toward operators that dominate estimated cycles or have low utilization. +The report should tell the same broad story as the INT8 LiteRT performance result: the model maps to the NPU path, MLIA reports NPU-scoped estimated metrics, and the advice points you toward operators that dominate estimated cycles or have low utilization. -Whether it is useful to you to use TOSA with MLIA, will depend on your workflow. Many developers will likely be using ExecuTorch or TFLite artifacts directly. +Whether it is useful to you to use TOSA with MLIA, will depend on your workflow. Many developers will likely be using ExecuTorch or LiteRT artifacts directly. ## What you have learned You have seen how the same MLIA CLI pattern applies to TOSA, and also how TOSA bridges model formats and backend compilation. -Next, you will look at how the ExecuTorch `.pt2` and `.pte` routes fit into the same MLIA workflow. +Next, you will look at how packaged ExecuTorch `.pte` artifacts fit into the same MLIA workflow. diff --git a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/5-executorch-pt2-pte.md b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/5-executorch-pte.md similarity index 53% rename from content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/5-executorch-pt2-pte.md rename to content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/5-executorch-pte.md index 7f5bb17fe9..3054546b87 100644 --- a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/5-executorch-pt2-pte.md +++ b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/5-executorch-pte.md @@ -1,19 +1,21 @@ --- title: Analyze ExecuTorch artifacts with Corstone +description: Run MLIA Corstone checks on packaged ExecuTorch PTE artifacts and compare model-wide NPU counters for Ethos-U55 and Ethos-U85. + weight: 6 ### FIXED, DO NOT MODIFY layout: "learningpathall" --- -## Understand the ExecuTorch routes +## Understand the ExecuTorch path -The previous sections used TensorFlow Lite and TOSA artifacts with Vela. You can also use MLIA if you are working with a PyTorch / ExecuTorch flow. +The previous sections used LiteRT and TOSA artifacts with Vela. You can also use MLIA if you are working with an ExecuTorch flow. -A `.pte` file is a portable ExecuTorch executable: the packaged artifact that ExecuTorch can load and run. A `.pt2` file is a PyTorch exported program. It is useful earlier in the workflow, before the model has been packaged for a specific ExecuTorch runtime path. +A `.pte` file is a portable ExecuTorch executable: the packaged artifact that ExecuTorch can load and run. -In this section, you first compare two packaged `.pte` artifacts with Corstone backends. Then you briefly look at the `.pt2` converter route that MLIA can use earlier in an ExecuTorch-oriented workflow. +In this section, you compare two packaged `.pte` artifacts with Corstone backends. ## Compare prebuilt .pte artifacts @@ -28,38 +30,14 @@ These are synthetic learning artifacts, not benchmark models. They use a small c For supported `.pte` workloads, MLIA performance analysis uses a Corstone backend. In this context, Corstone refers to an Arm reference subsystem platform, and the backend uses an FVP, or Fixed Virtual Platform. An FVP is a software model of a hardware platform. It lets you run a packaged artifact in a target-like environment and collect performance counters without needing a physical board on your desk. -For `.pte` artifacts, MLIA currently supports performance advice only. The artifact has already been packaged for an ExecuTorch deployment path, so Corstone is used to run the packaged program on an FVP and collect NPU counters. Use compatibility checks earlier in the flow, before packaging, when you are still asking whether operators and tensors can map to the target. +For `.pte` artifacts, MLIA currently supports performance analysis only. The artifact has already been packaged for an ExecuTorch deployment path, so Corstone is used to run the packaged program on an FVP and collect model-wide NPU counters. Corstone does not provide per-layer estimates or operator breakdowns. Use Vela-backed analysis earlier in the flow when you need layer-level performance estimates or compatibility checks. -Install the Corstone backends used in this section: - -```bash -mlia backend install corstone-320 -mlia backend install corstone-300 -``` +If the required Corstone backends are not yet installed, they will be after we run `mlia check` below. {{% notice Note %}} Corstone backend installation requires accepting a license. You will be prompted in the terminal to agree. {{% /notice %}} -If you do not install a Corstone backend before running the check command, MLIA can prompt you to install it during the check. - -{{% notice Note %}} -The Corstone FVP used by this backend may require the Python 3.9 shared library on the host. Install it before running the `.pte` checks: - -```bash -sudo apt install -y libpython3.9 -``` - -If your Ubuntu package repositories do not include `libpython3.9`, add the deadsnakes PPA and try again: - -```bash -sudo apt install -y software-properties-common -sudo add-apt-repository -y ppa:deadsnakes/ppa -sudo apt update -sudo apt install -y libpython3.9 -``` -{{% /notice %}} - Run the Ethos-U55 artifact with the Corstone-300 backend: ```bash @@ -78,7 +56,7 @@ mlia check pte/toy_conditional_select_int8_ethos_u85_256.pte \ --backend corstone-320 ``` -The reports should show Corstone running each `.pte` artifact and collecting NPU performance counters. Read the Corstone report as runtime counter evidence: +The reports should show Corstone running each `.pte` artifact and collecting model-wide NPU performance counters. Read the Corstone report as runtime counter evidence: | Field | What it tells you | | --- | --- | @@ -99,11 +77,11 @@ The comparison uses `ethos-u55-256` and `ethos-u85-256`, so both target profiles The U55 report shows `271,725` NPU active cycles and `1,363` NPU idle cycles. The U85 report shows `18,183` NPU active cycles and `873` NPU idle cycles, for `19,056` total NPU cycles. This difference reflects the combined effect of improvements in the U85 over the U55. -The Corstone report focuses on NPU counters. If part of the graph runs outside the NPU delegate, the CPU-side cost is not fully represented by the NPU cycle table. Corstone output is different from the earlier Vela estimates. Vela gave compiler-estimated layer-level advice before deployment. Corstone runs the packaged `.pte` artifact on an FVP and reports runtime-oriented NPU counters. For `.pte` inputs, the advice can be less detailed at the layer level, but the counter values are useful for comparing packaged artifacts under the same target profile and backend. +The Corstone report focuses on model-wide NPU counters. If part of the graph runs outside the NPU delegate, the CPU-side cost is not fully represented by the NPU cycle table. Corstone output is different from the earlier Vela estimates: Vela provides compiler-estimated layer-level metrics before deployment, while Corstone runs the packaged `.pte` artifact on an FVP and reports runtime-oriented NPU counters for the whole model. Use the Corstone counters to compare packaged artifacts under the same target profile and backend. ## Use Model Explorer for graph structure -Across this Learning Path, you used MLIA with TensorFlow Lite, TOSA, ExecuTorch `.pte`, and PyTorch exported program artifacts. MLIA answers target-aware questions about compatibility, estimated performance, runtime counters, and advice. To see how these artifacts can be visualized as graphs in Model Explorer, continue with the [Explore model artifacts with Model Explorer](/learning-paths/cross-platform/explore-model-artifacts-with-model-explorer/) Learning Path, which uses many of the same files from the model artifacts repository. +Across this Learning Path, you used MLIA with LiteRT, TOSA, and ExecuTorch `.pte` artifacts. MLIA answers target-aware questions about compatibility, estimated performance, runtime counters, and advice. To see how these artifacts can be visualized as graphs in Model Explorer, continue with the [Explore model artifacts with Model Explorer](/learning-paths/cross-platform/explore-model-artifacts-with-model-explorer/) Learning Path, which uses many of the same files from the model artifacts repository. Model Explorer can open some formats directly, while other formats use adapters. Use it alongside MLIA when you want to connect target advice with the graph structure that produced it. @@ -131,49 +109,8 @@ On Ethos-U85, that pattern is handled more cleanly by this target/backend flow, An ExecuTorch `.pte` file can contain work outside an accelerator delegate, but there is no guarantee that every non-delegated operator can run on the CPU runtime you deploy. CPU execution depends on the kernel libraries linked into that runtime and the operators, dtypes, layouts, and shapes they support. Cortex-M bare-metal runtimes are usually built with a smaller, more selective kernel set than Cortex-A runtimes, because Cortex-M systems have tighter memory and storage constraints. In both cases, actual CPU fallback support depends on which kernels are included in the runtime build. {{% /notice %}} -## Use the .pt2 converter route - -To analyze `.pt2` inputs, you need the PyTorch converter plugin: - -```bash -pip install mlia-converters-pytorch -``` - -This plugin registers transformer names used by MLIA, including: - -```output -pt2_to_tosa -pt2_to_pte -pte_to_delegate -``` - -These transformer names describe how MLIA can prepare PyTorch or ExecuTorch artifacts for downstream analysis. For example, when you give MLIA a `.pt2` file, the converter plugin can prepare the exported program for a target-specific analysis route instead of treating the `.pt2` file as a final deployment artifact. - -The model artifacts repository includes a PyTorch exported program: - -```output -pt2/toy_conditional_select_fp32.pt2 -``` - -This artifact contains the same small model pattern used by the packaged `.pte` examples. - -Run MLIA on the `.pt2` file: - -```bash -mlia check pt2/toy_conditional_select_fp32.pt2 \ - --target-profile ethos-u85-256 \ - --performance \ - --backend vela -``` - -{{% notice TODO %}} -Confirm installation and what the above does -{{% /notice %}} - -With `mlia-converters-pytorch` installed, MLIA can use the converter route to prepare the PyTorch exported program for the requested target analysis. - ## What you have learned -You have learned how `.tflite`, `.tosa`, `.pte`, and `.pt2` fit into MLIA workflows. You have also seen why PTE is useful for ExecuTorch artifact analysis, where TOSA can fit as an intermediate handoff, and why Model Explorer remains useful for graph structure. +You have learned how `.tflite`, `.tosa`, and `.pte` fit into MLIA workflows. You have also seen why PTE is useful for ExecuTorch artifact analysis, where TOSA can fit as an intermediate handoff, and why Model Explorer remains useful for graph structure. -Next, you will use the Python API to integrate MLIA into another workflow. +You can stop here if you only need the CLI workflow. Continue to the optional Python API section if you want to call MLIA from automation or another tool. diff --git a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/6-python-api.md b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/6-python-api.md index 149a922c84..3e41539d55 100644 --- a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/6-python-api.md +++ b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/6-python-api.md @@ -1,5 +1,7 @@ --- -title: Use the Python API +title: (Optional) Use the Python API + +description: Use the optional MLIA Python API to run compatibility checks, compare LiteRT models, and discover targets and backends programmatically. weight: 7 @@ -9,15 +11,15 @@ layout: "learningpathall" ## Use the API after the CLI -The MLIA CLI is the primary workflow for this Learning Path. It is the best way to learn the tool, inspect output, and debug your environment. +This section is optional. The MLIA CLI is the primary workflow for this Learning Path, and it is the best way to learn the tool, inspect output, and debug your environment. Use the Python API when you want another product, dashboard, workflow runner, or CI system to integrate MLIA. -The `mlia` Python package exposes the same advisor functionality used by the CLI. In this section, you use `run_advisor()` as the main API entry point, and helper functions such as `list_targets()`, `list_target_profiles()`, and `list_backends()` to discover what the installed environment supports. +The `mlia` Python package exposes the same advisor functionality used by the CLI. If you continue, you use `run_advisor()` as the main API entry point, and helper functions such as `list_targets()`, `list_target_profiles()`, and `list_backends()` to discover what the installed environment supports. ## Run MLIA from Python and compare two models -This example shows using the API to analyze two TFLite model variants, and then printing results and advice: +This example shows using the API to analyze two LiteRT model variants, and then printing results and advice: ```bash cat > compare_mlia_models.py <<'PY' @@ -91,5 +93,3 @@ Use discovery in integrations so your product can report what the current enviro ## What you have learned You have used the Python API to run the same kind of analysis you performed from the CLI. You have also seen how to compare model variants programmatically and why the API is useful for product integration or automation. - -Next, review where to go from here. diff --git a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/_index.md b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/_index.md index e93dac5fdf..790ca369a8 100644 --- a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/_index.md +++ b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/_index.md @@ -1,5 +1,5 @@ --- -title: Analyze ML models with Arm ML Inference Advisor +title: Analyze ML models with Arm ML Inference Advisor (MLIA) description: Learn how to use Arm ML Inference Advisor from the command line to check model compatibility, estimate performance, and identify target-aware model improvement opportunities using Ethos-U as the example target. @@ -8,12 +8,10 @@ minutes_to_complete: 45 who_is_this_for: This Learning Path is for ML developers who want to use Arm's ML Inference Advisor (MLIA) to evaluate whether a model is suitable for a target before moving into deployment, graph inspection, or runtime profiling. learning_objectives: - - Explain what MLIA does and where it fits in model preparation - - Use the MLIA CLI to discover targets, target profiles, and backends - - Run compatibility and performance analysis on model artifacts - - Interpret MLIA JSON output, metrics, unavailable fields, and advice - - Compare how TOSA, TensorFlow Lite, and ExecuTorch artifacts enter MLIA workflows - - Understand how to use the MLIA Python API to integrate MLIA with other tools and workflows. + - Use the MLIA CLI to discover target profiles and backends + - Run compatibility and performance analysis on LiteRT, TOSA, and ExecuTorch artifacts + - Interpret MLIA JSON output, advice, Vela estimates, and Corstone model-wide NPU counters + - (Optional) Call the MLIA Python API from automation or other tools prerequisites: - Ubuntu 22.04 LTS or another compatible Linux environment @@ -25,6 +23,10 @@ prerequisites: author: - Matt Cossins +generate_summary_faq: true +rerun_summary: false +rerun_faqs: false + ### Tags skilllevels: Introductory subjects: ML @@ -39,10 +41,9 @@ tools_software_languages: - MLIA - Vela - ExecuTorch - - PyTorch - Python - TOSA - - TensorFlow Lite + - LiteRT further_reading: - resource: @@ -53,10 +54,6 @@ further_reading: title: MLIA Ethos-U Plugin link: https://github.com/arm/mlia-ethos-u type: repository - - resource: - title: MLIA PyTorch Converter Plugin - link: https://github.com/arm/mlia-converters-pytorch - type: repository - resource: title: Arm ML model artifacts link: https://github.com/arm-education/ml-model-artifacts From 315bc1385a516c064aa95908c1cc91a784818dca Mon Sep 17 00:00:00 2001 From: Matt Cossins Date: Tue, 22 Sep 2026 08:53:17 +0100 Subject: [PATCH 3/4] Review updates --- .../1-overview.md | 16 +++++----- .../2-install-and-discover.md | 30 +++++++++---------- .../4-analyze-tosa-with-vela.md | 2 +- .../5-executorch-pte.md | 12 ++++---- .../_index.md | 2 +- 5 files changed, 31 insertions(+), 31 deletions(-) diff --git a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/1-overview.md b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/1-overview.md index 57d0adb36b..72406013bf 100644 --- a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/1-overview.md +++ b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/1-overview.md @@ -1,7 +1,7 @@ --- title: What is the ML Inference Advisor? -description: Understand where MLIA fits in model preparation and how it works with Vela, Model Explorer, and runtime profiling tools. +description: Understand where MLIA fits in model preparation and how Vela, Corstone FVP, Model Explorer, and runtime profiling tools support different checks. weight: 2 @@ -18,9 +18,9 @@ In this Learning Path, you use MLIA from the command line to check model compati MLIA is most useful before full deployment or runtime profiling, when you are asking questions such as: - Will this model map cleanly to my target? -- Which target profile should I use for early analysis? - Which operators or layers are likely to matter most for performance? - Is the model compute-bound, memory-bound, or affected by low MAC utilization? +- Is the model well-optimized for my Arm target hardware? - What should I investigate before building firmware or running on a board? MLIA does not make the final optimization decision for you. It gives target-aware evidence so you can decide what to change, what to measure next, and which workflow stage deserves attention. @@ -32,7 +32,6 @@ The MLIA CLI is the primary workflow in this Learning Path. You will use it to: - discover installed targets, target profiles, and backends - run compatibility checks - run performance analysis -- request JSON output - inspect advice and metrics If you want to automate the same checks, there is an optional Python API section at the end of this Learning Path. The API is useful when you want to embed MLIA results in another product, dashboard, CI job, or tool. @@ -41,14 +40,17 @@ If you want to automate the same checks, there is an optional Python API section MLIA is not a replacement for graph visualization or runtime profiling. It is an advisory layer that helps earlier in the model preparation workflow. -| Tool | Use it to answer | +| Tool or backend | Use it to answer | | --- | --- | | MLIA | Is this model suitable for my target, and what should I change? | +| Vela | Which operators are supported, and which layers dominate compiler-estimated cycles? | +| Corstone FVP | What NPU performance counters does a packaged `.pte` artifact produce for the whole model run on a virtual platform? | | Model Explorer | What does the generated model artifact graph look like? | -| Vela | How does the Ethos-U compiler map supported work onto the NPU? | | Runtime-specific profiling tools | What happened when the model actually ran? For example, use ETRecord, ETDump, and ExecuTorch Inspector for ExecuTorch deployments, or LiteRT benchmark and profiling tools for LiteRT deployments. | -For example, Vela-backed MLIA reports can tell you which layers dominate estimated cycles or have low MAC utilization. Model Explorer can show how an ExecuTorch `.pte` artifact is partitioned into delegate regions. Runtime-specific profiling tools can show behavior after you have a runnable deployment. +Vela-backed MLIA checks use compiler estimates. They can include operator-level breakdowns, such as which layers dominate estimated cycles or have low MAC utilization. Corstone-backed MLIA checks run a packaged `.pte` file on an FVP and report NPU performance counters for the whole model run. They do not provide per-layer estimates or operator breakdowns. + +Model Explorer can show how an ExecuTorch `.pte` artifact is partitioned into delegate regions. Runtime-specific profiling tools can show behavior after you have a runnable deployment. Use these tools together: @@ -62,9 +64,9 @@ MLIA can analyze different kinds of model artifacts depending on what workflow y | Format | Where it fits | | --- | --- | -| `.tosa` | Intermediate representation consumed by compiler/backend flows such as Ethos-U Vela. | | `.pte` | Serialized ExecuTorch program. Ethos-U `.pte` performance analysis uses Corstone backends. | | `.tflite` | LiteRT model format used in many Ethos-U and embedded ML workflows. | +| `.tosa` | Intermediate representation consumed by compiler/backend flows such as Ethos-U Vela. | ## What you have learned diff --git a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/2-install-and-discover.md b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/2-install-and-discover.md index 3c06913b55..24c718bd61 100644 --- a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/2-install-and-discover.md +++ b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/2-install-and-discover.md @@ -13,19 +13,17 @@ layout: "learningpathall" Use Ubuntu 22.04 LTS or another compatible Linux environment with Python 3.10 or later. -Some Python environments also require the Python development package, such as `libpython3.10-dev`, before installing MLIA packages. - Check that Git LFS is installed: ```bash git lfs version ``` -If this command fails, install Git LFS: +If this command fails, install Git LFS and the Python development package: ```bash sudo apt update -sudo apt install -y git-lfs +sudo apt install -y git-lfs python3.10-dev ``` ## Create a Python environment @@ -76,17 +74,17 @@ mlia target list For Ethos-U, typical bundled profiles include: -```output -ethos-u55-128 -ethos-u55-256 -ethos-u65-256 -ethos-u65-512 -ethos-u85-128 -ethos-u85-256 -ethos-u85-512 -ethos-u85-1024 -ethos-u85-2048 -``` +| Target profile | Ethos-U NPU | MACs per cycle | +| --- | --- | --- | +| `ethos-u55-128` | Ethos-U55 | 128 | +| `ethos-u55-256` | Ethos-U55 | 256 | +| `ethos-u65-256` | Ethos-U65 | 256 | +| `ethos-u65-512` | Ethos-U65 | 512 | +| `ethos-u85-128` | Ethos-U85 | 128 | +| `ethos-u85-256` | Ethos-U85 | 256 | +| `ethos-u85-512` | Ethos-U85 | 512 | +| `ethos-u85-1024` | Ethos-U85 | 1024 | +| `ethos-u85-2048` | Ethos-U85 | 2048 | In this Learning Path, the examples use one Ethos-U85 profile: @@ -104,7 +102,7 @@ Backends perform the work behind an MLIA analysis flow. List available and insta mlia backend list ``` -For this Ethos-U demonstration, you should expect Vela and Corstone backend options. Vela is used for compiler-oriented compatibility and performance analysis, including per-operator performance estimates for supported formats. Corstone backends are used for simulation-oriented performance flows, including supported ExecuTorch `.pte` workloads, and report model-wide NPU counters. +For this Ethos-U demonstration, you should expect Vela and Corstone backend options. You use Vela for the LiteRT and TOSA checks, and Corstone for the packaged ExecuTorch `.pte` checks later in this Learning Path. ```output Name Installed Installable diff --git a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/4-analyze-tosa-with-vela.md b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/4-analyze-tosa-with-vela.md index de032969a8..97fa26ef47 100644 --- a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/4-analyze-tosa-with-vela.md +++ b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/4-analyze-tosa-with-vela.md @@ -11,7 +11,7 @@ layout: "learningpathall" ## What is TOSA? -TOSA stands for Tensor Operator Set Architecture. It is an intermediate representation for machine learning graphs: a stable set of tensor operators that can sit between a model framework and a target backend. +TOSA stands for Tensor Operator Set Architecture. It is an intermediate representation for machine learning graphs: a stable set of tensor operators that can sit between a model framework and a target backend. The [TOSA specification](https://www.mlplatform.org/tosa/tosa_spec.html) defines the operator set and semantics. Instead of asking every backend to understand every framework operator directly, a conversion flow can lower supported parts of a model into TOSA. Backend tools can then analyze or compile that TOSA graph for a target. diff --git a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/5-executorch-pte.md b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/5-executorch-pte.md index 3054546b87..7744fb9537 100644 --- a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/5-executorch-pte.md +++ b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/5-executorch-pte.md @@ -1,7 +1,7 @@ --- title: Analyze ExecuTorch artifacts with Corstone -description: Run MLIA Corstone checks on packaged ExecuTorch PTE artifacts and compare model-wide NPU counters for Ethos-U55 and Ethos-U85. +description: Run MLIA Corstone checks on packaged ExecuTorch PTE artifacts and compare whole-model NPU performance counters for Ethos-U55 and Ethos-U85. weight: 6 @@ -30,7 +30,7 @@ These are synthetic learning artifacts, not benchmark models. They use a small c For supported `.pte` workloads, MLIA performance analysis uses a Corstone backend. In this context, Corstone refers to an Arm reference subsystem platform, and the backend uses an FVP, or Fixed Virtual Platform. An FVP is a software model of a hardware platform. It lets you run a packaged artifact in a target-like environment and collect performance counters without needing a physical board on your desk. -For `.pte` artifacts, MLIA currently supports performance analysis only. The artifact has already been packaged for an ExecuTorch deployment path, so Corstone is used to run the packaged program on an FVP and collect model-wide NPU counters. Corstone does not provide per-layer estimates or operator breakdowns. Use Vela-backed analysis earlier in the flow when you need layer-level performance estimates or compatibility checks. +For `.pte` artifacts, MLIA currently supports performance analysis only. The artifact has already been packaged for an ExecuTorch deployment path, so Corstone is used to run the packaged program on an FVP and collect NPU performance counters for the whole model run. Corstone does not provide per-layer estimates or operator breakdowns. Use Vela-backed analysis earlier in the flow when you need layer-level performance estimates or compatibility checks. If the required Corstone backends are not yet installed, they will be after we run `mlia check` below. @@ -56,7 +56,7 @@ mlia check pte/toy_conditional_select_int8_ethos_u85_256.pte \ --backend corstone-320 ``` -The reports should show Corstone running each `.pte` artifact and collecting model-wide NPU performance counters. Read the Corstone report as runtime counter evidence: +The reports should show Corstone running each `.pte` artifact and collecting NPU performance counters for the whole model run. Read the Corstone report as runtime counter evidence: | Field | What it tells you | | --- | --- | @@ -75,9 +75,9 @@ Use the reports to compare the NPU counters for each packaged artifact: The comparison uses `ethos-u55-256` and `ethos-u85-256`, so both target profiles use 256 MACs per cycle. However, Ethos-U85 is a newer and higher performance NPU than Ethos-U55, and the target information in the reports also shows other platform differences, such as accelerator clock and memory configuration. -The U55 report shows `271,725` NPU active cycles and `1,363` NPU idle cycles. The U85 report shows `18,183` NPU active cycles and `873` NPU idle cycles, for `19,056` total NPU cycles. This difference reflects the combined effect of improvements in the U85 over the U55. +The Ethos-U55 report shows `271,725` NPU active cycles and `1,363` NPU idle cycles. The Ethos-U85 report shows `18,183` NPU active cycles and `873` NPU idle cycles, for `19,056` total NPU cycles. This difference reflects the combined effect of improvements in Ethos-U85 over Ethos-U55. -The Corstone report focuses on model-wide NPU counters. If part of the graph runs outside the NPU delegate, the CPU-side cost is not fully represented by the NPU cycle table. Corstone output is different from the earlier Vela estimates: Vela provides compiler-estimated layer-level metrics before deployment, while Corstone runs the packaged `.pte` artifact on an FVP and reports runtime-oriented NPU counters for the whole model. Use the Corstone counters to compare packaged artifacts under the same target profile and backend. +Use the Corstone counters to compare packaged artifacts under the same target profile and backend. If part of the graph runs outside the NPU delegate, the CPU-side cost is not fully represented by the NPU cycle table. ## Use Model Explorer for graph structure @@ -85,7 +85,7 @@ Across this Learning Path, you used MLIA with LiteRT, TOSA, and ExecuTorch `.pte Model Explorer can open some formats directly, while other formats use adapters. Use it alongside MLIA when you want to connect target advice with the graph structure that produced it. -In the example above, not only is the U85 expected to perform better due to its platform differences, there is also a difference in the way the model delegates to the U85 vs U55. +In the example above, not only is Ethos-U85 expected to perform better due to its platform differences, there is also a difference in the way the model delegates to Ethos-U85 vs Ethos-U55. The model pattern is: diff --git a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/_index.md b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/_index.md index 790ca369a8..fee15b12f7 100644 --- a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/_index.md +++ b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/_index.md @@ -10,7 +10,7 @@ who_is_this_for: This Learning Path is for ML developers who want to use Arm's M learning_objectives: - Use the MLIA CLI to discover target profiles and backends - Run compatibility and performance analysis on LiteRT, TOSA, and ExecuTorch artifacts - - Interpret MLIA JSON output, advice, Vela estimates, and Corstone model-wide NPU counters + - Interpret MLIA JSON output, advice, Vela estimates, and Corstone whole-model NPU performance counters - (Optional) Call the MLIA Python API from automation or other tools prerequisites: From 01515dd6ba604d661b815aa4186133ce86aeedca Mon Sep 17 00:00:00 2001 From: Matt Cossins Date: Wed, 23 Sep 2026 08:09:10 +0100 Subject: [PATCH 4/4] Switch to draft --- .../analyze-ethos-u-models-with-mlia/_index.md | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/_index.md b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/_index.md index fee15b12f7..dbca1642e9 100644 --- a/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/_index.md +++ b/content/learning-paths/embedded-and-microcontrollers/analyze-ethos-u-models-with-mlia/_index.md @@ -1,6 +1,10 @@ --- title: Analyze ML models with Arm ML Inference Advisor (MLIA) +draft: true +cascade: + draft: true + description: Learn how to use Arm ML Inference Advisor from the command line to check model compatibility, estimate performance, and identify target-aware model improvement opportunities using Ethos-U as the example target. minutes_to_complete: 45