diff --git a/content/learning-paths/mobile-graphics-and-gaming/run-text-to-text-models-on-android/2-build-and-run.md b/content/learning-paths/mobile-graphics-and-gaming/run-text-to-text-models-on-android/2-build-and-run.md index 5bb5e25cca..8bfeee4d0c 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/run-text-to-text-models-on-android/2-build-and-run.md +++ b/content/learning-paths/mobile-graphics-and-gaming/run-text-to-text-models-on-android/2-build-and-run.md @@ -87,6 +87,7 @@ When you want to try another runtime or workload, return to the following table: | [SmolLM2 360M Instruct](https://huggingface.co/Arm/smollm2-360m-instruct-8da4w-xnnpack-executorch) | ExecuTorch 1.1.0 | Text generation | `smollm2-executorch` | | [BGE Base English v1.5](https://huggingface.co/Arm/bge-base-en-v1.5-int8-litert) | LiteRT 1.4.2 | Text embedding | `bge-base-litert` | | [TinyLlama 1.1B Chat](https://huggingface.co/Arm/tinyllama-1-1b-chat-onnx-genai-int4-kquantlast-emb-int8-vivo-x300) | ONNX Runtime GenAI | Text generation | `tinyllama-onnx-genai` | +| [Llama 3.2 1B Instruct](https://huggingface.co/Arm/llama-3-2-1b-instruct-onnx-genai-int4-kquantlast-emb-int8-vivo-x300) | ONNX Runtime GenAI | Text generation | `llama-3-2-onnx-genai` | The model and runtime change together in these examples, so their results and timings aren't controlled runtime benchmarks. diff --git a/content/learning-paths/mobile-graphics-and-gaming/run-text-to-text-models-on-android/4-compare-runtime-adapters.md b/content/learning-paths/mobile-graphics-and-gaming/run-text-to-text-models-on-android/4-compare-runtime-adapters.md index e338631573..da6866e310 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/run-text-to-text-models-on-android/4-compare-runtime-adapters.md +++ b/content/learning-paths/mobile-graphics-and-gaming/run-text-to-text-models-on-android/4-compare-runtime-adapters.md @@ -24,7 +24,7 @@ The downloaded model package isn't an Android executable. Gradle packages the ru ## Choose a text-generation path -The application includes two validated text-generation paths that share the interface and deployment flow. However, each needs an adapter because its runtime API, model format, tokenizer handling, and output contract differ. +The application includes three validated text-generation examples that share the interface and deployment flow. However, each needs an adapter because its runtime API, model format, tokenizer handling, and output contract differ. The following table describes the paths: @@ -32,6 +32,7 @@ The following table describes the paths: | --- | --- | --- | --- | --- | | SmolLM2 with ExecuTorch | `executorch` | `inference/models/smollm2/ExecuTorchTextGenerationAdapter` | `.pte`, configuration, tokenizer, and chat template | The recommended default and smallest validated generation example. | | TinyLlama with ONNX Runtime GenAI | `onnxruntime` | `inference/models/tinyllama/OnnxTextGenerationAdapter` | ONNX model directory, configuration, tokenizer, and chat template | A directory-based GenAI package with an additional Android AAR dependency. | +| Llama 3.2 with ONNX Runtime GenAI | `onnxruntime` | `inference/models/llama32/OnnxTextGenerationAdapter` | ONNX model directory, configuration, tokenizer, and chat template | The same GenAI runtime flow with the Llama 3.2 package. | The repository also contains `inference/models/bge/LiteRtEmbeddingAdapter` for the catalog runtime `litert`. That adapter runs the BGE text-embedding workload and returns vectors rather than generated text, so it is listed separately from the text-generation paths. @@ -53,43 +54,86 @@ This checks compatibility with the adapter and doesn't evaluate the model's lang ### Prepare the ONNX Runtime GenAI path -The TinyLlama example uses `com.microsoft.onnxruntime:onnxruntime-android:1.27.0` and the official ONNX Runtime GenAI Android AAR. First, complete the [ONNX Runtime GenAI Android Learning Path](https://learn.arm.com/learning-paths/mobile-graphics-and-gaming/build-android-chat-app-using-onnxruntime/) to build `onnxruntime-genai-release.aar`. +The ONNX Runtime GenAI examples use `com.microsoft.onnxruntime:onnxruntime-android:1.27.0` and the official ONNX Runtime GenAI Android AAR. Download the official pre-compiled ONNX Runtime GenAI Android AAR before you apply an ONNX Runtime GenAI adapter. -Create a separate project and download the complete model directory: +{{< tabpane code=true >}} + {{< tab header="macOS or Linux" language="bash" >}} +export WORK_DIR="$HOME/text-to-text-android" +mkdir -p "$WORK_DIR" + +export ONNX_GENAI_AAR="$WORK_DIR/onnxruntime-genai-release.aar" + +curl --fail --location \ + "https://github.com/microsoft/onnxruntime-genai/releases/download/v0.16.0/onnxruntime-genai-android-0.16.0.aar" \ + --output "$ONNX_GENAI_AAR" + +test -f "$ONNX_GENAI_AAR" + {{< /tab >}} + {{< tab header="Windows PowerShell" language="powershell" >}} +$WORK_DIR = Join-Path $HOME "text-to-text-android" +New-Item -ItemType Directory -Force -Path $WORK_DIR + +$ONNX_GENAI_AAR = Join-Path $WORK_DIR "onnxruntime-genai-release.aar" + +Invoke-WebRequest ` + -Uri "https://github.com/microsoft/onnxruntime-genai/releases/download/v0.16.0/onnxruntime-genai-android-0.16.0.aar" ` + -OutFile $ONNX_GENAI_AAR + +Test-Path $ONNX_GENAI_AAR + {{< /tab >}} +{{< /tabpane >}} + +Create a separate project and select one ONNX Runtime GenAI example: {{< tabpane code=true >}} {{< tab header="macOS or Linux" language="bash" >}} -cd .. +cd "$WORK_DIR" git clone https://github.com/arm-education/ai-portal-android-app-text-to-text.git \ ai-portal-android-app-text-to-text-onnx cd ai-portal-android-app-text-to-text-onnx -cp -R adapter-examples/tinyllama-onnx-genai/app/. app/ - +# Choose one validated ONNX Runtime GenAI example. +export ONNX_EXAMPLE="tinyllama-onnx-genai" export MODEL_ID="tinyllama-1-1b-chat-onnx-genai-int4-kquantlast-emb-int8-vivo-x300" +# Or use Llama 3.2 instead: +# export ONNX_EXAMPLE="llama-3-2-onnx-genai" +# export MODEL_ID="llama-3-2-1b-instruct-onnx-genai-int4-kquantlast-emb-int8-vivo-x300" + +cp -R "adapter-examples/$ONNX_EXAMPLE/app/." app/ +mkdir -p app/libs +cp "$ONNX_GENAI_AAR" app/libs/onnxruntime-genai-release.aar + export MODEL_DIR="../model-onnx" ../.hf-venv/bin/hf download "Arm/$MODEL_ID" --local-dir "$MODEL_DIR" {{< /tab >}} {{< tab header="Windows PowerShell" language="powershell" >}} -Set-Location .. +Set-Location $WORK_DIR git clone https://github.com/arm-education/ai-portal-android-app-text-to-text.git ` ai-portal-android-app-text-to-text-onnx Set-Location ai-portal-android-app-text-to-text-onnx -Copy-Item -Path "adapter-examples\tinyllama-onnx-genai\app\*" ` +# Choose one validated ONNX Runtime GenAI example. +$ONNX_EXAMPLE = "tinyllama-onnx-genai" +$MODEL_ID = "tinyllama-1-1b-chat-onnx-genai-int4-kquantlast-emb-int8-vivo-x300" +# Or use Llama 3.2 instead: +# $ONNX_EXAMPLE = "llama-3-2-onnx-genai" +# $MODEL_ID = "llama-3-2-1b-instruct-onnx-genai-int4-kquantlast-emb-int8-vivo-x300" + +Copy-Item -Path "adapter-examples\$ONNX_EXAMPLE\app\*" ` -Destination "app" -Recurse -Force +New-Item -ItemType Directory -Force -Path "app\libs" +Copy-Item -Path $ONNX_GENAI_AAR -Destination "app\libs\onnxruntime-genai-release.aar" -Force -$MODEL_ID = "tinyllama-1-1b-chat-onnx-genai-int4-kquantlast-emb-int8-vivo-x300" $MODEL_DIR = "..\model-onnx" ..\.hf-venv\Scripts\hf.exe download "Arm/$MODEL_ID" --local-dir $MODEL_DIR {{< /tab >}} {{< /tabpane >}} -Copy the generated AAR to `app/libs/onnxruntime-genai-release.aar` before continuing. Keep every file in the downloaded ONNX model directory together, including external model data, `genai_config.json`, tokenizer files, and the chat template. +The commands above copy the generated AAR to `app/libs/onnxruntime-genai-release.aar`. Keep every file in the downloaded ONNX model directory together, including external model data, `genai_config.json`, tokenizer files, and the chat template. ## Build and run the selected alternative -The ONNX Runtime GenAI alternative now follows the same workflow: build the APK, install the APK, copy the chosen model package to the directory named by its catalog ID, and start the application. +The selected ONNX Runtime GenAI alternative now follows the same workflow: build the APK, install the APK, copy the chosen model package to the directory named by its catalog ID, and start the application. Continue in the same terminal so that `MODEL_ID` and `MODEL_DIR` remain set: