Skip to content

Arm backend: Use a writable HF cache in test-arm-backend-no-driver - #23268

Open
psiddh wants to merge 5 commits into
pytorch:mainfrom
psiddh:sidart/arm-ci-writable-hf-cache
Open

psiddh wants to merge 5 commits into
pytorch:mainfrom
psiddh:sidart/arm-ci-writable-hf-cache

Conversation

@psiddh

@psiddh psiddh commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

The mt-l-x86iavx512-8-64 OSDC runner mounts HF_HOME read-only at
/mnt/hf_cache. The prequantized NSS/NFRU tests added in #23236 and #23237 call
hf_hub_download() for checkpoints that aren't in that cache, so they fail with:

OSError: [Errno 30] Read-only file system: 
'/mnt/hf_cache/hub/models--Arm--neural-frame-rate-upscaling'

This redirects HF_HOME to a writable dir in test-arm-backend-no-driver, using
the same workaround as cuda-perf.yml and cuda-windows.yml.
test-arm-backend-vkml runs on a different runner type and doesn't need it.

The job no longer uses the shared read-only cache, so HF assets are downloaded
each run.

Test plan:

CI. `test_nss_prequantized_tosa_INT` and 

test_nfru_prequantized_tosa_INT should pass in test-arm-backend-no-driver (test_pytest_models_tosa).

The change: it's the six lines from cuda-perf.yml, added after conda activate in
the no-driver job only. The workflow YAML parses; actionlint isn't installed here,
so I couldn't run it. Nothing has run it in CI yet, so this PR's
test_pytest_models_tosa run is the real test.

Watch the runtime: that job took about 13 minutes in the last run. If
re-downloading the HF files makes it much slower, the better long-term fix is
asking test-infra to add the new checkpoints to the shared cache.

cc @digantdesai @freddan80 @per @zingo @oscarandersson8218 @mansnils @Sebastian-Larsson @robell @rascani

The mt-l-x86iavx512-8-64 OSDC runner mounts HF_HOME read-only at
/mnt/hf_cache, so the prequantized NSS/NFRU tests added in pytorch#23236 and
pytorch#23237 fail with EROFS when hf_hub_download fetches checkpoints that are
not already cached. Redirect HF_HOME to a writable dir, matching
cuda-perf.yml and cuda-windows.yml.

Authored with Claude Code.
Copilot AI balanced review requested due to automatic review settings September 29, 2026 21:37
@pytorch-bot

pytorch-bot Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23268

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 Cancelled Job, 6 Unrelated Failures

As of commit 33b29c9 with merge base 273cb33 (image):

CANCELLED JOB - The following job was cancelled. Please retry:

FLAKY - The following jobs failed but were likely due to flakiness present on trunk:

BROKEN TRUNK - The following jobs failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 29, 2026
@psiddh
psiddh requested review from robert-kalmar and zingo and removed request for Copilot September 29, 2026 21:37
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@zingo zingo left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice, superthanks for the help!
We migth want/need to do that on all test-arm-backend-* based job in pull and trunk also.

@zingo

zingo commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

We migth want/need to do that on all test-arm-backend-* based job in pull and trunk also.

or not? :

test-arm-backend-vkml runs on a different runner type and doesn't need it.

@zingo zingo added partner: arm For backend delegation, kernels, demo, etc. from the 3rd-party partner, Arm module: arm Issues related to arm backend ciflow/trunk labels Sep 29, 2026
@psiddh

psiddh commented Sep 29, 2026

Copy link
Copy Markdown
Contributor Author

We migth want/need to do that on all test-arm-backend-* based job in pull and trunk also.

or not? :

test-arm-backend-vkml runs on a different runner type and doesn't need it.

yeah, you're right for trunk. test-arm-backend-vkml in trunk.yml runs on an OSDC runner and picks up the *_prequantized_vgf_INT tests, so it'll hit the same EROFS, i think. I'll add the redirect there and to trunk test-arm-backend-ethos-u (same runner and HF setup). The pull vkml job is on linux.2xlarge.memory, and the other Arm jobs don't touch HF, so I'm leaving those alone.

The trunk vkml job also runs on an OSDC runner, and test_pytest_models_vkml
picks up the prequantized NSS/NFRU VGF tests, so it hits the same EROFS on
/mnt/hf_cache.

Authored with Claude Code.
Copilot AI balanced review requested due to automatic review settings September 29, 2026 22:37

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

HF_HUB_CACHE remains pointed at the read-only mount, so hf_hub_download() can still fail.

Review effort: Balanced
Findings: 2 High severity

Open (2)
What changed in this PR

Redirects Arm CI Hugging Face downloads to writable temporary caches.

Changes:

  • Adds RUNNER_TEMP cache paths with /tmp fallback.
  • Applies the workaround to pull and trunk Arm jobs.
File Description
.github/​workflows/​pull.yml Updates the no-driver Arm job cache.
.github/​workflows/​trunk.yml Updates the VKML Arm job cache.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread .github/workflows/pull.yml Outdated
Comment on lines +871 to +873
export HF_HOME="${RUNNER_TEMP:-/tmp}/hf_cache"
mkdir -p "${HF_HOME}" 2>/dev/null || export HF_HOME=/tmp/hf_cache
mkdir -p "${HF_HOME}"
Comment thread .github/workflows/trunk.yml Outdated
Comment on lines +347 to +349
export HF_HOME="${RUNNER_TEMP:-/tmp}/hf_cache"
mkdir -p "${HF_HOME}" 2>/dev/null || export HF_HOME=/tmp/hf_cache
mkdir -p "${HF_HOME}"
Copilot AI balanced review requested due to automatic review settings September 29, 2026 22:41
@psiddh

psiddh commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor Author

Ok I merged #23265 and this PR has label : 'ciflow/trunk'. , I also rebased this PR on the latest mainline..(should have the fix #23265)

Lets monitor the CI , all the issues related arm jobs should be resolved.

cc @zingo @metascroy @digantdesai @rascani

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The unresolved HF_HUB_CACHE override still directs downloads to the read-only mount.

Review effort: Balanced
Findings: 2 High severity

Open (2)

@zingo

zingo commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator

FYI @Michiel-Olieslagers

linux_job_v3 already sets a writable HF_HOME and points HF_HUB_CACHE at the
read-only /mnt/hf_cache/hub, which hf_hub_download prefers, so the redirect
had no effect. Missing checkpoints are seeded via the ci-refresh-hf-cache label.

Authored with Claude Code.
Copilot AI balanced review requested due to automatic review settings September 30, 2026 00:26

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot wasn't able to review any files in this pull request. Check if the Files changed in this pull request are included in default exclusions.

Still open (2)

@zingo

zingo commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator

This seems to fix that problem
I now only de separate problem
FAILED backends/arm/test/models/test_nss.py::test_nss_prequantized_vgf_INT[random_data] - AssertionError: Output 0 does not match reference output.
Given atol: 0.03249606287479401, rtol: 0.001.
Output tensor shape: torch.Size([1, 36, 136, 240]), dtype: torch.float32
Difference: max: 0.043307095766067505, abs: 0.043307095766067505, mean abs error: 0.00012670702797030749.
-- Model vs. Reference --
Numel: 1175040, 1175040
Median: 0.05118110030889511, 0.05118110030889511
Mean: 0.12287650634025973, 0.12287394988535973
Max: 0.9921259880065918, 0.9921259880065918
Min: 0.0, 0.0
FAILED backends/arm/test/models/test_nfru.py::test_nfru_vgf_quant[real_data-qat] - AssertionError: Output 0 does not match reference output.
Given atol: 0.7856333971023559, rtol: 0.001.
Output tensor shape: torch.Size([1, 4, 270, 480]), dtype: torch.float32
Difference: max: 0.5856342315673828, abs: 1.1712665557861328, mean abs error: 0.008185762909160536.
-- Model vs. Reference --
Numel: 518400, 518400
Median: -1.4640834331512451, -1.4640834331512451
Mean: -3.139178010777421, -3.1391497683884966
Max: 17.276185989379883, 17.276185989379883
Min: -43.33687210083008, -43.33687210083008
== 2 failed, 141 passed, 5 skipped, 89 warnings, 4 rerun in 685.05s (0:11:25) ==

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci-refresh-hf-cache ciflow/trunk CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. module: arm Issues related to arm backend partner: arm For backend delegation, kernels, demo, etc. from the 3rd-party partner, Arm

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants