Conversation
The mt-l-x86iavx512-8-64 OSDC runner mounts HF_HOME read-only at /mnt/hf_cache, so the prequantized NSS/NFRU tests added in pytorch#23236 and pytorch#23237 fail with EROFS when hf_hub_download fetches checkpoints that are not already cached. Redirect HF_HOME to a writable dir, matching cuda-perf.yml and cuda-windows.yml. Authored with Claude Code.
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23268
Note: Links to docs will display an error until the docs builds have been completed. ❌ 1 Cancelled Job, 6 Unrelated FailuresAs of commit 33b29c9 with merge base 273cb33 ( CANCELLED JOB - The following job was cancelled. Please retry:
FLAKY - The following jobs failed but were likely due to flakiness present on trunk:
BROKEN TRUNK - The following jobs failed but were present on the merge base:👉 Rebase onto the `viable/strict` branch to avoid these failures
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This PR needs a
|
zingo
left a comment
There was a problem hiding this comment.
Nice, superthanks for the help!
We migth want/need to do that on all test-arm-backend-* based job in pull and trunk also.
or not? :
|
yeah, you're right for trunk. test-arm-backend-vkml in trunk.yml runs on an OSDC runner and picks up the *_prequantized_vgf_INT tests, so it'll hit the same EROFS, i think. I'll add the redirect there and to trunk test-arm-backend-ethos-u (same runner and HF setup). The pull vkml job is on linux.2xlarge.memory, and the other Arm jobs don't touch HF, so I'm leaving those alone. |
The trunk vkml job also runs on an OSDC runner, and test_pytest_models_vkml picks up the prequantized NSS/NFRU VGF tests, so it hits the same EROFS on /mnt/hf_cache. Authored with Claude Code.
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
HF_HUB_CACHE remains pointed at the read-only mount, so hf_hub_download() can still fail.
Review effort: Balanced
Findings: 2
Open (2)
What changed in this PR
Redirects Arm CI Hugging Face downloads to writable temporary caches.
Changes:
- Adds
RUNNER_TEMPcache paths with/tmpfallback. - Applies the workaround to pull and trunk Arm jobs.
| File | Description |
|---|---|
.github/workflows/pull.yml |
Updates the no-driver Arm job cache. |
.github/workflows/trunk.yml |
Updates the VKML Arm job cache. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| export HF_HOME="${RUNNER_TEMP:-/tmp}/hf_cache" | ||
| mkdir -p "${HF_HOME}" 2>/dev/null || export HF_HOME=/tmp/hf_cache | ||
| mkdir -p "${HF_HOME}" |
| export HF_HOME="${RUNNER_TEMP:-/tmp}/hf_cache" | ||
| mkdir -p "${HF_HOME}" 2>/dev/null || export HF_HOME=/tmp/hf_cache | ||
| mkdir -p "${HF_HOME}" |
linux_job_v3 already sets a writable HF_HOME and points HF_HUB_CACHE at the read-only /mnt/hf_cache/hub, which hf_hub_download prefers, so the redirect had no effect. Missing checkpoints are seeded via the ci-refresh-hf-cache label. Authored with Claude Code.
…psiddh/executorch into sidart/arm-ci-writable-hf-cache
There was a problem hiding this comment.
Copilot wasn't able to review any files in this pull request. Check if the Files changed in this pull request are included in default exclusions.
Still open (2)
- 🔴 .github/workflows/trunk.yml#L349 — This
linux_job_v3.yml@​mainjob inheritsHF_HUB_CACHE=/mnt/hf_cache/hub. Because… - 🔴 .github/workflows/pull.yml#L873 —
linux_job_v3.yml@​mainexportsHF_HUB_CACHE=/mnt/hf_cache/hubbefore this script runs.…
|
This seems to fix that problem |

The
mt-l-x86iavx512-8-64OSDC runner mountsHF_HOMEread-only at/mnt/hf_cache. The prequantized NSS/NFRU tests added in #23236 and #23237 callhf_hub_download()for checkpoints that aren't in that cache, so they fail with:This redirects
HF_HOMEto a writable dir intest-arm-backend-no-driver, usingthe same workaround as
cuda-perf.ymlandcuda-windows.yml.test-arm-backend-vkmlruns on a different runner type and doesn't need it.The job no longer uses the shared read-only cache, so HF assets are downloaded
each run.
Test plan:
test_nfru_prequantized_tosa_INTshould pass intest-arm-backend-no-driver (test_pytest_models_tosa).The change: it's the six lines from cuda-perf.yml, added after conda activate in
the no-driver job only. The workflow YAML parses; actionlint isn't installed here,
so I couldn't run it. Nothing has run it in CI yet, so this PR's
test_pytest_models_tosa run is the real test.
Watch the runtime: that job took about 13 minutes in the last run. If
re-downloading the HF files makes it much slower, the better long-term fix is
asking test-infra to add the new checkpoints to the shared cache.
cc @digantdesai @freddan80 @per @zingo @oscarandersson8218 @mansnils @Sebastian-Larsson @robell @rascani