Skip to content

Name the dev dependency groups after the vLLM line they install - #138

Draft
ItzikVa wants to merge 3 commits into
mainfrom
bugfix/gpu-leg-dep-group-name
Draft

ItzikVa wants to merge 3 commits into
mainfrom
bugfix/gpu-leg-dep-group-name

Conversation

@ItzikVa

@ItzikVa ItzikVa commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator

The vllm26 GPU leg has failed before pytest on every run since #130, and not for any reason visible in this repo. The cluster harness derives the vLLM line it asserts against from the dependency-group NAME it is handed via --dep-group, and it maps the bare name dev to a frozen 0.19.x expectation:

dep group: dev          (job tag v19, expect vllm 0.19.x)
dep group: dev-vllm20   (job tag v20, expect vllm 0.20.x)

#130 moved dev onto the 0.26 line but left the matrix passing dev, so the pod installed 0.26 correctly and then killed itself on the harness's guard:

torch 2.11.0+cu130 vllm 0.26.0
FATAL expected vllm 0.19.x for group dev but got 0.26.0

That is an 83-second failure with no pytest output, which is why it reads as infrastructure rather than as a test result. The vllm27 leg was never affected -- dev-vllm27 names its line, so the guard derives 0.27.x and agrees.

Rename dev to dev-vllm26 so every dev group states its line, and update the matrix, the CPU job's sync, CONTRIBUTING.md and docs/AUDIO.md to match. dev is removed rather than kept as an alias so the name that silently means "0.19.x" cannot be passed again by accident.

Rename rather than addition for a second reason: an extra conflicting group multiplies uv's resolution-marker space, costing +3093 net lines of uv.lock for a provably identical resolution. The rename is diff-balanced (3493/3493), and the resolved set is unchanged either way -- 275 packages, zero version differences, verified by comparing every (name, version) pair before and after. dev-vllm26 exports vllm==0.26.0 / torch==2.11.0 / transformers==5.17.0 / pytest==9.0.3, dev-vllm27 still exports vllm==0.27.1, and --group dev is now a hard error.

Note for the next line added: the harness's mapping is keyed on the group name, so a dev-vllm30 leg needs the same dev-vllmNN form to resolve.

The vllm26 GPU leg has failed before pytest on every run since #130, and not
for any reason visible in this repo. The cluster harness derives the vLLM line
it asserts against from the dependency-group NAME it is handed via
--dep-group, and it maps the bare name `dev` to a frozen 0.19.x expectation:

    dep group: dev          (job tag v19, expect vllm 0.19.x)
    dep group: dev-vllm20   (job tag v20, expect vllm 0.20.x)

#130 moved `dev` onto the 0.26 line but left the matrix passing `dev`, so the
pod installed 0.26 correctly and then killed itself on the harness's guard:

    torch 2.11.0+cu130 vllm 0.26.0
    FATAL expected vllm 0.19.x for group dev but got 0.26.0

That is an 83-second failure with no pytest output, which is why it reads as
infrastructure rather than as a test result. The vllm27 leg was never affected
-- `dev-vllm27` names its line, so the guard derives 0.27.x and agrees.

Rename `dev` to `dev-vllm26` so every dev group states its line, and update
the matrix, the CPU job's sync, CONTRIBUTING.md and docs/AUDIO.md to match.
`dev` is removed rather than kept as an alias so the name that silently means
"0.19.x" cannot be passed again by accident.

Rename rather than addition for a second reason: an extra conflicting group
multiplies uv's resolution-marker space, costing +3093 net lines of uv.lock
for a provably identical resolution. The rename is diff-balanced
(3493/3493), and the resolved set is unchanged either way -- 275 packages,
zero version differences, verified by comparing every (name, version) pair
before and after. `dev-vllm26` exports vllm==0.26.0 / torch==2.11.0 /
transformers==5.17.0 / pytest==9.0.3, `dev-vllm27` still exports vllm==0.27.1,
and `--group dev` is now a hard error.

Note for the next line added: the harness's mapping is keyed on the group
name, so a `dev-vllm30` leg needs the same `dev-vllmNN` form to resolve.

Signed-off-by: itzikvaisman <yizhak.vaisman@ibm.com>
The previous commit's comments said the cluster harness derives the expected
vLLM line from the dep-group NAME. It does not. Run logs show it looks the
name up in a table of its own:

    dep group: dev          (job tag v19,         expect vllm 0.19.x)
    dep group: dev-vllm20   (job tag v20,         expect vllm 0.20.x)
    dep group: dev-vllm27   (job tag dev-vllm27,  expect vllm any.x)

The third line is the one that matters: an unknown name is not an error, it
falls back to `tag = the name itself, expect any.x` -- no assertion. Only names
the table knows carry a version claim, and `dev`'s was frozen at 0.19.x.

This makes the rename a stronger fix than described, not a weaker one. It does
not rely on the harness parsing "26" out of `dev-vllm26` -- it cannot, the
table is explicit -- it simply stops addressing the one stale entry and lands
on the no-assertion fallback. A future `dev-vllm30` leg therefore needs no
change on the image side either, which the previous message flagged as an open
risk.

It also surfaces a trade-off worth recording: `any.x` means no leg asserts its
vLLM line any more. That assertion was already absent for dev-vllm27 and was
wrong for dev, so nothing working is lost, but the real follow-up is for the
harness table to gain genuine dev-vllm26/27/30 entries. Until then the
per-line `vllm26`/`vllm27` version bounds in pyproject.toml are what pin the
line.

Comments only -- no dependency, workflow or lockfile change. Verified by
re-running `uvx uv@0.8.4 lock`, which left uv.lock byte-identical.

Signed-off-by: itzikvaisman <yizhak.vaisman@ibm.com>
CPU Tests went red on this branch's first commit while the suite itself passed.
Step-level conclusions, identical on 3.11 and 3.12:

    success   Run uv sync --frozen --group dev-vllm26 --extra hf --extra compose
    success   Run CPU tests
    failure   Upload coverage

codecov-action@v4 exits nonzero on a token/rate-limit problem of its own BEFORE
it consults `fail_ci_if_error: false`, so the flag that already declared
"upload problems must not gate CI" never gets a chance to apply. Moved to v5
(v4 is deprecated) and added `continue-on-error: true`, which enforces that
intent at the workflow level where the action cannot override it.

Not masking a failure: `tests/unit/` is untouched here, the install step
resolved as intended, and the test step is green on both interpreters. The
same fix is already on #135, where v5 then uploaded successfully -- evidence
that v4 itself was at fault rather than the token. It is duplicated here
because this PR branches off main and must be able to go green on its own;
whichever lands first makes the other's hunk a no-op.

This also confirms the rename in the previous commits: `uv sync --frozen
--group dev-vllm26` is a passing step in real CI on both Python versions.

Signed-off-by: itzikvaisman <yizhak.vaisman@ibm.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant