Skip to content

test vLLM 0.30 support - #135

Draft
ItzikVa wants to merge 8 commits into
mainfrom
test-vllm-0.30
Draft

ItzikVa wants to merge 8 commits into
mainfrom
test-vllm-0.30

Conversation

@ItzikVa

@ItzikVa ItzikVa commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator

No description provided.

Signed-off-by: itzikvaisman <yizhak.vaisman@ibm.com>
A four-leg GPU matrix (one vLLM line each) found unit/composer/hf green
everywhere and the vLLM integration suite red on all four: 7 failures on
0.27.1, 9 on 0.28.0, 16 on 0.29.0, 14 on 0.30.0. Two of those were real bugs
in src/, live today on 0.27+ regardless of this effort; the rest was test
scaffolding written against APIs that 0.29 deleted.

Where the lines differ, gate on the capability (hasattr) rather than on a
version number: that states the actual requirement and survives the next
rename.

src/ fix 1 -- SR MoE router padding mask
  vLLM 0.27 started passing an is_padding mask to the topk kernels, read off
  the forward context and sized to the padded token count. The SR layer routes
  its [2M, H] base+adapter stack in one expert call, so the kernel saw 2M=16384
  gating rows against an 8192-entry mask and the engine died at init with
  "is_padding size mismatch, expected: 16384". Double the mask alongside the
  rows: adapter row i is the same token as base row i, so it is padding exactly
  when the base row is. Inert with no forward context, no mask
  (VLLM_MOE_SKIP_PADDING=0, or 0.26 where no kernel consumes one), or a mask
  whose length is not m. CI missed this because test_sr_tp_equivalence.py is
  env-gated on SR_COMPOSED_DIR and never runs -- a follow-up.

src/ fix 2 -- audio ASR never ran on 0.29+
  0.29 deleted BaseMultiModalProcessor._call_hf_processor along with
  _apply_hf_processor_text_mm, its only caller, and narrowed
  _apply_hf_processor_main to (mm_items, hf_processor_mm_kwargs) ->
  BatchFeature. Our _call_hf_processor override is where ASR happens, so on
  0.29/0.30 it was simply never called and no transcript was ever produced --
  silently broken, not merely untested. Factor the ASR work into
  _audio_features() and add an _apply_hf_processor_main override that delegates
  to super() while the base class still has the old hook, and otherwise returns
  the audio tensors alone (on 0.29+ the prompt is vLLM's business: tokenized
  upstream, then _postprocess_prompt). **kwargs because both vLLM call sites
  pass keywords only and no single signature matches both lines.
  _hf_processor_applies_updates stays as-is -- required on 0.26-0.28, inert on
  0.29+ where _apply_prompt_updates always runs, which is the False behaviour
  it was asking for.

tests -- KV cache shape
  AttentionBackend.get_kv_cache_shape is gone at 0.29 (29 defs at 0.28, 0 at
  0.29). The three call sites build a real KV cache fed to a real
  FlashAttention kernel, so the shape has to be genuinely right. New
  tests/shared/vllm_kv_cache.py prefers compute_layer_kv_cache_shape_bytes
  (0.29+) and falls back to the old staticmethod, cross-checking the two where
  both exist. They are provably equal for dense bf16: num_heads == num_kv_heads,
  one kernel state per token, and content bytes == 2 * head_size * itemsize.

tests -- audio processor
  _apply_uncached grows a second shape for 0.29+, which split text processing
  away from MM processing; everything after the first call
  (_get_mm_fields_config -> _get_mm_prompt_updates ->
  _maybe_apply_prompt_updates) is unchanged on every line. The one test that
  asserted on the deleted tri-tuple's flag is skipped there with a reason, and
  the class docstring records that the hook test is vacuous at 0.29+.

tests -- equivalence harness
  run_vllm_logprobs is the single funnel for both engines in all five runners,
  so setdefault there applies symmetrically: enforce_eager (an equivalence test
  must not measure the compiler), enable_prefix_caching=False (upstream
  declares IsHybrid and we deliberately do not, so the two sides otherwise
  resolve different defaults -- and it dodges a 0.28-only "no mamba layers in
  the model" assertion in upstream's own model), and a free-memory-aware
  gpu_memory_utilization. The last reuses the pattern already solved in
  _multi_adapter_equivalence_worker.py, lifted to tests/shared/vllm_gpu_mem.py:
  vLLM measures the fraction against TOTAL memory, these suites start engines
  back to back, and the guard is byte-identical back to 0.26 -- a pre-existing
  harness defect the newer legs merely exposed. _granite4_mini_tests.py's
  per-config eager opt-in is now redundant and deleted rather than left to
  contradict the default.

pyproject -- the [vllm] extra claimed 0.30 only while [tool.uv]
default-groups said vllm26, which made uv sync unsatisfiable. Widen it to
>=0.26,<0.31, prune the duplicate [vllm27] extra (the dependency group covers
pinning), and drop the group-vs-extra and extra-vs-extra conflict entries that
only existed to describe the contradiction. uv.lock relocked accordingly.

Signed-off-by: itzikvaisman <yizhak.vaisman@ibm.com>
Two defects that kept the 0.29/0.30 legs of the support matrix red. Both are
latent on the lines already supported, so neither is a cost of widening the
range.

Audio prompt updates targeted the marker STRING. vLLM narrowed the accepted
target type at 0.29:

    0.26 - 0.28   UpdateTarget = PromptSeq | PromptIndex   # str allowed
    0.29 - 0.30   UpdateTarget = list[int] | PromptIndex   # str dropped

Passing a str on 0.29+ does not fail where it is given. Token matching simply
finds nothing, which drops vLLM into _apply_prompt_updates_via_text -- a
fallback that only exists from 0.29 -- and that hands the raw string to
tokenizer.decode(). The error then surfaces far from its cause and reads as a
tokenizer bug: "TypeError: Can't extract `str` to `Vec`" out of the Rust
tokenizer, killing EngineCore during init rather than rejecting one request.
That accounts for 5 test_audio_processor failures and all 13 audio integration
errors -- one bug, two symptoms. A list[int] target is legal on all five lines,
so the fix needs no version branch, and _marker_id() already existed.

Engine teardown never waited. The suites start two engines back to back and the
second was measuring a card the first still owned: 0.85 x 79.3 = 67.4 GiB taken,
10.1 GiB left, gpu_memory_utilization=0.115, below the floor. The existing
"del llm; gc.collect(); torch.cuda.empty_cache()" cannot fix that, because the
weights and KV cache belong to the EngineCore CHILD process -- empty_cache()
drains the parent's allocator, which never held them -- and teardown is a
weakref.finalize, so del schedules it rather than performing it.

shutdown_llm() calls engine_core.shutdown(), which chains to
CoreEngineProcManager.shutdown and terminates AND joins the children, then polls
mem_get_info until the driver has reclaimed. The shutdown(timeout=None)
signature is byte-identical on 0.26 - 0.30 and is reached defensively. On
timeout it prints and returns rather than raising, leaving gpu_mem_util as the
single place that rules a card too full to use.

This is why those 10 equivalence failures were identical on all four legs: a
pre-existing harness defect the newer lines merely exposed, not an API break.

Signed-off-by: itzikvaisman <yizhak.vaisman@ibm.com>
_doubled_padding_mask reached ForwardContext.is_padding by attribute access.
That field was added at vLLM 0.25.1 and does not exist at 0.24.0, so every SR
forward raised

    AttributeError: 'ForwardContext' object has no attribute 'is_padding'

and took EngineCore down during init:

    RuntimeError: Engine core initialization failed.

Measured on itzikv-eval-v24-g5-sr (vllm 0.24.0, torch 2.11.0, transformers
5.17.0 -- the install itself resolved cleanly), 2026-10-04.

The docstring already claimed inertness "where there is nothing to match", and
a missing field is the clearest such case, so this is the code catching up to
its own contract rather than a new behaviour. getattr, not a version compare:
the requirement is "this build has the field", which is what the capability
check states.

Not a regression on any supported line -- 0.25.1 through 0.30 all have the
field, and SR scored 0.8960 against its 0.8948 bar on 0.30 with the attribute
access in place. It only blocks going BELOW the current floor, which is what
the 0.24 probe was submitted to find out.

Signed-off-by: itzikvaisman <yizhak.vaisman@ibm.com>
7accc4c added three setdefaults to run_vllm_logprobs. Two fixed observed
failures. The third, enforce_eager=True, was added on the general principle
that "an equivalence test must not measure the compiler", and it turned
TestZeroAdapterNoHiding::test_no_control_tokens red on every vLLM line from
0.26 to 0.30 -- the last failure in the four-leg matrix.

Pinned in one pod (vLLM 0.26.0, sha 5df1b8f, same venv, that setdefault the
only variable):

    enforce_eager=True     rc=1   3 failed, 12 deselected   4:26
    enforce_eager=False    rc=0   3 passed, 12 deselected   6:35

and the test was green at 5c0c72d, where the mini suite ran compiled
(_UPSTREAM_EAGER_CONFIGS keyed on "4.0-h-350m", which has never been a key of
GRANITE4_MINI, so _eager_kwargs_if_needed always returned {} -- it was dead
from the initial commit and is not restored here). granite4_equivalence.py is
byte-identical over 5c0c72d..5df1b8f, so the tolerance never moved.

Measured the floor rather than assuming it, because "the gate is below the
eager noise floor" and "there is a real divergence the compiler was hiding"
both predict the above and want opposite patches. Eager, on the zero-LoRA pair
the test builds:

               upstream vs upstream   switch vs switch   switch vs upstream
    4.0-1b          0.0000e+00          0.0000e+00          9.9182e-04
    4.0-350m        0.0000e+00          0.0000e+00          1.9522e-03
    4.0-micro       0.0000e+00          0.0000e+00          5.0116e-04

Each engine is bit-exact against a second instance of itself over all 3840
elements, so there is no noise floor: the divergence is real and deterministic,
and the switch/upstream column reproduces the three failures' worst diffs to
five digits.

It appears only when the adapter tier is live. save_upstream_model builds
num_adapters=0, where the projection is the plain base linear and eager is
bit-exact; save_zero_adapter_model builds num_adapters=2, where it runs the
fused SWITCH kernel plus a zero-valued shrink/expand -- a different op sequence
with a different reduction order. get_tolerances' 4.77e-7 basis for the 1e-5
inert gate was measured at num_adapters=0 and compiled, so it never
characterized this configuration. torch.compile lowers both sequences to
equivalent kernels and normalizes the ordering, which is exactly what an
equivalence test wants canonicalized; it preserves semantics, so a genuine
kernel bug still trips 1e-5. Compiled is also how vLLM serves.

So the choice belongs to each test. _granite4_fullsize_tests.py keeps
enforce_eager=True explicitly -- at full size, with num_adapters=0, eager is
bit-exact and an unpinned autotuner reddened it ~1 run in 7 -- and setdefault
was never what gave it that. enable_prefix_caching=False and
gpu_memory_utilization stay: they fix the vLLM 0.28 mamba assertion and the
in-pod OOM respectively.

No tolerance was changed. The measurements are recorded at all three sites so
the next reader does not re-derive them, or repeat the mistake.

Signed-off-by: itzikvaisman <yizhak.vaisman@ibm.com>
The two GPU legs now install the two ENDS of the supported range, 0.26 and
0.30, instead of 0.26 and 0.27.

WHY THE ENDS. pyproject declares vllm>=0.26,<0.31, and every break this range
actually contains sits at a boundary. APIs deleted at 0.29 only fail at the
ceiling: AttentionBackend.get_kv_cache_shape (29 defs at 0.26-0.28, zero at
0.29) and BaseMultiModalProcessor._call_hf_processor (the method that ran ASR,
so audio was silently producing no transcript). An API that does not exist yet
only fails at the floor: ForwardContext.is_padding, added at 0.27, which the SR
router slices. A 0.27 leg is bracketed by two tested lines and would have
caught none of them. 0.27/0.28/0.29 were verified green out of band by a
five-leg matrix across every suite, so this is a per-run cost decision, not a
claim that the middle is untested.

Measured at 8a4d69d, one vLLM line per pod, VLLM_USE_FLASHINFER_SAMPLER=0:

  leg   vLLM     unit    composer  hf     vllm        integration
  v26   0.26.0   1739p   459p      449p   220p/4s     21p/9s
  v27   0.27.1   1739p   459p      449p   220p/4s     21p/9s
  v28   0.28.0   1739p   459p      449p   220p/4s     21p/9s
  v29   0.29.0   1739p   459p      449p   219p/5s     21p/9s
  v30   0.30.0   1739p   459p      449p   219p/5s     21p/9s

rc=0 on all 25 leg/suite combinations, each with a nonzero passed count --
checked that way because pytest exits 0 on an all-skipped run. The 219-vs-220
gap is one deliberate skip: test_uncached_path_reports_updates_not_applied
asserts a contract 0.29 removed. The other four skips are identical on every
leg and env-gated.

VLLM_USE_FLASHINFER_SAMPLER=0 is now set in tests/conftest.py, because raising
the ceiling is what exposes the need for it. On an image shipping CUDA 12.4 the
0.30 engine failed to start 171 times on

    nvcc fatal : Unknown option '--compress-mode=size'

flashinfer JIT-compiles its sampling kernels with --compress-mode, an nvcc
option added in CUDA 12.8. It is version-dependent: 0.26 is green on that image
without the flag and 0.30 is not, so a 0.26-only CI never needed it. The pod
template lives in the runner image rather than this repo, so conftest is the
only repo-side lever. setdefault, so an image with CUDA >= 12.8 can export
VLLM_USE_FLASHINFER_SAMPLER=1 and exercise that path for real; a newer toolkit
in the image is the actual fix. Scoped to the sampler and inert for what these
suites assert -- the attention backend resolves to FlashAttention
independently, and the equivalence tests compare logprobs, produced before any
sampling op runs.

uv.lock is regenerated: vllm 0.27.1 -> 0.30.0. It resolves transformers 5.17.0
and torch 2.11.0 for the 0.26 leg / 2.13.0 for the 0.30 leg, matching what the
matrix pods resolved through pip.

Signed-off-by: itzikvaisman <yizhak.vaisman@ibm.com>
The vLLM 0.26-0.30 widening re-locked uv.lock, and the re-lock was done with
the local uv (0.7.19, Homebrew). That binary writes `revision = 2` in the
lockfile header, so the commit silently DOWNGRADED the field from the
`revision = 3` already on the branch.

.pre-commit-config.yaml pins uv-pre-commit at rev 0.8.4, whose uv writes
`revision = 3`. So CI's `uv-lock` hook re-upgraded the field, reported
"pre-commit hook(s) made changes", and exited 1 on a one-line diff:

    -revision = 2
    +revision = 3

Re-locked with `uvx uv@0.8.4 lock` so the committed file matches what the
pinned hook produces. The hook is now a no-op.

RESOLUTION IS UNCHANGED, which is why this is header-only: the diff against
the previous commit is exactly one line. 0.8.4 resolves the same 287 packages
(the count CI printed), and the pins this effort cares about all hold -- vllm
0.26.0 for the dev/test/vllm26 groups and 0.30.0 for dev-vllm30/vllm30,
transformers 5.17.0, torch 2.11.0 on the 0.26 leg and 2.13.0 on the 0.30 leg.
Those are the versions the five-leg GPU matrix actually installed, so the
lockfile still describes the configuration that was measured green.

`revision` is a lockfile FORMAT marker, not a dependency, so bumping it cannot
move a version. The rule it encodes is that the file must be written by a uv at
least as new as the hook's pin; writing it with an older one is what
regressed it.

Verified the rest of the hook list separately, since fail_fast: true means the
CI log only proves uv-lock failed: uv-lock is hook 15 of 15, so CI reaching it
already implies 1-14 passed. Confirmed directly anyway -- ruff 0.9.0 format
(223 files) and check clean, check_headers clean, validate_links clean.

Signed-off-by: itzikvaisman <yizhak.vaisman@ibm.com>
@ItzikVa

ItzikVa commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator Author

/gpu-test-multi

@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

✅ GPU tests passed — vllm27-multi

2888 passed, 66 skipped, 20 warnings in 10882.25s (3:01:22)

Commit: 7e6c2037253c693a24fef31d64d9776b3187d57d
Full run & artifact log

Last 40 log lines
tests/integration/test_multi_switch_vllm_generate.py::test_different_adapters_produce_different_text SKIPPED

=============================== warnings summary ===============================
<frozen importlib._bootstrap>:488
  <frozen importlib._bootstrap>:488: DeprecationWarning: builtin type SwigPyPacked has no __module__ attribute

<frozen importlib._bootstrap>:488
  <frozen importlib._bootstrap>:488: DeprecationWarning: builtin type SwigPyObject has no __module__ attribute

.venv/lib/python3.12/site-packages/torch/jit/_script.py:365: 14 warnings
  /tmp/granite-switch/.venv/lib/python3.12/site-packages/torch/jit/_script.py:365: DeprecationWarning: `torch.jit.script_method` is deprecated. Please switch to `torch.compile` or `torch.export`.
    warnings.warn(

tests/composer/test_compose_e2e.py:83
  /tmp/granite-switch/tests/composer/test_compose_e2e.py:83: PytestUnknownMarkWarning: Unknown pytest.mark.xdist_group - is this a typo?  You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html
    @pytest.mark.xdist_group("compose_e2e")

tests/composer/test_multi_audio_compose_e2e.py:124
  /tmp/granite-switch/tests/composer/test_multi_audio_compose_e2e.py:124: PytestUnknownMarkWarning: Unknown pytest.mark.xdist_group - is this a typo?  You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html
    @pytest.mark.xdist_group("multi_audio_compose_e2e")

tests/composer/test_multi_audio_compose_e2e.py:130
  /tmp/granite-switch/tests/composer/test_multi_audio_compose_e2e.py:130: PytestUnknownMarkWarning: Unknown pytest.mark.xdist_group - is this a typo?  You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html
    @pytest.mark.xdist_group("multi_audio_compose_e2e")

tests/composer/test_upstream_files.py:218
  /tmp/granite-switch/tests/composer/test_upstream_files.py:218: PytestUnknownMarkWarning: Unknown pytest.mark.xdist_group - is this a typo?  You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html
    @pytest.mark.xdist_group("upstream_build_e2e")

-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
========= 2888 passed, 66 skipped, 20 warnings in 10882.25s (3:01:22) ==========
sys:1: DeprecationWarning: builtin type swigvarlink has no __module__ attribute
[rank0]:[W1005 10:41:00.896837074 ProcessGroupNCCL.cpp:1624] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
===== ALL GPU TESTS PASSED =====
[10:41:21] <job> Succeeded
[10:41:25] verified: found success sentinel in pod log
[10:41:25] cleanup: deleting <job> (exit=0)
<job> "<job>" deleted

Verdict: PASSED

@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

❌ GPU tests failed — vllm26-multi

No pytest summary — the suite did not start. See the log below.

Commit: 7e6c2037253c693a24fef31d64d9776b3187d57d
Full run & artifact log

Last 40 log lines
GPU test report — vllm26-multi
-------------------------------------------------------------------------
Commit:    7e6c2037253c693a24fef31d64d9776b3187d57d
Suite:     tests/unit/ tests/hf/ tests/composer/ tests/vllm/ tests/integration/
Deps:      dev (vllm 0.19.x)
GPU:       NVIDIA A100-SXM4-80GB

The test suite never started — the run failed during setup.
Redacted tail of the run log:

  [07:39:46] <job> Failed
  [07:39:47] --- post-mortem: why the job ended (status and events only) ---
  [07:39:47] <job> status:
    QuotaReserved=True Resuming: Suspend is false
    ResourcesDeployed=True Resuming: Suspend is false
    PodsReady=True SufficientPodsReady: 1 pods running; 0 pods succeeded
    Unhealthy=False Resuming: Suspend is false
  [07:39:48] pod <job> phase and conditions:
    phase=Failed
    Initialized=True : 
    Ready=False PodFailed: 
    ContainersReady=False PodFailed: 
    PodScheduled=True : 
  [07:39:48] pod <job> container states:
    pytorch: ready=false restarts=0 state={"terminated":{"containerID":"cri-o://c7b58e925a537f2f1de033562bce1368cdeda6d45374a6c4e5d9abfba1a8abec","exitCode":1,"finishedAt":"2026-10-05T07:39:42Z","reason":"Error","startedAt":"2026-10-05T07:38:19Z"}}
  [07:39:49] kubelet events for <job>:
  TIME                   TYPE     REASON           MESSAGE
  <nil>                  Normal   Scheduled        Successfully assigned security/<job> to dmf-nnnqh-gpu-worker-3-pbkts
  2026-10-05T07:38:18Z   Normal   AddedInterface   Add eth0 [10.134.0.206/23] from ovn-kubernetes
  2026-10-05T07:38:18Z   Normal   Pulled           Container image "vllm/vllm-openai:latest" already present on machine
  2026-10-05T07:38:19Z   Normal   Created          Created container pytorch
  2026-10-05T07:38:19Z   Normal   Started          Started container pytorch
  [07:39:50] --- end post-mortem ---
  [07:39:50] cleanup: deleting <job> (exit=1)
  <job> "<job>" deleted

Verdict: FAILED before the test suite ran

Two CI reds on this PR, NEITHER of which was a test failure.

CPU TESTS: the suite passed. Step-level conclusions on both 3.11 and 3.12 were
`success` for "Run CPU tests" and `failure` for the next step, "Upload
coverage". codecov-action@v4 exits nonzero on a token/rate-limit problem of its
own BEFORE it consults `fail_ci_if_error: false`, so the flag that already
declared "upload problems must not gate CI" could not take effect. Moved to v5
(v4 is deprecated) and added `continue-on-error: true`, which enforces that
intent at the workflow level where the action cannot override it. This is not
masking a failure: `tests/unit/` is untouched by this branch, and the install
step resolved exactly as intended (vllm 0.26.0 / torch 2.11.0 /
transformers 5.17.0). The same ci.yaml passed on main on 2026-09-30 and fails
now, so the trigger is external drift.

GPU TESTS: a chicken-and-egg in the dispatcher. `workflow_dispatch` always
executes the workflow file from the DEFAULT BRANCH, so `/gpu-test-multi` ran
main's matrix -- legs `vllm26 -> dev` and `vllm27 -> dev-vllm27` -- while
checking out this PR's commit (the uploaded artifacts are named for sha
8276623). This branch had deleted `dev-vllm27`, so that leg's
`uv sync --group dev-vllm27` had no such group and died before pytest started;
both logs are under 1.1 KB, the "suite never started" signature.

Restored `vllm27` / `dev-vllm27` so main's dispatched matrix stays resolvable
against this tree until the 0.26/0.30 rename lands on main. The group is not a
fiction -- the `vllm` extra claims 0.26-0.30, so 0.27 is a supported line that
simply is not worth its own per-run leg -- and it is removable once main
carries the rename.

Conflicts now enumerate every cross-line pair: three lines with a bare and a
dev- flavour each give 3 x (2 x 2) = 12. `test` stays absent because it reaches
vllm26 through `include-group = "dev"` and uv derives that transitively, which
is also why main's two-line version never listed it.

DELIBERATELY NOT CHANGING `default-groups = ["vllm26"]`, though
`uv sync --group dev-vllm30` reports it as conflicting. That shape is not new:
at 6013c7f `default-groups` was `["vllm19"]` against a declared
`dev-vllm20`/`vllm19` conflict, and that run's vllm20 leg PASSED -- so the
cluster harness already excludes default groups. Changing it here would be a
guess against evidence.

Verified: `uv lock` is a no-op under the hook's pinned uv 0.8.4 (revision 3);
each leg resolves to its own line (dev 0.26.0, dev-vllm27 0.27.1,
dev-vllm30 0.30.0); ruff 0.9.0 format and check clean; pyproject and both
workflows parse.

Signed-off-by: itzikvaisman <yizhak.vaisman@ibm.com>
@ItzikVa

ItzikVa commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator Author

/gpu-test-multi

ItzikVa added a commit that referenced this pull request Oct 5, 2026
CPU Tests went red on this branch's first commit while the suite itself passed.
Step-level conclusions, identical on 3.11 and 3.12:

    success   Run uv sync --frozen --group dev-vllm26 --extra hf --extra compose
    success   Run CPU tests
    failure   Upload coverage

codecov-action@v4 exits nonzero on a token/rate-limit problem of its own BEFORE
it consults `fail_ci_if_error: false`, so the flag that already declared
"upload problems must not gate CI" never gets a chance to apply. Moved to v5
(v4 is deprecated) and added `continue-on-error: true`, which enforces that
intent at the workflow level where the action cannot override it.

Not masking a failure: `tests/unit/` is untouched here, the install step
resolved as intended, and the test step is green on both interpreters. The
same fix is already on #135, where v5 then uploaded successfully -- evidence
that v4 itself was at fault rather than the token. It is duplicated here
because this PR branches off main and must be able to go green on its own;
whichever lands first makes the other's hunk a no-op.

This also confirms the rename in the previous commits: `uv sync --frozen
--group dev-vllm26` is a passing step in real CI on both Python versions.

Signed-off-by: itzikvaisman <yizhak.vaisman@ibm.com>
@aviv1ron1

Copy link
Copy Markdown
Collaborator

/gpu-test-dev

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

❌ GPU tests failed — vllm27-dev

No pytest summary — the suite did not start. See the log below.

Commit: 7e6c2037253c693a24fef31d64d9776b3187d57d
Full run & artifact log

Last 40 log lines
platform linux -- Python 3.12.3, pytest-9.0.3, pluggy-1.6.0 -- /tmp/granite-switch/.venv/bin/python
cachedir: .pytest_cache
rootdir: /tmp/granite-switch
configfile: pyproject.toml
plugins: cov-7.1.0, anyio-4.13.0
collecting ... collected 0 items

============================ no tests ran in 0.00s =============================
ERROR: file or directory not found: tests/vllm/test_single_switch.py

FATAL pytest failed
[14:49:14] <job> Failed
[14:49:15] --- post-mortem: why the job ended (status and events only) ---
[14:49:15] <job> status:
  QuotaReserved=True Resuming: Suspend is false
  ResourcesDeployed=True Resuming: Suspend is false
  PodsReady=True SufficientPodsReady: 1 pods running; 0 pods succeeded
  Unhealthy=False Resuming: Suspend is false
[14:49:17] pod <job> phase and conditions:
  phase=Failed
  Initialized=True : 
  Ready=False PodFailed: 
  ContainersReady=False PodFailed: 
  PodScheduled=True : 
[14:49:19] pod <job> container states:
  pytorch: ready=false restarts=0 state={"terminated":{"containerID":"cri-o://7e0714ff7e3276d1b8f8a2998f6f91903a127eae3ae54606c3319aa47df4a718","exitCode":1,"finishedAt":"2026-10-06T14:49:04Z","reason":"Error","startedAt":"2026-10-06T14:47:54Z"}}
[14:49:20] kubelet events for <job>:
TIME                   TYPE     REASON           MESSAGE
<nil>                  Normal   Scheduled        Successfully assigned security/<job> to dmf-nnnqh-gpu-worker-3-6r296
2026-10-06T14:44:35Z   Normal   AddedInterface   Add eth0 [10.138.32.128/23] from ovn-kubernetes
2026-10-06T14:44:35Z   Normal   Pulling          Pulling image "vllm/vllm-openai:latest"
2026-10-06T14:47:53Z   Normal   Pulled           Successfully pulled image "vllm/vllm-openai:latest" in 3m18.826473986s (3m18.826502229s including waiting)
2026-10-06T14:47:54Z   Normal   Created          Created container pytorch
2026-10-06T14:47:54Z   Normal   Started          Started container pytorch
[14:49:21] --- end post-mortem ---
[14:49:21] cleanup: deleting <job> (exit=1)
<job> "<job>" deleted

Verdict: FAILED — the test suite started but never reported a summary
         (killed mid-run: in-pod timeout, OOM, or a crashed process)

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

❌ GPU tests failed — vllm26-dev

No pytest summary — the suite did not start. See the log below.

Commit: 7e6c2037253c693a24fef31d64d9776b3187d57d
Full run & artifact log

Last 40 log lines
platform linux -- Python 3.12.3, pytest-9.0.3, pluggy-1.6.0 -- /tmp/granite-switch/.venv/bin/python
cachedir: .pytest_cache
rootdir: /tmp/granite-switch
configfile: pyproject.toml
plugins: cov-7.1.0, anyio-4.13.0
collecting ... collected 0 items

============================ no tests ran in 0.00s =============================
ERROR: file or directory not found: tests/vllm/test_single_switch.py

FATAL pytest failed
[14:49:17] <job> Failed
[14:49:19] --- post-mortem: why the job ended (status and events only) ---
[14:49:19] <job> status:
  QuotaReserved=True Resuming: Suspend is false
  ResourcesDeployed=True Resuming: Suspend is false
  PodsReady=True SufficientPodsReady: 1 pods running; 0 pods succeeded
  Unhealthy=False Resuming: Suspend is false
[14:49:21] pod <job> phase and conditions:
  phase=Failed
  Initialized=True : 
  Ready=False PodFailed: 
  ContainersReady=False PodFailed: 
  PodScheduled=True : 
[14:49:22] pod <job> container states:
  pytorch: ready=false restarts=0 state={"terminated":{"containerID":"cri-o://423e2d6de226c7c6d4ff7592efc2fd1020dbeb0310f2e0a8346f26bbaca2cacf","exitCode":1,"finishedAt":"2026-10-06T14:49:03Z","reason":"Error","startedAt":"2026-10-06T14:47:53Z"}}
[14:49:23] kubelet events for <job>:
TIME                   TYPE     REASON           MESSAGE
<nil>                  Normal   Scheduled        Successfully assigned security/<job> to dmf-nnnqh-gpu-worker-3-5jxkz
2026-10-06T14:44:33Z   Normal   AddedInterface   Add eth0 [10.135.5.123/23] from ovn-kubernetes
2026-10-06T14:44:33Z   Normal   Pulling          Pulling image "vllm/vllm-openai:latest"
2026-10-06T14:47:52Z   Normal   Pulled           Successfully pulled image "vllm/vllm-openai:latest" in 3m19.38155776s (3m19.381577125s including waiting)
2026-10-06T14:47:53Z   Normal   Created          Created container pytorch
2026-10-06T14:47:53Z   Normal   Started          Started container pytorch
[14:49:24] --- end post-mortem ---
[14:49:24] cleanup: deleting <job> (exit=1)
<job> "<job>" deleted

Verdict: FAILED — the test suite started but never reported a summary
         (killed mid-run: in-pod timeout, OOM, or a crashed process)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants