Repository navigation
Conversation
Signed-off-by: itzikvaisman <yizhak.vaisman@ibm.com>
A four-leg GPU matrix (one vLLM line each) found unit/composer/hf green everywhere and the vLLM integration suite red on all four: 7 failures on 0.27.1, 9 on 0.28.0, 16 on 0.29.0, 14 on 0.30.0. Two of those were real bugs in src/, live today on 0.27+ regardless of this effort; the rest was test scaffolding written against APIs that 0.29 deleted. Where the lines differ, gate on the capability (hasattr) rather than on a version number: that states the actual requirement and survives the next rename. src/ fix 1 -- SR MoE router padding mask vLLM 0.27 started passing an is_padding mask to the topk kernels, read off the forward context and sized to the padded token count. The SR layer routes its [2M, H] base+adapter stack in one expert call, so the kernel saw 2M=16384 gating rows against an 8192-entry mask and the engine died at init with "is_padding size mismatch, expected: 16384". Double the mask alongside the rows: adapter row i is the same token as base row i, so it is padding exactly when the base row is. Inert with no forward context, no mask (VLLM_MOE_SKIP_PADDING=0, or 0.26 where no kernel consumes one), or a mask whose length is not m. CI missed this because test_sr_tp_equivalence.py is env-gated on SR_COMPOSED_DIR and never runs -- a follow-up. src/ fix 2 -- audio ASR never ran on 0.29+ 0.29 deleted BaseMultiModalProcessor._call_hf_processor along with _apply_hf_processor_text_mm, its only caller, and narrowed _apply_hf_processor_main to (mm_items, hf_processor_mm_kwargs) -> BatchFeature. Our _call_hf_processor override is where ASR happens, so on 0.29/0.30 it was simply never called and no transcript was ever produced -- silently broken, not merely untested. Factor the ASR work into _audio_features() and add an _apply_hf_processor_main override that delegates to super() while the base class still has the old hook, and otherwise returns the audio tensors alone (on 0.29+ the prompt is vLLM's business: tokenized upstream, then _postprocess_prompt). **kwargs because both vLLM call sites pass keywords only and no single signature matches both lines. _hf_processor_applies_updates stays as-is -- required on 0.26-0.28, inert on 0.29+ where _apply_prompt_updates always runs, which is the False behaviour it was asking for. tests -- KV cache shape AttentionBackend.get_kv_cache_shape is gone at 0.29 (29 defs at 0.28, 0 at 0.29). The three call sites build a real KV cache fed to a real FlashAttention kernel, so the shape has to be genuinely right. New tests/shared/vllm_kv_cache.py prefers compute_layer_kv_cache_shape_bytes (0.29+) and falls back to the old staticmethod, cross-checking the two where both exist. They are provably equal for dense bf16: num_heads == num_kv_heads, one kernel state per token, and content bytes == 2 * head_size * itemsize. tests -- audio processor _apply_uncached grows a second shape for 0.29+, which split text processing away from MM processing; everything after the first call (_get_mm_fields_config -> _get_mm_prompt_updates -> _maybe_apply_prompt_updates) is unchanged on every line. The one test that asserted on the deleted tri-tuple's flag is skipped there with a reason, and the class docstring records that the hook test is vacuous at 0.29+. tests -- equivalence harness run_vllm_logprobs is the single funnel for both engines in all five runners, so setdefault there applies symmetrically: enforce_eager (an equivalence test must not measure the compiler), enable_prefix_caching=False (upstream declares IsHybrid and we deliberately do not, so the two sides otherwise resolve different defaults -- and it dodges a 0.28-only "no mamba layers in the model" assertion in upstream's own model), and a free-memory-aware gpu_memory_utilization. The last reuses the pattern already solved in _multi_adapter_equivalence_worker.py, lifted to tests/shared/vllm_gpu_mem.py: vLLM measures the fraction against TOTAL memory, these suites start engines back to back, and the guard is byte-identical back to 0.26 -- a pre-existing harness defect the newer legs merely exposed. _granite4_mini_tests.py's per-config eager opt-in is now redundant and deleted rather than left to contradict the default. pyproject -- the [vllm] extra claimed 0.30 only while [tool.uv] default-groups said vllm26, which made uv sync unsatisfiable. Widen it to >=0.26,<0.31, prune the duplicate [vllm27] extra (the dependency group covers pinning), and drop the group-vs-extra and extra-vs-extra conflict entries that only existed to describe the contradiction. uv.lock relocked accordingly. Signed-off-by: itzikvaisman <yizhak.vaisman@ibm.com>
Two defects that kept the 0.29/0.30 legs of the support matrix red. Both are
latent on the lines already supported, so neither is a cost of widening the
range.
Audio prompt updates targeted the marker STRING. vLLM narrowed the accepted
target type at 0.29:
0.26 - 0.28 UpdateTarget = PromptSeq | PromptIndex # str allowed
0.29 - 0.30 UpdateTarget = list[int] | PromptIndex # str dropped
Passing a str on 0.29+ does not fail where it is given. Token matching simply
finds nothing, which drops vLLM into _apply_prompt_updates_via_text -- a
fallback that only exists from 0.29 -- and that hands the raw string to
tokenizer.decode(). The error then surfaces far from its cause and reads as a
tokenizer bug: "TypeError: Can't extract `str` to `Vec`" out of the Rust
tokenizer, killing EngineCore during init rather than rejecting one request.
That accounts for 5 test_audio_processor failures and all 13 audio integration
errors -- one bug, two symptoms. A list[int] target is legal on all five lines,
so the fix needs no version branch, and _marker_id() already existed.
Engine teardown never waited. The suites start two engines back to back and the
second was measuring a card the first still owned: 0.85 x 79.3 = 67.4 GiB taken,
10.1 GiB left, gpu_memory_utilization=0.115, below the floor. The existing
"del llm; gc.collect(); torch.cuda.empty_cache()" cannot fix that, because the
weights and KV cache belong to the EngineCore CHILD process -- empty_cache()
drains the parent's allocator, which never held them -- and teardown is a
weakref.finalize, so del schedules it rather than performing it.
shutdown_llm() calls engine_core.shutdown(), which chains to
CoreEngineProcManager.shutdown and terminates AND joins the children, then polls
mem_get_info until the driver has reclaimed. The shutdown(timeout=None)
signature is byte-identical on 0.26 - 0.30 and is reached defensively. On
timeout it prints and returns rather than raising, leaving gpu_mem_util as the
single place that rules a card too full to use.
This is why those 10 equivalence failures were identical on all four legs: a
pre-existing harness defect the newer lines merely exposed, not an API break.
Signed-off-by: itzikvaisman <yizhak.vaisman@ibm.com>
_doubled_padding_mask reached ForwardContext.is_padding by attribute access.
That field was added at vLLM 0.25.1 and does not exist at 0.24.0, so every SR
forward raised
AttributeError: 'ForwardContext' object has no attribute 'is_padding'
and took EngineCore down during init:
RuntimeError: Engine core initialization failed.
Measured on itzikv-eval-v24-g5-sr (vllm 0.24.0, torch 2.11.0, transformers
5.17.0 -- the install itself resolved cleanly), 2026-10-04.
The docstring already claimed inertness "where there is nothing to match", and
a missing field is the clearest such case, so this is the code catching up to
its own contract rather than a new behaviour. getattr, not a version compare:
the requirement is "this build has the field", which is what the capability
check states.
Not a regression on any supported line -- 0.25.1 through 0.30 all have the
field, and SR scored 0.8960 against its 0.8948 bar on 0.30 with the attribute
access in place. It only blocks going BELOW the current floor, which is what
the 0.24 probe was submitted to find out.
Signed-off-by: itzikvaisman <yizhak.vaisman@ibm.com>
7accc4c added three setdefaults to run_vllm_logprobs. Two fixed observed failures. The third, enforce_eager=True, was added on the general principle that "an equivalence test must not measure the compiler", and it turned TestZeroAdapterNoHiding::test_no_control_tokens red on every vLLM line from 0.26 to 0.30 -- the last failure in the four-leg matrix. Pinned in one pod (vLLM 0.26.0, sha 5df1b8f, same venv, that setdefault the only variable): enforce_eager=True rc=1 3 failed, 12 deselected 4:26 enforce_eager=False rc=0 3 passed, 12 deselected 6:35 and the test was green at 5c0c72d, where the mini suite ran compiled (_UPSTREAM_EAGER_CONFIGS keyed on "4.0-h-350m", which has never been a key of GRANITE4_MINI, so _eager_kwargs_if_needed always returned {} -- it was dead from the initial commit and is not restored here). granite4_equivalence.py is byte-identical over 5c0c72d..5df1b8f, so the tolerance never moved. Measured the floor rather than assuming it, because "the gate is below the eager noise floor" and "there is a real divergence the compiler was hiding" both predict the above and want opposite patches. Eager, on the zero-LoRA pair the test builds: upstream vs upstream switch vs switch switch vs upstream 4.0-1b 0.0000e+00 0.0000e+00 9.9182e-04 4.0-350m 0.0000e+00 0.0000e+00 1.9522e-03 4.0-micro 0.0000e+00 0.0000e+00 5.0116e-04 Each engine is bit-exact against a second instance of itself over all 3840 elements, so there is no noise floor: the divergence is real and deterministic, and the switch/upstream column reproduces the three failures' worst diffs to five digits. It appears only when the adapter tier is live. save_upstream_model builds num_adapters=0, where the projection is the plain base linear and eager is bit-exact; save_zero_adapter_model builds num_adapters=2, where it runs the fused SWITCH kernel plus a zero-valued shrink/expand -- a different op sequence with a different reduction order. get_tolerances' 4.77e-7 basis for the 1e-5 inert gate was measured at num_adapters=0 and compiled, so it never characterized this configuration. torch.compile lowers both sequences to equivalent kernels and normalizes the ordering, which is exactly what an equivalence test wants canonicalized; it preserves semantics, so a genuine kernel bug still trips 1e-5. Compiled is also how vLLM serves. So the choice belongs to each test. _granite4_fullsize_tests.py keeps enforce_eager=True explicitly -- at full size, with num_adapters=0, eager is bit-exact and an unpinned autotuner reddened it ~1 run in 7 -- and setdefault was never what gave it that. enable_prefix_caching=False and gpu_memory_utilization stay: they fix the vLLM 0.28 mamba assertion and the in-pod OOM respectively. No tolerance was changed. The measurements are recorded at all three sites so the next reader does not re-derive them, or repeat the mistake. Signed-off-by: itzikvaisman <yizhak.vaisman@ibm.com>
The two GPU legs now install the two ENDS of the supported range, 0.26 and 0.30, instead of 0.26 and 0.27. WHY THE ENDS. pyproject declares vllm>=0.26,<0.31, and every break this range actually contains sits at a boundary. APIs deleted at 0.29 only fail at the ceiling: AttentionBackend.get_kv_cache_shape (29 defs at 0.26-0.28, zero at 0.29) and BaseMultiModalProcessor._call_hf_processor (the method that ran ASR, so audio was silently producing no transcript). An API that does not exist yet only fails at the floor: ForwardContext.is_padding, added at 0.27, which the SR router slices. A 0.27 leg is bracketed by two tested lines and would have caught none of them. 0.27/0.28/0.29 were verified green out of band by a five-leg matrix across every suite, so this is a per-run cost decision, not a claim that the middle is untested. Measured at 8a4d69d, one vLLM line per pod, VLLM_USE_FLASHINFER_SAMPLER=0: leg vLLM unit composer hf vllm integration v26 0.26.0 1739p 459p 449p 220p/4s 21p/9s v27 0.27.1 1739p 459p 449p 220p/4s 21p/9s v28 0.28.0 1739p 459p 449p 220p/4s 21p/9s v29 0.29.0 1739p 459p 449p 219p/5s 21p/9s v30 0.30.0 1739p 459p 449p 219p/5s 21p/9s rc=0 on all 25 leg/suite combinations, each with a nonzero passed count -- checked that way because pytest exits 0 on an all-skipped run. The 219-vs-220 gap is one deliberate skip: test_uncached_path_reports_updates_not_applied asserts a contract 0.29 removed. The other four skips are identical on every leg and env-gated. VLLM_USE_FLASHINFER_SAMPLER=0 is now set in tests/conftest.py, because raising the ceiling is what exposes the need for it. On an image shipping CUDA 12.4 the 0.30 engine failed to start 171 times on nvcc fatal : Unknown option '--compress-mode=size' flashinfer JIT-compiles its sampling kernels with --compress-mode, an nvcc option added in CUDA 12.8. It is version-dependent: 0.26 is green on that image without the flag and 0.30 is not, so a 0.26-only CI never needed it. The pod template lives in the runner image rather than this repo, so conftest is the only repo-side lever. setdefault, so an image with CUDA >= 12.8 can export VLLM_USE_FLASHINFER_SAMPLER=1 and exercise that path for real; a newer toolkit in the image is the actual fix. Scoped to the sampler and inert for what these suites assert -- the attention backend resolves to FlashAttention independently, and the equivalence tests compare logprobs, produced before any sampling op runs. uv.lock is regenerated: vllm 0.27.1 -> 0.30.0. It resolves transformers 5.17.0 and torch 2.11.0 for the 0.26 leg / 2.13.0 for the 0.30 leg, matching what the matrix pods resolved through pip. Signed-off-by: itzikvaisman <yizhak.vaisman@ibm.com>
ItzikVa
force-pushed
the
test-vllm-0.30
branch
from
October 5, 2026 07:02
0c5cc47 to
b40a225
Compare
The vLLM 0.26-0.30 widening re-locked uv.lock, and the re-lock was done with
the local uv (0.7.19, Homebrew). That binary writes `revision = 2` in the
lockfile header, so the commit silently DOWNGRADED the field from the
`revision = 3` already on the branch.
.pre-commit-config.yaml pins uv-pre-commit at rev 0.8.4, whose uv writes
`revision = 3`. So CI's `uv-lock` hook re-upgraded the field, reported
"pre-commit hook(s) made changes", and exited 1 on a one-line diff:
-revision = 2
+revision = 3
Re-locked with `uvx uv@0.8.4 lock` so the committed file matches what the
pinned hook produces. The hook is now a no-op.
RESOLUTION IS UNCHANGED, which is why this is header-only: the diff against
the previous commit is exactly one line. 0.8.4 resolves the same 287 packages
(the count CI printed), and the pins this effort cares about all hold -- vllm
0.26.0 for the dev/test/vllm26 groups and 0.30.0 for dev-vllm30/vllm30,
transformers 5.17.0, torch 2.11.0 on the 0.26 leg and 2.13.0 on the 0.30 leg.
Those are the versions the five-leg GPU matrix actually installed, so the
lockfile still describes the configuration that was measured green.
`revision` is a lockfile FORMAT marker, not a dependency, so bumping it cannot
move a version. The rule it encodes is that the file must be written by a uv at
least as new as the hook's pin; writing it with an older one is what
regressed it.
Verified the rest of the hook list separately, since fail_fast: true means the
CI log only proves uv-lock failed: uv-lock is hook 15 of 15, so CI reaching it
already implies 1-14 passed. Confirmed directly anyway -- ruff 0.9.0 format
(223 files) and check clean, check_headers clean, validate_links clean.
Signed-off-by: itzikvaisman <yizhak.vaisman@ibm.com>
Collaborator
Author
|
/gpu-test-multi |
✅ GPU tests passed —
|
❌ GPU tests failed —
|
Two CI reds on this PR, NEITHER of which was a test failure. CPU TESTS: the suite passed. Step-level conclusions on both 3.11 and 3.12 were `success` for "Run CPU tests" and `failure` for the next step, "Upload coverage". codecov-action@v4 exits nonzero on a token/rate-limit problem of its own BEFORE it consults `fail_ci_if_error: false`, so the flag that already declared "upload problems must not gate CI" could not take effect. Moved to v5 (v4 is deprecated) and added `continue-on-error: true`, which enforces that intent at the workflow level where the action cannot override it. This is not masking a failure: `tests/unit/` is untouched by this branch, and the install step resolved exactly as intended (vllm 0.26.0 / torch 2.11.0 / transformers 5.17.0). The same ci.yaml passed on main on 2026-09-30 and fails now, so the trigger is external drift. GPU TESTS: a chicken-and-egg in the dispatcher. `workflow_dispatch` always executes the workflow file from the DEFAULT BRANCH, so `/gpu-test-multi` ran main's matrix -- legs `vllm26 -> dev` and `vllm27 -> dev-vllm27` -- while checking out this PR's commit (the uploaded artifacts are named for sha 8276623). This branch had deleted `dev-vllm27`, so that leg's `uv sync --group dev-vllm27` had no such group and died before pytest started; both logs are under 1.1 KB, the "suite never started" signature. Restored `vllm27` / `dev-vllm27` so main's dispatched matrix stays resolvable against this tree until the 0.26/0.30 rename lands on main. The group is not a fiction -- the `vllm` extra claims 0.26-0.30, so 0.27 is a supported line that simply is not worth its own per-run leg -- and it is removable once main carries the rename. Conflicts now enumerate every cross-line pair: three lines with a bare and a dev- flavour each give 3 x (2 x 2) = 12. `test` stays absent because it reaches vllm26 through `include-group = "dev"` and uv derives that transitively, which is also why main's two-line version never listed it. DELIBERATELY NOT CHANGING `default-groups = ["vllm26"]`, though `uv sync --group dev-vllm30` reports it as conflicting. That shape is not new: at 6013c7f `default-groups` was `["vllm19"]` against a declared `dev-vllm20`/`vllm19` conflict, and that run's vllm20 leg PASSED -- so the cluster harness already excludes default groups. Changing it here would be a guess against evidence. Verified: `uv lock` is a no-op under the hook's pinned uv 0.8.4 (revision 3); each leg resolves to its own line (dev 0.26.0, dev-vllm27 0.27.1, dev-vllm30 0.30.0); ruff 0.9.0 format and check clean; pyproject and both workflows parse. Signed-off-by: itzikvaisman <yizhak.vaisman@ibm.com>
Collaborator
Author
|
/gpu-test-multi |
ItzikVa
added a commit
that referenced
this pull request
Oct 5, 2026
CPU Tests went red on this branch's first commit while the suite itself passed.
Step-level conclusions, identical on 3.11 and 3.12:
success Run uv sync --frozen --group dev-vllm26 --extra hf --extra compose
success Run CPU tests
failure Upload coverage
codecov-action@v4 exits nonzero on a token/rate-limit problem of its own BEFORE
it consults `fail_ci_if_error: false`, so the flag that already declared
"upload problems must not gate CI" never gets a chance to apply. Moved to v5
(v4 is deprecated) and added `continue-on-error: true`, which enforces that
intent at the workflow level where the action cannot override it.
Not masking a failure: `tests/unit/` is untouched here, the install step
resolved as intended, and the test step is green on both interpreters. The
same fix is already on #135, where v5 then uploaded successfully -- evidence
that v4 itself was at fault rather than the token. It is duplicated here
because this PR branches off main and must be able to go green on its own;
whichever lands first makes the other's hunk a no-op.
This also confirms the rename in the previous commits: `uv sync --frozen
--group dev-vllm26` is a passing step in real CI on both Python versions.
Signed-off-by: itzikvaisman <yizhak.vaisman@ibm.com>
Collaborator
|
/gpu-test-dev |
❌ GPU tests failed —
|
❌ GPU tests failed —
|
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.