Skip to content

test: validate AMD power exporters on MI300X, MI325X, and MI355X - #3434

Draft
cquil11 wants to merge 4 commits into
mainfrom
validate/amd-power-profiles-20260925
Draft

cquil11 wants to merge 4 commits into
mainfrom
validate/amd-power-profiles-20260925

Conversation

@cquil11

@cquil11 cquil11 commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

Tests srt-slurm #31 at 4f95eee1b9f50fc0dadbc7163f93800680bca0a7 as a patch on the current InferenceX runtime.

One Qwen3.5 FP8 fixed-sequence point per cluster: TP8, concurrency 4, 8k1k. MI355X uses a test-only TP8 override; these results are not for performance publication.

Runs the official AMD exporter on port 19500 to avoid existing cluster services, requires native power telemetry, and writes the single-node measurement window using the existing InferenceX helper. Python setup uses runner scratch because MI325X's home-directory mount is unavailable.

Validation

Cluster Run Result
MI300X Single-node Passed; 40/40 requests, 8-GPU power validation
MI325X Single-node Passed; 40/40 requests, 8-GPU power validation
MI355X Single-node Passed; 40/40 requests, 8-GPU power validation
MI355X Disaggregated 1P1D Failed telemetry readiness: exporter omitted GPU power readings

The disaggregated smoke uses the existing Qwen3.5 FP8 SGLang/MoRI recipe, with test-only TP8 prefill and one concurrency, to validate all 16 GPUs across two nodes. Required native telemetry is enabled.

Do not merge this smoke-test branch.

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant