Skip to content

[HLSL] Add LinAlg CopyConvert and Convert coverage#8671

Draft
JoeCitizen wants to merge 9 commits into
microsoft:mainfrom
JoeCitizen:linalg-hlk-convert-coverage
Draft

[HLSL] Add LinAlg CopyConvert and Convert coverage#8671
JoeCitizen wants to merge 9 commits into
microsoft:mainfrom
JoeCitizen:linalg-hlk-convert-coverage

Conversation

@JoeCitizen

Copy link
Copy Markdown
Collaborator

Summary

  • require CopyConvert source and destination component support at the same wave size, add rectangular F16-to-F32 and transposed F32-to-F16 cases, and verify the source matrix remains unchanged
  • add exact I16-to-I32 widening and F32-to-I16 RTNE-plus-saturation Convert coverage
  • add independently capability-gated E4M3FN and E5M2 packed encoding plus F16 round-trip coverage using a host-derived FP8 oracle

D3D12 exposes no dedicated Convert capability query, so the FP8 cases use shape-independent MatrixConstruction support as the closest component-type proxy. Each advertised format executes; the method is not applicable only when neither format is advertised.

Validation

  • built the Release ExecHLSLTests target
  • ran 30 focused stack cases on WARP 1.65535.20-preview, Agility SDK 1.721.2-preview, and ExperimentalShaders=*: 29 passed and the FP8 method skipped because WARP advertises neither FP8 format
  • compiled both E4M3FN and E5M2 shader variants independently
  • compiled HlslExecTestUtils.cpp and LinAlgTests.cpp against the preview D3D12 headers
  • clang-format 17.0.1 and git diff --check pass

No physical GPU or HLK lab execution is claimed.

Stack

This draft is stacked on PR #8670, which is stacked on PR #8669, PR #8668, PR #8667, PR #8666, PR #8665, and PR #8662. Until those ancestors land, this diff contains their commits as well. The CopyConvert/Convert change itself is commit 49b913926.

This remains a draft for named human review. The reviewer should verify the mixed-type capability intersection, source-preservation oracle, RTNE/saturation expectations, FP8 host encoding, and composite applicability policy before requesting maintainer review.

Refs #7841
Refs #8650
Refs #8653
Refs #8546
Refs #8564

Assisted-by: GitHub Copilot

Jack Elliott and others added 9 commits July 23, 2026 14:52
Use the shared MatrixUse parameter for the OuterProduct result and set it to Accumulator, matching proposal 0035 and the public dx::linalg API. Add a host-side invariant to prevent the legacy A-use declaration from returning.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Create SRV buffers without UAV flags and transition them for both pixel and non-pixel shader access. Use a direct resource-initialization list so the graphics-only pixel state is legal.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add typed F16, F32, I32, and U32 matrix data with safe byte encoding, rectangular row/column-major storage mapping, and explicit exact, permitted-result, or excluded comparison policy. Cover offsets and padded strides with independent host goldens, and migrate the existing CopyConvert tests onto the oracle.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add ABI-checked wrappers for the six D3D12 Linear Algebra capability
query categories and explicit applicability classification. Gate the
rectangular F32 CopyConvert case using concrete supported wave sizes.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Compile capability-gated CopyConvert coverage at the exact wave size whose MatrixConstruction support was queried. Keep mandatory baseline cases on the existing ranged WaveSize attribute.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add rectangular Length/GetCoordinate/GetElement coverage and the specified Get/Set out-of-bounds behaviour. Capture thread-local matrix records without UAV races and gate optional F32 cases at the exact queried wave size.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add bounded raw descriptor-table bindings and independent whole-buffer
oracles for LinAlg descriptor operations. Cover non-zero offsets, padded
strides, row/column-major transfer, descriptor bounds, and capability-gated
atomic accumulation.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add race-free Wave and ThreadGroup group-shared transfer coverage for row/column-major layouts, non-zero offsets, padded strides, and exact whole-buffer guards. Add capability-gated Wave atomic accumulation with coordinate-derived values, while keeping cross-component conversion out of scope pending runtime conformance.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add mixed F16/F32 CopyConvert cases and verify that conversion leaves the source matrix unchanged.

Cover exact integer widening, RTNE plus saturating float narrowing, and capability-gated FP8 encoding and round-trip semantics with independent host oracles.

Assisted-by: GitHub Copilot
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: New

Development

Successfully merging this pull request may close these issues.

1 participant