Skip to content

[HLSL] Add LinAlg OuterProduct and VectorAccumulate coverage#8675

Draft
JoeCitizen wants to merge 23 commits into
microsoft:mainfrom
JoeCitizen:linalg-hlk-outer-vector-coverage
Draft

[HLSL] Add LinAlg OuterProduct and VectorAccumulate coverage#8675
JoeCitizen wants to merge 23 commits into
microsoft:mainfrom
JoeCitizen:linalg-hlk-outer-vector-coverage

Conversation

@JoeCitizen

@JoeCitizen JoeCitizen commented Jul 24, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • add capability-gated non-uniform Thread OuterProduct coverage for F16 8x16 and F32 4x8, while retaining and activating the existing F16 16x16 case
  • accumulate in DXIL OuterProductOptimal layout and use the preview D3D12 matrix-conversion API for tight RowMajor host readback, with exact checked CPU expectations
  • migrate the existing F16 length-four VectorAccumulate case to the shared harness and add F16/F32 length-eight cases with non-zero destinations and checked host multiplication/addition oracles
  • replace the released-SDK OuterProduct skip with SDK-neutral preview ABI mirrors, validated against the preview headers
  • require OuterProduct cases to advertise both ThreadOuterProduct and RWByteAddressBuffer atomic accumulation, assert every field offset in the conversion ABI mirrors, and guard every vector destination with typed trailing sentinels

The shader and host APIs intentionally use their respective enum values for OuterProductOptimal: DXIL layout value 4 and D3D12 conversion layout value 3.

Validation

  • built the Release ExecHLSLTests target
  • compiled both changed implementation units against the preview D3D12 headers, activating exhaustive ABI size, alignment, field-offset and enum assertions
  • ran the capability-policy host test plus all 3 OuterProduct methods on the compatible WARP/Agility runtime: 4 passed; each OuterProduct case positively queried both operation capabilities
  • ran all 3 VectorAccumulate methods on the compatible WARP decoder: 3 passed, including exact verification of typed trailing guards
  • ran the complete 60-case LinAlg selection with the compatible local runtime: 57 passed, the unchanged FP8 and ColumnMajor capability cases skipped, and only the already-documented mandatory native-F32/SInt8 MatVec WARP defect failed
  • inspected emitted DXIL for Thread scope, Accumulator use, non-uniform shapes, layout value 4 and handle-first VectorAccumulate operands
  • clang-format 17.0.1, git diff --check and a focused correctness review pass

Runtime compatibility notes

Public preview WARP currently decodes VectorAccumulate using the obsolete vector-first operand order, while current DXC emits the specified handle-first order. Agility SDK 1.721 also rejects ConvertLinearAlgebraMatrix on a compute command list; the local OuterProduct run used the documented DIRECT-list compatibility workaround. Neither disposable runtime change is included here, and the tests retain the current compiler/runtime contract rather than hiding either defect.

  • after the MatVec, Wave, ThreadGroup, and outer-vector review corrections were propagated through the complete stack, the 66-method suite on the compatible local runtime reported 61 passed, 0 failed, and 5 capability-backed skips

No physical GPU or packaged-HLK qualification is claimed.

Stack

This draft is stacked on PR #8674, which is stacked on PR #8673, PR #8672, PR #8671, PR #8670, PR #8669, PR #8668, PR #8667, PR #8666, PR #8665 and PR #8662. Until those ancestors land, this diff contains their commits as well. The OuterProduct and VectorAccumulate implementation is commit 8ac77765e6; review correction 353ca9a02 adds complete applicability, ABI-offset, and trailing-guard validation.

This remains a draft for named human review. The reviewer should verify the complete preview ABI mirrors, distinct DXIL/host layout values, intersected OuterProduct/atomic capability tuples, conversion/readback path, VectorAccumulate operand contract, trailing guards, and independent exact oracles before requesting maintainer review.

Refs #7841
Refs #8563
Refs #8565

Assisted-by: GitHub Copilot

Jack Elliott and others added 3 commits July 23, 2026 14:52
Use the shared MatrixUse parameter for the OuterProduct result and set it to Accumulator, matching proposal 0035 and the public dx::linalg API. Add a host-side invariant to prevent the legacy A-use declaration from returning.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Create SRV buffers without UAV flags and transition them for both pixel and non-pixel shader access. Use a direct resource-initialization list so the graphics-only pixel state is legal.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add typed F16, F32, I32, and U32 matrix data with safe byte encoding, rectangular row/column-major storage mapping, and explicit exact, permitted-result, or excluded comparison policy. Cover offsets and padded strides with independent host goldens, and migrate the existing CopyConvert tests onto the oracle.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Jack Elliott and others added 3 commits July 25, 2026 12:59
Handle packed row or column byte-count overflow before using the result, and include raw F32 bits in exact mismatch diagnostics. Cover adjacent float bit patterns in the host oracle test.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add ABI-checked wrappers for the six D3D12 Linear Algebra capability
query categories and explicit applicability classification. Gate the
rectangular F32 CopyConvert case using concrete supported wave sizes.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Compile capability-gated CopyConvert coverage at the exact wave size whose MatrixConstruction support was queried. Keep mandatory baseline cases on the existing ranged WaveSize attribute.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
@JoeCitizen
JoeCitizen force-pushed the linalg-hlk-outer-vector-coverage branch from f686e57 to 5043c64 Compare July 25, 2026 01:03
Jack Elliott added 2 commits July 25, 2026 14:10
Validate multiplication support flags per operation, exhaustively check the preview D3D12 ABI mirrors, and preserve query-backed optional skips in HLK mode.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add rectangular Length/GetCoordinate/GetElement coverage and the specified Get/Set out-of-bounds behaviour. Capture thread-local matrix records without UAV races and gate optional F32 cases at the exact queried wave size.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
@JoeCitizen
JoeCitizen force-pushed the linalg-hlk-outer-vector-coverage branch from 5043c64 to 6c86300 Compare July 25, 2026 02:27
Jack Elliott and others added 2 commits July 25, 2026 14:51
Seed OOB Get outputs with non-zero sentinels and require every lane in the selected wave to execute and write the specified zero result.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add bounded raw descriptor-table bindings and independent whole-buffer
oracles for LinAlg descriptor operations. Cover non-zero offsets, padded
strides, row/column-major transfer, descriptor bounds, and capability-gated
atomic accumulation.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
@JoeCitizen
JoeCitizen force-pushed the linalg-hlk-outer-vector-coverage branch from 6c86300 to 7dcabd4 Compare July 25, 2026 02:55
Jack Elliott added 2 commits July 25, 2026 15:21
Reject invalid raw-buffer views, conflicting shader-visible resource heaps, and ambiguous root-parameter bindings before ShaderOp execution.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add race-free Wave and ThreadGroup group-shared transfer coverage for row/column-major layouts, non-zero offsets, padded strides, and exact whole-buffer guards. Add capability-gated Wave atomic accumulation with coordinate-derived values, while keeping cross-component conversion out of scope pending runtime conformance.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
@JoeCitizen
JoeCitizen force-pushed the linalg-hlk-outer-vector-coverage branch from 7dcabd4 to 3ab14c3 Compare July 25, 2026 03:23
Jack Elliott and others added 2 commits July 25, 2026 15:48
Extend each group-shared backing array by four typed sentinel elements so transfer and accumulation tests verify writes do not overrun the matrix extent.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add mixed F16/F32 CopyConvert cases and verify that conversion leaves the source matrix unchanged.

Cover exact integer widening, RTNE plus saturating float narrowing, and capability-gated FP8 encoding and round-trip semantics with independent host oracles.

Assisted-by: GitHub Copilot
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
@JoeCitizen
JoeCitizen force-pushed the linalg-hlk-outer-vector-coverage branch from 3ab14c3 to e4c1fca Compare July 25, 2026 03:51
Jack Elliott added 2 commits July 25, 2026 18:37
Feed host-derived packed FP8 bytes through an SRV for decode so the F16 result cannot false-pass through a folded shader encode/decode chain.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Refactor MatVec execution tests around independent matrix, vector, bias, and output resources with host-derived exact expectations.

Add required interpreted input tuples, non-uniform layout coverage, unsigned output, and independent bias validation behind the runtime ThreadVectorMatrixMultiply capability query. The mandatory native F32-to-SInt8 case remains active and exposes the current preview WARP conversion defect.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
@JoeCitizen
JoeCitizen force-pushed the linalg-hlk-outer-vector-coverage branch from e4c1fca to 3ead6b0 Compare July 25, 2026 06:40
Jack Elliott added 2 commits July 25, 2026 19:50
Separate native F32 inputs from hand-derived SInt8 values so MatVec exercises RTNE saturation, and use high-bit UInt8 lanes to distinguish unsigned packed interpretation.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add capability-gated Wave matrix multiply, multiply-accumulate, and B-use accumulate cases with independent exact host oracles. Make the accumulator-layout query select an observable A-use or B-use execution path.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
@JoeCitizen
JoeCitizen force-pushed the linalg-hlk-outer-vector-coverage branch from 3ead6b0 to cb9e552 Compare July 25, 2026 07:52
Jack Elliott added 2 commits July 25, 2026 20:17
Require multiply-only case data to leave the accumulator vector empty so malformed inputs cannot pass validation and then be silently ignored.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add capability-gated ThreadGroup matrix multiply and multiply-accumulate cases with typed group-shared staging and exact host-derived results. Select and compile at the concrete wave and thread-group sizes advertised for each type and shape.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
@JoeCitizen
JoeCitizen force-pushed the linalg-hlk-outer-vector-coverage branch from cb9e552 to 7c0f49a Compare July 25, 2026 08:18
Jack Elliott added 2 commits July 25, 2026 22:43
Prefer the smallest advertised multi-wave thread-group size when available so ThreadGroup operations cannot pass by behaving only at Wave scope. Add typed trailing guards to the group-shared result store and verify the complete guarded readback.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add non-uniform Thread OuterProduct coverage with exact host readback through the preview matrix-conversion ABI. Add length-eight F16/F32 VectorAccumulate cases with non-zero destinations, capability-gate both operation families, and remove the released-SDK OuterProduct skip through SDK-neutral ABI mirrors.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
@JoeCitizen
JoeCitizen force-pushed the linalg-hlk-outer-vector-coverage branch from 7c0f49a to 8ac7776 Compare July 25, 2026 10:46
Require OuterProduct cases to advertise both the outer-product operation and descriptor accumulation they execute. Assert every field offset in the preview matrix-conversion ABI mirrors, and add typed trailing guards to all vector-accumulation outputs.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: New

Development

Successfully merging this pull request may close these issues.

1 participant