Skip to content

Tensor transpose CPU+GPU - #145

Merged
abagusetty merged 19 commits into
NWChemEx:mainfrom
abagusetty:tt_inhouse
Sep 30, 2026
Merged

abagusetty merged 19 commits into
NWChemEx:mainfrom
abagusetty:tt_inhouse

Conversation

@abagusetty

Copy link
Copy Markdown
Contributor

Replaces librett & hptt with GPU stream-aware pipeline with inhouse tensor transpose kernels

abagusetty and others added 9 commits July 9, 2024 23:22
Replaces the librett-based tensor transpose in assign_gpu with a
hand-written, rank-agnostic reorder kernel supporting CUDA/HIP/DPCPP.

- New kernels/gpu_reorder.cpp: linear-index N-D axis-permuting transpose,
  one thread per output element. Supports ranks 1..8 (tamm::maxrank),
  unlike the previous 4D-only ArrayFire-derived attempt.
- Column-major (first-axis-fastest) strides, verified byte-identical to
  librett's TensorConv convention (TensorTester reference) for all ranks,
  arbitrary permutations, and double/float/complex types.
- Runs entirely on the caller's cuda stream / hip stream / in-order sycl
  queue (handle.first). No plan construction, no host round-trips, no
  internal synchronization -- addresses librett's synchronization issues.
- Complex values transported as trivially-copyable POD bytes (device-safe;
  matches librett's opaque-byte handling).
- Removes broken transpose_inplace.cpp (wrong in-place 2D kernel; would
  not compile) and the transpose_inplace declaration.
- multiply.hpp: drops librett include, calls gpu::transpose_reorder.
Drop the librett dependency: new transpose_reorder kernel (CUDA/HIP/SYCL)
driven by allocation-free ReorderSpec metadata. GPU output transpose keeps
librett overwrite semantics (GEMM beta accumulates the running total, so
accumulating again in the transpose double-counts). Full ExaChem CI 19/19.
Host row-major transpose_reorder_cpu (general alpha/beta, odometer loop,
no per-element division) reusing the gpu_reorder math. Rewires assign,
BlockAssignPlan and drops the HPTT dependency. Unit tests: 117522
assertions green; h2o/butanol2 CPU exact.
@abagusetty
abagusetty marked this pull request as ready for review September 22, 2026 14:26
@abagusetty
abagusetty merged commit 8019a83 into NWChemEx:main Sep 30, 2026
3 checks passed
@abagusetty
abagusetty deleted the tt_inhouse branch September 30, 2026 18:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants