microLLM-rocm is organized as a small runtime with explicit component boundaries. Each
component has a public interface under include/microllm, an implementation under
src, and corresponding tests under tests.
If you are integrating the built SDK from another repository, start with the CMake Config beginner guide (中文).
- Build the repository.
- Run the appropriate test suite.
- Read the repository layout and dependency rules.
- For operator work, follow operator development.
- For performance work, follow profiling.
- For cross-framework work, use alignment experiments.
- For RCCL training, read distributed training.
- For model memory/throughput matrices, read single-GPU benchmarking.
- For matched Python/PyTorch data, read PyTorch performance comparison.
- For context/batch/KV-cache inference evidence, read the simple inference-matrix guide.
- For device-resident greedy token collection, read the GPU token-history guide.
- For unequal cached positions in one batch, read the divergent KV-row guide.
- For inserting a new prompt into one empty shared-cache slot, read the slot row-prefill guide.
- For token-level refill across fixed shared-cache slots, read the continuous slot-scheduler guide.
- For removing inactive dummy model rows, read the active-row compaction guide.
- For batching real rows with different cache positions, read the positions-aware decode guide.
- For collecting an uncontaminated scheduler trace, read the continuous-only profile guide.
- For merging tiny token/position/row uploads, read the packed decode-metadata guide.
- For batching equal-length prompt admission, read the batched slot-prefill guide.
- For official Qwen/DeepSeek short/long context, slot, KV-cache and memory evidence, read the continuous serving matrix guide.
- For a fixed-request 1/2/4/8-slot efficiency and Cache sweep, read the continuous slot-sweep guide.
- For locating a low-margin cross-slot token divergence, read the continuous divergence guide.
- For swapping and duplicating B2 prefill rows, read the prefill row-audit guide.
- For complete-logit and per-block B1/B2 error growth, read the prefill layer-drift guide.
- For the block-zero Attention/FFN substage split, read the block-zero drift guide.
- For cast/gate/up/SwiGLU/down FFN drift, read the BF16 FFN drift guide.
- For M32/M64 hipBLASLt candidate intersection, read the BF16 algorithm-inventory guide.
- For the version-local same-algorithm counterfactual, read the same BF16 algorithm guide.
- For request TTFT/completion and slot tradeoffs, read the request latency guide.
- For sharing weights while splitting KV capacity by request length, read the length-bucketed KV-cache guide.
- For delayed arrivals, skewed bucket traffic and tail latency, read the continuous arrival guide.
- For FP32/BF16 cache policy and its numerical gates, read the KV-cache dtype guide.
- For delayed multi-request serving semantics, read the serving scheduler guide.
- For the measured optimization loop, read the 0→1 optimization log.
- For separating FP8 weight and activation error, read the FP8 error-attribution guide.
- For halving AdamW moment state without changing FP32 master weights, read the BF16 AdamW moment guide.
- For the one-byte symmetric weight format, scale Tensor, safetensors layout and current no-INT8-GEMM boundary, read the INT8 weight-format guide.
- For models whose Attention width differs from hidden size and which normalize Q/K before RoPE, read the explicit head-dimension and QK-Norm guide.
- CPU reference behavior is the numerical oracle inside the engine.
- PyTorch is the independent external oracle for the supported FP32 domain.
- Readable HIP remains available when an optimized implementation is introduced.
- Optional bindings and vendor libraries do not become core dependencies.
- A benchmark result is accepted only with its environment, warm-up, repetitions, correctness regression, and end-to-end measurement.
- “Implemented”, “smoke-tested”, “reference-trained”, and “released” are different evidence states.
The versioned C ABI is the most stable integration boundary. The C++ API is pre-1.0 and may evolve, but public changes require shape/error tests and documentation. Python and PyTorch integrations are optional adapters over the same engine.