Skip to content

Latest commit

 

History

History
99 lines (91 loc) · 5.59 KB

File metadata and controls

99 lines (91 loc) · 5.59 KB

Developer guide

microLLM-rocm is organized as a small runtime with explicit component boundaries. Each component has a public interface under include/microllm, an implementation under src, and corresponding tests under tests.

If you are integrating the built SDK from another repository, start with the CMake Config beginner guide (中文).

Start here

  1. Build the repository.
  2. Run the appropriate test suite.
  3. Read the repository layout and dependency rules.
  4. For operator work, follow operator development.
  5. For performance work, follow profiling.
  6. For cross-framework work, use alignment experiments.
  7. For RCCL training, read distributed training.
  8. For model memory/throughput matrices, read single-GPU benchmarking.
  9. For matched Python/PyTorch data, read PyTorch performance comparison.
  10. For context/batch/KV-cache inference evidence, read the simple inference-matrix guide.
  11. For device-resident greedy token collection, read the GPU token-history guide.
  12. For unequal cached positions in one batch, read the divergent KV-row guide.
  13. For inserting a new prompt into one empty shared-cache slot, read the slot row-prefill guide.
  14. For token-level refill across fixed shared-cache slots, read the continuous slot-scheduler guide.
  15. For removing inactive dummy model rows, read the active-row compaction guide.
  16. For batching real rows with different cache positions, read the positions-aware decode guide.
  17. For collecting an uncontaminated scheduler trace, read the continuous-only profile guide.
  18. For merging tiny token/position/row uploads, read the packed decode-metadata guide.
  19. For batching equal-length prompt admission, read the batched slot-prefill guide.
  20. For official Qwen/DeepSeek short/long context, slot, KV-cache and memory evidence, read the continuous serving matrix guide.
  21. For a fixed-request 1/2/4/8-slot efficiency and Cache sweep, read the continuous slot-sweep guide.
  22. For locating a low-margin cross-slot token divergence, read the continuous divergence guide.
  23. For swapping and duplicating B2 prefill rows, read the prefill row-audit guide.
  24. For complete-logit and per-block B1/B2 error growth, read the prefill layer-drift guide.
  25. For the block-zero Attention/FFN substage split, read the block-zero drift guide.
  26. For cast/gate/up/SwiGLU/down FFN drift, read the BF16 FFN drift guide.
  27. For M32/M64 hipBLASLt candidate intersection, read the BF16 algorithm-inventory guide.
  28. For the version-local same-algorithm counterfactual, read the same BF16 algorithm guide.
  29. For request TTFT/completion and slot tradeoffs, read the request latency guide.
  30. For sharing weights while splitting KV capacity by request length, read the length-bucketed KV-cache guide.
  31. For delayed arrivals, skewed bucket traffic and tail latency, read the continuous arrival guide.
  32. For FP32/BF16 cache policy and its numerical gates, read the KV-cache dtype guide.
  33. For delayed multi-request serving semantics, read the serving scheduler guide.
  34. For the measured optimization loop, read the 0→1 optimization log.
  35. For separating FP8 weight and activation error, read the FP8 error-attribution guide.
  36. For halving AdamW moment state without changing FP32 master weights, read the BF16 AdamW moment guide.
  37. For the one-byte symmetric weight format, scale Tensor, safetensors layout and current no-INT8-GEMM boundary, read the INT8 weight-format guide.
  38. For models whose Attention width differs from hidden size and which normalize Q/K before RoPE, read the explicit head-dimension and QK-Norm guide.

Engineering rules

  • CPU reference behavior is the numerical oracle inside the engine.
  • PyTorch is the independent external oracle for the supported FP32 domain.
  • Readable HIP remains available when an optimized implementation is introduced.
  • Optional bindings and vendor libraries do not become core dependencies.
  • A benchmark result is accepted only with its environment, warm-up, repetitions, correctness regression, and end-to-end measurement.
  • “Implemented”, “smoke-tested”, “reference-trained”, and “released” are different evidence states.

Public stability

The versioned C ABI is the most stable integration boundary. The C++ API is pre-1.0 and may evolve, but public changes require shape/error tests and documentation. Python and PyTorch integrations are optional adapters over the same engine.