Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
40 changes: 35 additions & 5 deletions apps/hf_infer.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -1650,11 +1650,41 @@ int main(int argc, char** argv) {
model.to(device);
microllm::model::LoadWeightsOptions load_options;
load_options.mapping = microllm::model::qwen_style_weight_mapping(external.model);
load_options.aliases =
microllm::model::qwen3_tied_weight_aliases(external.model);
// qwen3_tied_weight_aliases declares "lm_head.weight" as a mandatory
// redundant tensor whenever qk_norm+tie_embeddings hold, matching
// checkpoints observed for dense Qwen3 -- but not every real tied
// checkpoint actually serializes that redundant tensor (a Qwen3-MoE
// fixture prepared for M7 does not). Only apply the alias when the
// weight file genuinely contains it, so a legitimately absent
// redundant tensor doesn't get strict-mode-reported as "missing".
const auto declared_aliases = microllm::model::qwen3_tied_weight_aliases(external.model);
if (!declared_aliases.empty()) {
const auto weights_path = std::filesystem::path(command.weights);
const auto has_tensor = [&](const std::string& name) {
if (weights_path.extension() == ".json") {
const auto index = microllm::io::inspect_safetensors_index(weights_path);
return index.contains(name);
}
const auto tensors = microllm::io::inspect_safetensors(weights_path);
return std::any_of(tensors.begin(), tensors.end(),
[&](const auto& info) { return info.name == name; });
};
for (const auto& [source_alias, target] : declared_aliases) {
if (has_tensor(source_alias)) load_options.aliases[source_alias] = target;
}
}
const auto load_start = std::chrono::steady_clock::now();
const auto report = model.load_safetensors(command.weights, load_options);
const auto load_finish = std::chrono::steady_clock::now();
// external.model.weight_bytes() throws for MoE configs (the exact
// per-expert layout is a weight-loading concern, not a config-formula
// one -- see ModelConfig::parameter_count()). model.parameter_count()
// sums the model's own live named_parameters(), which already
// reflects the real loaded weights for both dense and MoE models, so
// it is used as the FP32-resident-bytes basis for every model rather
// than only working around MoE here.
const auto stored_weight_bytes =
static_cast<std::uint64_t>(model.parameter_count()) * sizeof(float);
microllm::model::Bf16FfnPreparationReport bf16_report;
microllm::model::Bf16FfnDecodeUpPreparationReport
bf16_decode_up_report;
Expand Down Expand Up @@ -2036,7 +2066,7 @@ int main(int argc, char** argv) {
device, "host CPU", "host"}
: microllm::runtime::device_info(device);
const auto resident_weight_bytes =
external.model.weight_bytes(sizeof(float)) -
stored_weight_bytes -
bf16_report.fp32_bytes_released +
bf16_report.bf16_bytes_retained -
bf16_attention_report.fp32_bytes_released +
Expand Down Expand Up @@ -2948,7 +2978,7 @@ int main(int argc, char** argv) {
<< ",\"int8_device_amax_tensors\":"
<< int8_report.device_amax_tensors
<< ",\"resident_weight_bytes\":"
<< external.model.weight_bytes(sizeof(float)) -
<< stored_weight_bytes -
bf16_report.fp32_bytes_released +
bf16_report.bf16_bytes_retained -
bf16_attention_report.fp32_bytes_released +
Expand All @@ -2966,7 +2996,7 @@ int main(int argc, char** argv) {
<< "\""
<< ",\"parameter_count\":" << model.parameter_count()
<< ",\"fp32_weight_bytes\":"
<< external.model.weight_bytes(sizeof(float))
<< stored_weight_bytes
<< ",\"loaded_tensors\":" << report.loaded.size()
<< ",\"token_count\":" << ids.size()
<< ",\"batch\":" << command.batch
Expand Down
21 changes: 21 additions & 0 deletions data/model_fixtures.toml
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,27 @@ inference_token_ids = [1]
expected_generated_tokens = [14582, 25, 16246, 264]
training_tokens = "1,2,3,4"

[[model]]
id = "qwen3-moe-tiny-random"
repo_id = "amd-quark/tiny-random-qwen3_moe"
source = "https://huggingface.co/amd-quark/tiny-random-qwen3_moe"
revision = "ca2aaa9a82e9f86c09c2785b405932fe6c806c90"
# Synthetic structural-test checkpoint (real Qwen3MoeForCausalLM architecture,
# randomly initialized weights): no LICENSE file and no license tag on the
# Hugging Face repo, and the same is true of every comparable tiny-random
# Qwen3-MoE test repo checked while preparing this fixture (M7). Weights are
# never committed to this repository regardless -- only this pinned-revision
# registry entry is -- and there is no pretrained/copyrighted training content
# here to license in the first place, only random initialization values.
license = "unspecified (synthetic random-weight test checkpoint; no LICENSE file published on Hugging Face)"
license_url = "https://huggingface.co/amd-quark/tiny-random-qwen3_moe"
parameter_count = 20334464
tensor_count = 44
tokenizer_family = "qwen2"
inference_token_ids = [9707, 1879]
expected_generated_tokens = [1879, 1879, 1879, 1879]
training_tokens = "1,2,3,4"

[[model]]
id = "deepseek-r1-distill-qwen-1.5b"
repo_id = "deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B"
Expand Down
64 changes: 64 additions & 0 deletions docs/OPERATOR_CONTRACTS.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -59,6 +59,70 @@ FP32 `[1,N]`。真实K/N完整输出、零payload传输和少一次逻辑分配
反量化和`FusedDecode`按列读取。CPU/HIP完整值与零传输门保留,但官方Qwen精度失败,所以该
原语不能成为默认模型策略。

## MoE 路由(CPU/HIP readable 前向 + CPU backward + 真实 model forward 已实现;ROCm 硬件门尚未跑)

CPU reference 见 [M1 CPU reference](development/2026-09-04-m1-qwen3-moe-cpu-reference.md);
readable HIP kernel 见 [M2 HIP readable kernels](development/2026-09-04-m2-qwen3-moe-hip-kernels.md);
CPU backward + autograd 见 [M3 autograd backward nodes](development/2026-09-04-m3-qwen3-moe-autograd-backward.md);
config parsing 见 [M4 config parsing](development/2026-09-04-m4-qwen3-moe-config-parsing.md);
weight loading 见 [M5 weight loading](development/2026-09-04-m5-qwen3-moe-weight-loading.md);
model-level full-graph gate 见 [M6 model forward](development/2026-09-04-m6-qwen3-moe-model-forward.md);
真实 checkpoint 门见 [M7 real checkpoint](development/2026-09-05-m7-qwen3-moe-real-checkpoint.md);
契约设计过程见 [M0 operator contracts](development/2026-09-04-m0-qwen3-moe-op-contracts.md)。
M2 的 `.hip` kernel 代码在没有 ROCm 工具链的机器上写成,**从未编译或在真实 AMD GPU 上跑过**——
C++ dispatch 层(`ops.cpp` 里的 `device().is_hip()` 分支)已经在 `MICROLLM_ENABLE_HIP=OFF`
下编译验证过,但 kernel 本身的三方 CPU/HIP/PyTorch oracle 门(`tests/ops/hip_ops_test.cpp`)
要等有 AMD GPU 时才能真正跑,见 M2 记录里的说明。M3/M6 新增的 backward 目前也都是 CPU-only
(HIP 分支显式 `throw`),已经用 conda 环境的 PyTorch(transformers 5.8.0 也在同一环境)跑过真实的
`TorchOps.OperatorParity` 三方门,M6/M7 的门直接对照未修改的 `Qwen3MoeSparseMoeBlock.forward()`。

**M6→M7 修正(推翻了 M6 自己的修正)**:M6 曾经把内部权重表示改成"融合 gate_up_proj",
依据是读了 transformers 5.8.0 的**当前内存模块源码**(`Qwen3MoeExperts` 用一个打包
`nn.Parameter`)。M7 准备 fixture 时下载了一个真实 checkpoint
(`amd-quark/tiny-random-qwen3_moe`)并另外查了官方 `Qwen/Qwen3-30B-A3B` 的
safetensors index,发现**磁盘上的真实文件格式**其实是 per-expert 独立 tensor
(`mlp.experts.E.{gate_proj,up_proj,down_proj}.weight`)——跟 M5 最初的设计一致,
不是 M6 猜的融合格式。用 transformers 实际 load 这个 checkpoint 并检查发现:它在
load 时会用 `torch.cat([gate_proj,up_proj],dim=0)` 把磁盘上的 per-expert tensor
拼成内存里的打包 Parameter——这是 transformers 自己的内部转换层,本仓库不需要
复刻它,因为一旦格式理解对了,本仓库从来没有"打包表示"这个问题要解决。
内部表示已改回 per-expert 独立 `Linear`(`WeightMapping` 恢复成逐 tensor 的
`Transpose2D`,不需要新的 loader 能力),`moe_split_gate_up`/`_backward` 已删除
(不是弃用——现有证据下它解决的问题在任何真实 checkpoint 里都不存在),换成
`moe_stack_experts`/`_backward_one`(见下表):把 N 个分开的同形状 tensor 堆成
`moe_expert_ffn` 要的 `[num_experts,rows,cols]`,不需要转置(本仓库自己的 `Linear`
权重布局 `[input,output]` 已经和 `moe_expert_ffn` 的 per-expert 约定一致;只有
HF 的 nn.Linear 风格 `[output,input]` per-expert tensor 才需要在加载时 `Transpose2D`,
这一点从 M5 起就没变过)。

第一版故意选择"计算全部专家、掩码非选中"的朴素路径(O(num_experts) 而不是
O(k) 的 gather/dispatch),把稀疏 dispatch 留给之后的性能里程碑,此处只对
数值正确性负责。

| microLLM(草案) | 输入和输出 shape | PyTorch oracle | FP32 阈值 | 必测非法输入 |
|---|---|---|---|---|
| `moe_router_top_k` | logits `[tokens,num_experts]` → indices `[tokens,k]`(Int32)、weights `[tokens,k]`(FP32) | `p=torch.softmax(logits.float(),-1); w,idx=torch.topk(p,k,-1); if norm_topk_prob: w=w/w.sum(-1,keepdim=True)` | `2e-6,2e-5` | 非二维 logits、`k<=0`、`k>num_experts`、非 FP32、HIP 非连续 |
| `moe_expert_ffn` | input `[tokens,dim]`,expert_indices `[tokens,k]`(Int32),gate/up weight `[num_experts,dim,ffn_dim]`,down weight `[num_experts,ffn_dim,dim]` → `[tokens,num_experts,dim]` | 对每个专家 `e`:`F.silu(x@gate[e]) * (x@up[e]) @ down[e]`,再用 `expert_indices` 生成的 one-hot mask 把未选中专家的输出置零;对全部专家逐个计算(不做 gather) | `2e-6,2e-5` | input/weight `dim` 或 `ffn_dim` 不一致、`expert_indices` 越界、`k>num_experts`、非 FP32 |
| `moe_combine` | expert_output `[tokens,num_experts,dim]`,expert_indices `[tokens,k]`,expert_weights `[tokens,k]` → `[tokens,dim]` | `sum_j expert_weights[:,j:j+1] * expert_output.gather(1, expert_indices[:,j].view(-1,1,1).expand(-1,1,dim)).squeeze(1)` | 默认 `1e-6,1e-5` | indices/weights 的 `k` 不一致、indices 越界、shape/device 不同 |
| `moe_stack_experts`(M7,checkpoint 格式适配器,不是路由原语) | num_experts 个 `[rows,cols]` tensor → `[num_experts,rows,cols]` | `torch.stack(experts, dim=0)` | 精确(纯拷贝,无浮点运算) | 空列表、shape/dtype 不一致、非 FP32 |

## MoE 路由反向(M3,CPU only;HIP backward 显式 throw)

| microLLM backward | 返回 shape | PyTorch reference | 说明 |
|---|---|---|---|
| `moe_router_top_k_backward` | logits 梯度 `[tokens,num_experts]` | `p=softmax(logits); w=p.gather(-1,idx); (renorm); (w*seed).sum().backward()` | indices 不可微,只作为常量;由于 softmax 分母耦合全部专家,**每个 logit 都会收到非零梯度**,不是稀疏的——这是正确的稠密 softmax Jacobian,不是 bug。内部复用 `softmax`/`softmax_backward`,不是独立公式 |
| `moe_expert_ffn_backward` | `{input, gate_weight, up_weight, down_weight}` 四个梯度 | 对每个专家的 dense SwiGLU FFN 反向,`torch.autograd` 直接对 forward 求导 | 前向对未选中 `(token,expert)` 输出恒为零,其局部 Jacobian 对该专家权重也恒为零,因此专家权重梯度只由选中过该专家的 token 贡献——效果等价于 `embedding_backward` 的 scatter-add,但这里是 mask-multiply 的自然结果,不是手写 scatter |
| `moe_combine_backward` | `{expert_output, expert_weights}` 两个梯度 | 对 `moe_combine` 公式直接求导 | `expert_output` 梯度是真正的 scatter-add:只有被 `moe_combine` 前向实际读取过的 `(token,expert)` 位置才非零 |
| `moe_stack_experts_backward_one` | 单个专家的梯度 `[rows,cols]` | PyTorch 对 `torch.stack` 执行 autograd(等价于取出对应切片) | `autograd::moe_stack_experts` 让每个专家的 backward 只从堆叠梯度里读回自己那一片,天然不重叠,不需要联合 backward |

`autograd::moe_router_top_k`(返回 `MoeRouterResult{Tensor indices; Value weights;}`)、
`autograd::moe_expert_ffn`、`autograd::moe_combine` 是对应的 Value 级别封装,`indices`
参数(或返回值)始终是普通 `Tensor`,从不是 `Value`,与 `embedding()` 的 `indices` 约定一致。
测试标准:反向 reference 必须是 PyTorch 对 forward 执行 `autograd`(`test_operator_parity.py`
里的 `graph_moe_*` 系列 case),不能手写一份"答案公式"当作 oracle;未被选中的专家权重/
expert_output 行必须严格为零(`EXPECT_EQ`,不是 `EXPECT_NEAR`),测试严格度与
`cross_entropy` 的 ignored-row 梯度测试、`causal_softmax` 的 future-mask 梯度测试一致。

## BF16 weight gradient

`bf16_weight_gradient(input_fp32, output_gradient_fp32)` 接受两个连续二维 FP32 Tensor:
Expand Down
57 changes: 57 additions & 0 deletions docs/development/2026-09-04-m0-qwen3-moe-op-contracts.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,57 @@
# 2026-09-04 — M0 Qwen3 MoE operator contracts

## Contract

This is the paper-only milestone from the Qwen3 MoE plan: define the public shape,
dtype, and PyTorch-oracle contract for three new operators before touching `ops.h`,
`coverage_manifest.json`, or any `.cpp`/`.hip` file. No code lands in this change.
The full contract text lives in
[`docs/OPERATOR_CONTRACTS.zh-CN.md`](../OPERATOR_CONTRACTS.zh-CN.md), under the "MoE
路由" heading; this record captures the decision context and the next-milestone
boundary. (As of M1, that heading and its implementation status have moved on from
this record's "M0, not implemented" framing — see
[the M1 CPU reference record](2026-09-04-m1-qwen3-moe-cpu-reference.md).)

## Observed state before this change

`grep -r moe src/ include/` finds nothing. The repository supports Qwen2 and Qwen3
as dense-only architectures; there is no router, no top-k operator, and no per-expert
weight loading path anywhere in the tree.

## Three operators proposed

- `moe_router_top_k`: router logits `[tokens, num_experts]` to `(indices[tokens,k]
Int32, weights[tokens,k] FP32)`. Oracle is softmax over the full expert row, then
`torch.topk`, with an optional renormalization over just the k selected weights
(`norm_topk_prob`), matching the Qwen3 MoE routing block.
- `moe_expert_ffn`: deliberately naive first version. Computes the SwiGLU FFN for
every expert against every token (`O(num_experts)`, not `O(k)`), then masks out
experts that were not selected for a given token. No gather or dispatch yet — that
is explicitly deferred to a later performance milestone.
- `moe_combine`: weighted sum of the k selected expert outputs back down to
`[tokens, dim]`, using the router weights from `moe_router_top_k`.

## Decisions

- Contracts are written before any implementation, per the correctness-first
operator workflow (`docs/dev/operator-development.md` step 1).
- The naive compute-all-experts-then-mask shape is chosen over sparse dispatch for
the first version so correctness can be validated independently of a gather
kernel. Grouped/sparse expert dispatch is explicitly deferred to the performance
goal, not silently assumed here.
- Backward is out of scope for M0/M1/M2; the contract only records that top-k
selection itself is not differentiable and that gradient must flow like
`embedding_backward`'s scatter-add — only visited `(token, expert)` rows receive
gradient, everything else must be exactly zero. This becomes binding once M3
(autograd backward nodes) starts.
- Every new op name will need a matching entry in `tests/coverage_manifest.json` the
moment `ops.h` gains the declaration; that sync is scheduled for M1, not here.

## Current boundary

This change is contracts only. It does not add CPU reference code, HIP kernels,
config parsing, weight loading, or coverage manifest entries. The next milestone
(M1) implements the naive CPU float32 reference for these three operators, gated by
a PyTorch oracle (`torch.topk` + masked-softmax-renorm + manual masked SwiGLU loop)
with exact-value comparison, and registers all three names in
`tests/coverage_manifest.json` so `scripts/audit_test_coverage.py` does not fail CI.
64 changes: 64 additions & 0 deletions docs/development/2026-09-04-m1-qwen3-moe-cpu-reference.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
# 2026-09-04 — M1 Qwen3 MoE CPU reference operators

## Contract

Implement the naive CPU float32 reference for the three MoE operators contracted in
M0 (`docs/development/2026-09-04-m0-qwen3-moe-op-contracts.md`), gated by a PyTorch
oracle and hand-computed values, before any HIP kernel exists.

## Implemented operators

- `moe_router_top_k`: softmax over the full expert row, then top-k selection, with
an optional renormalization over just the k selected weights;
- `moe_expert_ffn`: evaluates the SwiGLU FFN for every expert against every token
(`O(num_experts)` per token, not `O(k)`) and masks out experts not selected for a
given token — no gather/dispatch;
- `moe_combine`: weighted sum of the k selected expert outputs back to
`[tokens, dim]`.

All three reject a HIP-device Tensor explicitly (`"... has no HIP kernel yet; CPU
reference only until M2"`) rather than silently falling back to host compute.

## Verification

```text
CPU Debug (MICROLLM_ENABLE_HIP=OFF): 288/288 microllm_tests passed, including the
three new CpuOpsTest.Moe* cases with hand-computed expected values (router top-2
softmax reduces to sigmoid(+-1) for the chosen two-value case; expert_ffn/combine
use small identity-matrix weights so the SwiGLU and weighted-sum arithmetic is
exact by hand).
scripts/audit_test_coverage.py: pass tensor_ops=202 (was 199) graph_api=45 test_files=159
```

The three-way CPU/PyTorch oracle (`tests/torch/operator_oracle.cpp` emit cases plus
matching `python/tests/test_operator_parity.py` `record()` calls and
`invalid_moe_*` rejection names) was run against a CUDA-built PyTorch from a local
conda env (`conda activate torch`; `torch==2.11.0+cu130`) by pointing
`CMAKE_PREFIX_PATH` at its `torch/share/cmake` directory. `find_package(Torch)`
succeeds against a CUDA build without a working CUDA runtime being present, and
`microllm_operator_oracle` itself has no torch dependency at all (it only links
`microllm::model`/`microllm::training`), so this validates real values, not just
configuration:

```text
$ MICROLLM_OPERATOR_ORACLE=.../microllm_operator_oracle python python/tests/test_operator_parity.py -v
test_every_declared_invalid_shape_or_dtype_is_rejected ... ok
test_every_numeric_case_has_a_pytorch_reference_and_matching_shape ... ok
test_forward_and_backward_values_match_pytorch ... ok
test_model_graph_snapshot_has_real_topology ... ok
Ran 4 tests in 1.349s — OK
```

`test_forward_and_backward_values_match_pytorch` is the one that matters for this
milestone: it confirms `moe_router_top_k`'s softmax+top-k+renorm, `moe_expert_ffn`'s
masked all-experts SwiGLU, and `moe_combine`'s weighted sum are bit-for-bit
consistent (within the default `3e-5` tolerance) with independently written PyTorch
formulas, not just internally self-consistent. This closes the verification gap
this record originally carried — the earlier draft of this paragraph said the
parity test could not be run in this environment; it has now been run.

## Current boundary

No `.hip` kernel, no autograd node, no config parsing, and no weight loading exist
for these three ops yet. The next milestone (M2) implements the readable HIP kernels
and the three-way CPU/HIP/PyTorch oracle comparison in `tests/ops/hip_ops_test.cpp`.
Loading