Repository navigation
Conversation
Example 12 fixed the expert count at 32 and took per-expert row counts from an in-source 24x32 table, so the reported GFLOP/s could only be measured at one fill without editing the source. At a fixed --n/--k the printed rate is dominated by that fill: on Arc Pro B70 the N=5760 K=2880 leg reports a median 82.1 TFLOP/s with the built-in table versus 14.7 TFLOP/s at a uniform 30 rows/expert, while kernel time moves only 2.17 -> 3.25 ms. Routed MoE inference operates near the low end, so being able to sweep fill makes the number self-describing. Add --experts and --rows_per_expert, and print the effective expert count, fill mode and mean rows/expert. Defaults are unchanged: without the new flags the built-in table is used verbatim and the example reports 'Mean rows/expert : 250.824', matching the table literal, so existing invocations including the CMake test arguments behave as before. Signed-off-by: Rohit <knowrohit.work@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Feature
examples/12_xe20_moe_gemm_cute_interfacegains two options:--experts=<int>— number of experts per layer (default 32, unchanged)--rows_per_expert=<int>— give every expert this many rows, replacing thebuilt-in measured table (default 0 = use the built-in table, unchanged)
It also prints the effective configuration, so a result is self-describing:
Follow-up to #855.
Use Case
The example previously fixed
num_expertsat 32 and drew per-expert row countsfrom a fixed 24x32 in-source table, so the
GFLOPSit prints could only beobtained at one fill unless you edited the source.
That matters because at fixed
--n/--kthe printed rate is dominated by fill.On Arc Pro B70 (Battlemage G31), 32 experts,
--n=2880 --k=2880 --num_layers=24 --verify=0, median over 24 layers of theN=5760, K=2880leg:From uniform 30 to uniform 250 the FLOP count rises 8.3x and the reported rate
rises 7.7x, while kernel time rises 1.09x — the leg is bound by streaming expert
weights at these fills, so the rate largely reports
sum_e M_e. Routed MoEinference operates near the low end of that table (top_k over tens to hundreds
of experts), which is exactly the region that previously required a source edit
to reach.
API
Additive only. Both options are parsed with the existing
cutlass::CommandLineidiom next to
--n/--k/--num_layers/--verify, and validated with the sameerror = truepattern used byexamples/10_bmg_grouped_gemm_mixed_dtype:--expertsmust be positive--rows_per_expertmust be non-negative--expertsabove 32 requires--rows_per_expert, since the built-in tableonly has 32 columns; the diagnostic says so explicitly
main()gained the standardif (options.error) { ... return -1; }block thatsibling examples already use.
The per-layer row array became a
std::vector<int>because the expert count isnow a runtime value; the literal table is untouched and is now a
static const int [max_layers][kDefaultTableExperts]that seeds the vector when--rows_per_expertis not supplied.Example
Testing
Built and run on Arc Pro B70 (Battlemage G31), driver
1.15.39122+12, NEO26.27.39122.12, oneAPI 2026.1, JITspir64,SYCL_UR_USE_LEVEL_ZERO_V2=0.cmake --build build-sycl --target 12_xe20_moe_gemm_cute_interface.Mean rows/expert : 250.824. I independently brace-parsed the 24x32 tableliteral out of the source (768 entries, every inner group verified length 32)
and its mean is 250.8 — so the vector path reproduces the table exactly. The
CMake test arguments
--n=2880 --k=2880 --num_layers=24are unaffected andneed no change.
--experts=85without--rows_per_expertprints thediagnostic and aborts with
-1:--experts=85 --rows_per_expert={30,120,230}at--n=2048 --k=3072, all run to completion.Not tested: PVC, and
verify=1at--experts> 32 (the correctness referencepath was not touched, but I only exercised it at the default expert count).
ToDo
None that I'm aware of. If you'd rather have the flag spelled
--fill, or wantthe CMake test arguments extended to cover a second fill point in CI, say so and
I'll adjust.