Measure real XMX (matrix engine) utilization on Intel Arc GPUs, per precision, while any workload runs.
Intel's tooling makes this surprisingly hard. xpu-smi/xpumcli don't implement
the EU/XMX metrics for Arc B-series. VTune's GPU plugin can't attach to an
already-running process, and its instruction-count mode costs 100-200x slowdown.
unitrace reads the right counters but has to launch the target itself.
xmxmon takes a different approach: it opens a Level Zero metric streamer on the device and samples hardware counters device-wide. It never launches, wraps, injects into, or even knows about the workload process. That makes it backend-agnostic — it works the same for any framework that touches the GPU (SYCL, PyTorch XPU, OpenVINO, oneDNN, your own Level Zero code), including ones running inside other containers.
The counters that matter:
XVE_INST_EXECUTED_XMX_INT2 XVE_INST_EXECUTED_XMX_FP16
XVE_INST_EXECUTED_XMX_INT4 XVE_INST_EXECUTED_XMX_BF16
XVE_INST_EXECUTED_XMX_INT8
plus XVE_ACTIVE, per-ALU-pipe instruction counts, occupancy, L3, and memory
bandwidth — all from the same VectorEngineProfile group.
Developed against Arc Pro B70 and B50. Other Arc / Data Center GPU Max parts that
expose VectorEngineProfile should work; check with xmxmon --list.
- An Intel GPU with the compute runtime installed (
libze1,libze-intel-gpu1) intel-metrics-discoveryandintel-metrics-library- Performance counters unlocked for non-root:
Persist with a file in
sudo sysctl -w dev.xe.observation_paranoid=0 # Xe driver sudo sysctl -w dev.i915.perf_stream_paranoid=0 # i915 driver
/etc/sysctl.d/. Without this, opening the streamer fails. - To build on the host:
g++andlibze-dev. The container image handles all of the above except the sysctl, which is a host setting.
docker build -t xmxmon .
# What devices and metric groups do I have?
docker run --rm --device /dev/dri xmxmon --list
# Sample device 0 for 60s while your workload runs, then summarize:
docker run --rm --device /dev/dri -v "$PWD:/data" xmxmon \
--device 0 --group VectorEngineProfile --period-ms 100 \
--duration 60 --out /data/run.csv
docker run --rm -v "$PWD:/data" --entrypoint /app/xmx-summary.py xmxmon /data/run.csvOnly --device /dev/dri is needed — no privileged mode, no host PID namespace,
no changes to the container running your workload.
For anything beyond a one-off capture, run the daemon instead. It keeps samplers alive across reboots and lets you start and stop captures on demand, so you don't have to predict when to attach:
cp xmxmon.yaml.example xmxmon.yaml # your copy; gitignored, survives git pull
mkdir -p captures
docker compose up -d # starts, and comes back after a rebootNo docker compose? Some distributions (including Ubuntu's docker.io
package) ship Docker without the Compose plugin — docker compose version
will say unknown command. Either install it (sudo apt install docker-compose-v2, or Docker's own docker-ce packages), or skip Compose
entirely — this is the same thing in one command:
docker run -d --name xmxmon --restart unless-stopped \
--device /dev/dri \
-p 127.0.0.1:9143:9143 \
-v "$PWD/xmxmon.yaml:/etc/xmxmon.yaml:ro" \
-v "$PWD/captures:/data/captures" \
--entrypoint /app/xmxmond.py \
xmxmonCheck it, watch it live, then wrap a benchmark with a tagged capture:
curl -s localhost:9143/now # is it alive?
./xmxmon-tui.py http://localhost:9143 # live terminal view
curl -s -X POST localhost:9143/capture -d '{"name":"my-bench","device":0}'
./run-my-benchmark.sh
curl -s -X POST localhost:9143/capture/stop -d '{"device":0}'
curl -s localhost:9143/captures # where the file landedBetween captures it idles at a coarse sampling period, so leaving it up costs
very little. Edit your xmxmon.yaml to change devices, periods, or to enable the
web UI and Prometheus exporter — see Configuration and
Daemon mode below, and restart the container to pick up changes.
Stop it with docker compose down (or docker rm -f xmxmon if you started it
with docker run).
Note that while the daemon is running it owns the device's performance counters, so stop it before using the one-shot CLI on that same GPU.
g++ -O2 -o xmxmon xmxmon.cpp -lze_loader
./xmxmon --list # devices + metric groups
./xmxmon --device 0 --group VectorEngineProfile --list-metrics # + descriptions
./xmxmon --device 0 --group VectorEngineProfile --period-ms 100 --out run.csv
# Ctrl-C when done — rows are written as they arrive, nothing is lost.
./xmx-summary.py run.csvxmx-summary.py answers the first question directly:
Real capture of a PyTorch XPU fp16 4096³ matmul loop, sampled from outside the container running it:
460 reports total, 411 with GPU_BUSY > 1% (stats below are over busy reports only)
--- XMX (matrix engine) ---
XVE_INST_EXECUTED_XMX_INT2 total 0 avg 0 peak 0
XVE_INST_EXECUTED_XMX_INT4 total 0 avg 0 peak 0
XVE_INST_EXECUTED_XMX_INT8 total 0 avg 0 peak 0
XVE_INST_EXECUTED_XMX_FP16 total 5.226e+12 avg 1.272e+10 peak 1.63e+10 <-- ACTIVE
XVE_INST_EXECUTED_XMX_BF16 total 0 avg 0 peak 0
VERDICT: XMX IS being used (see precisions above)
--- utilization ---
GPU_BUSY avg 99.23 peak 100
XVE_ACTIVE avg 50.27 peak 50.91
...
Two things worth knowing before you interpret a capture:
Sample a representative workload. Matrix engines light up on matrix-shaped work. An LLM decoding one token at a time is doing matrix-vector products, and XMX will read as legitimately idle — that says nothing about whether the backend uses XMX for batched work. Capture a batch/prefill-heavy phase before concluding anything.
Mixed-precision instructions are counted under their wider operand. A DPAS with,
say, 8-bit activations against 2-bit weights increments the INT8 counter, not INT2.
So XMX_INT2 staying at zero does not mean 2-bit data never reached the matrix
engine — reaching that counter requires both operands to be 2-bit.
VectorEngineProfile answers "is the matrix engine being used, and how
efficiently" — but a device exposes a dozen groups (xmxmon --list), and only
one can be sampled at a time. Three of them profile a compute workload from
angles the XMX counters can't see, and xmxmon now ships a matched view for each:
the TUI hero column, the web-UI tiles and charts, the --detailed overhead
ratios, and a Grafana dashboard all follow whichever group a device samples.
| Group | Answers | Highlights |
|---|---|---|
VectorEngineProfile |
matrix-engine use and operand-prep tax | XMX per precision, prep/XMX, arithmetic intensity |
ComputeBasic |
where the whole pipeline spends time | XVE pipes + L1/L3 hit rates + VRAM + PCIe/sysmem + dispatch, all at once |
MemoryProfile |
the memory subsystem in depth (the decode bottleneck) | read/write bandwidth, 32B-vs-64B coalescing, L3 flow-control saturation, copy engine |
DeviceCacheProfile |
L3 behaviour and who drives it | L3 reads attributed to load/store vs instruction vs sampler, and the VRAM traffic behind misses |
Select per device in xmxmon.yaml — the group: default plus an optional
groups: map that overrides individual cards, so one GPU can report XMX while
another reports memory or cache:
group: VectorEngineProfile
groups:
1: MemoryProfile # device 1 samples the memory subsystem insteadSwitch on the fly. You don't have to edit the config and restart to change
lens. The web UI has a group dropdown in each device header, with a "sync all"
checkbox beside it — check it (it links across every card) and a dropdown change
switches all GPUs at once. The TUI's g key
opens a picker: choose the device — with "apply to all devices" as the first
entry — then the group. Both hit POST /group. The change is in-memory only, so the next
daemon start returns to whatever xmxmon.yaml specifies. Two caveats, both from
the hardware: counters are device-wide, so the switch affects every viewer of
that GPU, not just you; and it is refused while that device is capturing (stop the
capture first, so the ndjson keeps a single schema).
The picker offers every group the device exposes (the daemon enumerates them
at startup), with the four profiled groups plus VectorEngineStalls listed
first. EuStallSampling and TestOa are omitted — they use a different sampling
mechanism, not the time-based streamer, and would fail to open. A group without a
curated view (e.g. RenderBasic) still samples fine; its metrics show through the
default view and the raw-counter panel, just without a purpose-built layout.
When you switch all devices at once, only groups common to every card are offered; a group unique to one card is struck through (TUI) or disabled (web UI, while "sync all" is checked), so a mixed-architecture fleet can never be sent a group one of its cards lacks. Per-device selection still offers each card's full set, and the daemon validates every switch against the target device's own list regardless.
Everything downstream keys off the active group automatically. xmx-summary.py
detects the group from a capture's columns, so offline analysis needs no flag.
There is nothing to configure beyond the group itself; a metric absent from the
active group is skipped rather than shown as zero.
Each profiled group has a hand-built, import-ready dashboard next to the default
one: grafana-dashboard.json (VectorEngineProfile), grafana-dashboard-compute.json,
grafana-dashboard-memory.json, and grafana-dashboard-cache.json. Every other
switchable group (stalls, render, depth, RT, …) has an auto-generated
grafana-dashboard-<group>.json built straight from its metric list — functional
(hero, levels, bandwidth, counters), just not hand-tuned. All read the same
generic xmxmon_* series, so enabling the exporter is all it takes.
Adding a fourth view is a data change, not a code change — see
docs/metric-groups.md for the profile format.
xmxmond.py keeps samplers running continuously and adds an HTTP API, so
benchmark scripts can turn high-rate capture on and off around a run
(docker compose up -d, as in Leave it running above).
Published on 127.0.0.1:9143 only. Always available:
| Endpoint | Purpose |
|---|---|
GET /now |
latest aggregate snapshot (JSON) |
POST /capture |
{"name":"run1","device":0,"duration_s":600} — switch to high-rate sampling, tee to a tagged ndjson in captures/, auto-revert |
POST /capture/stop |
{"device":0} — end early |
GET /captures |
running and finished captures |
POST /group |
{"device":0,"group":"MemoryProfile"} — switch metric group at runtime; reverts to config on restart |
Omit device on either capture call to act on every configured device at once.
A capture ends when its duration_s elapses or you stop it, whichever comes
first; the sampler then drops back to the idle period on its own.
./xmxmon-tui.py http://localhost:9143Live bars for XMX per precision, utilization, and bandwidth, with peak-hold markers.
It's a plain HTTP client, so it works fine over ssh and doesn't need GPU access
itself. Press d for the overhead detail (below), r to add raw counters,
q to quit; start expanded with --detailed.
Detail renders as a second column beside the live bars and fits an 80×24 terminal with two GPUs. Below 72 columns it stacks instead, and the metric grid reflows to whatever width is available.
Raw counters say what happened; the detail view says how much of it was useful
work. Available in the TUI (d), the web UI ("show detail"), and
xmx-summary.py --detailed, all computing the same ratios:
| Metric | Reads as |
|---|---|
prep work / XMX |
operand preparation (bit manipulation, integer and float ALU, extended math) per unit of matrix work. A dense fp16 matmul measures ~0.03 — treat that as the floor. Ratios in the single digits mean most instructions are unpacking and scaling, not multiplying. |
XMX per VRAM byte |
arithmetic intensity on the matrix engine. Low means memory-fed, not compute-fed. |
L3 hit rate |
low values mean prepared operands are spilling to VRAM instead of being reused in cache. |
memory ops / XMX |
load/store pressure per unit of matrix work. |
barrier share |
synchronization cost as a share of issued instructions. |
kernel dispatches |
launch overhead; high rates with little work per launch mean dispatch-bound. |
divergent issue, icache miss, multi-pipe active |
branch divergence, oversized/spilling kernels, and instruction-level parallelism. |
When the active group is VectorEngineStalls, the same view lists stall reasons
(stall: sbid, stall: sendwr, …) instead — why the engine was idle. Raw
per-second counters appear underneath the ratios so any number can be checked
against its inputs.
Both are off by default. Enable in xmxmon.yaml:
wui: true # live charts at GET /
prometheus: true # GET /metrics for scrapingNeither endpoint has authentication, so leave the bind address on loopback unless you've put something in front of it.
With prometheus: true, scrape it:
- job_name: xmxmon
scrape_interval: 5s
static_configs:
- targets: ['127.0.0.1:9143']Exported series:
xmxmon_rate_per_s{device,metric}— counters, already converted to per-secondxmxmon_gauge{device,metric}— levels (percentages, frequency)xmxmon_derived{device,metric}— the same overhead ratios the TUI and summary show (prep_work_per_xmx,xmx_per_vram_byte,l3_hit_rate, …), so the dashboard graphs identical numbers rather than re-deriving them in PromQLxmxmon_capturing{device}— 0/1
grafana-dashboard.json imports as a full dashboard: hero stat tiles with
sparklines, XMX-by-precision and memory as stacked areas, the operand-prep tax
and L3 hit rate as threshold-coloured lines, arithmetic intensity for the
compute-vs-memory story, and a shared crosshair so a spike in one panel lines up
with the dip it caused in another. A collapsed "Instruction mix & micro-overhead"
row holds ALU-pipe composition, stall reasons, and dispatch rate. On import Grafana asks
which Prometheus datasource to use; pick yours (it defaults to your default
Prometheus). The device filter defaults to "All".
Seeing "No data"? Check, in order: prometheus: true is set in your
xmxmon.yaml and the daemon was restarted (curl localhost:9143/metrics should
return text, not 404); the Prometheus target is up (Status → Targets, job
xmxmon); and the dashboard's datasource variable points at the Prometheus that
scrapes it. The exporter being off is the most common cause — nothing downstream
can show data that was never exported.
Configuration lives in two files you own, both created by copying a tracked
template and both gitignored — so your settings survive git pull and never end
up in a commit:
cp xmxmon.yaml.example xmxmon.yaml
cp docker-compose.override.yml.example docker-compose.override.yml # optionalxmxmon.yaml holds daemon settings: devices, metric group, sampling periods,
bind address, capture directory, and the two feature flags. See
xmxmon.yaml.example for the annotated list. On the host, point at any copy with
XMXMOND_CONFIG=/path/to.yaml.
Because only one metric group can be active per device, a groups: map assigns
different groups to different cards — letting one GPU report XMX and operand-prep
counters while another reports stall reasons.
docker-compose.override.yml holds machine-specific deployment changes — port
exposure, device pinning, capture volume. Compose merges it automatically when
present, so docker compose up -d needs no extra arguments. Leave it out
entirely and the tracked defaults apply.
Never edit docker-compose.yml or the .example files directly; upstream owns
those, and local edits there are exactly what git pull will fight with.
- One metric group per device at a time. That's a hardware constraint, not a
tool limitation. To see another group on a device, switch it (config, or
POST /group/ the TUI and web UI pickers at runtime) — you can't sample two at once. To compare two groups simultaneously, put them on two different devices. - One streamer per device. While the daemon is running it owns the device's counters; stop it before running the standalone CLI against the same GPU.
- Counters are device-wide. Two workloads sharing a GPU are summed together. Pin the workload under test to its own device.
VectorEngineProfileis not available on integrated GPUs — check--list.
docker run --rm --entrypoint /app/check-driver-updates.sh xmxmonRead-only: prints installed versions of the Intel GPU userspace packages and whether newer ones are available. Exits non-zero if updates exist. Installs nothing.
Apache-2.0. See LICENSE.