domainmgr, zedmanager, zedagent: SMT/NUMA-aware CPU placement for pinned applications - #6335
Draft
rucoder wants to merge 15 commits into
Draft
domainmgr, zedmanager, zedagent: SMT/NUMA-aware CPU placement for pinned applications#6335rucoder wants to merge 15 commits into
rucoder wants to merge 15 commits into
Conversation
Adds a standalone package that reads the machine's socket, physical-core, SMT-sibling, NUMA and L3 structure straight from sysfs, in pure Go with no CGO and no external topology library. Device classes EVE targets often ship neither, and the native dependency previously considered for this proved fragile on client and non-server SKUs. SMT siblings are identified by a shared (socket, core_id) key, which is authoritative on every architecture we target. Grouping by a cache id would be wrong: on Intel hybrid parts an efficiency-core module exposes one shared L2 across four distinct physical cores with no SMT, which such a key would model as a single four-thread core. Discovery degrades to a flat model rather than failing when sysfs cannot be read, so a caller always has a usable topology and simply loses the locality guarantees it cannot substantiate. The package deliberately depends on nothing else in pillar so the allocator, the hardware inventory and a future cluster-side consumer can share it. Signed-off-by: Mikhail Malyshev <mike.malyshev@gmail.com>
Replaces the CPU allocator with one that understands the machine's topology, so a workload asking for dedicated CPUs can be given whole physical cores in a NUMA-local, SMT-aware way rather than an arbitrary set of logical CPUs. Placement is computed for the whole set of pinned workloads at once and ordered by how constrained each one is -- whole-core-SMT first, since it can only use a core that really has two hardware threads and on a hybrid or SMT-disabled machine most cores cannot, then one-per-core, then anything thread-granular. The result is therefore a function of the request set rather than of the order requests arrive in. Allocating incrementally meant whichever workload activated first won the scarce cores, so a flexible workload could take the only SMT-capable core and leave a workload that needs one unplaceable -- and the same set of workloads could land differently on each boot. Plan does not mutate the allocator: the caller reserves an assignment when the workload actually starts, which is what lets a workload that has not started yet, or starts late, still claim the CPUs set aside for it. Score ranks an assignment by what actually costs performance -- NUMA nodes spanned, then last-level caches -- and deliberately not by which CPU indices were used. Many assignments share the best score, so comparing indices would report a workload as mis-placed merely because its first-choice CPUs were taken, and demand a restart that changes nothing. A core is withheld when any of its siblings is reserved for EVE. That costs capacity, so the shortage message says as much: handing out a core whose sibling runs housekeeping would reintroduce exactly the interference whole-core placement is bought to remove. The shortage also carries how many cores were needed against how many were free, so a caller can explain the refusal without computing a second, differently-filtered count. PoolUtilization reports the housekeeping, dedicated and isolated pools with both their CPU sets and their whole-core counts. Free threads alone answer "will it fit?" wrongly: threads left on partially-owned cores cannot satisfy a request for whole cores. Signed-off-by: Mikhail Malyshev <mike.malyshev@gmail.com>
Introduces the device-internal representation of a workload's CPU placement intent, mirroring the Kubernetes CPUManager and Topology Manager terms the controller API uses: cpu policy, full-pcpus-only, threads per core, NUMA policy, IO placement, isolation tier and disruption policy. Intent is kept deliberately separate from the allocator's vocabulary. Intent says what a workload needs; the allocator decides which host CPUs it gets. Keeping them apart means the wire format never dictates the placement mechanism, and it lets the two sources of intent -- the controller and the operator-editable /persist override -- resolve into one representation. The zero value means no policy was sent, so VmConfig.CPUsPinned alone keeps deciding and behaviour is unchanged for a controller that sets none of this. DomainStatus gains the resulting guest topology, the per-vCPU host CPU mapping, the emulator CPU set and the achieved placement quality. Quality is status rather than an error: a sub-optimally placed workload runs normally, and whether the improvement is worth a restart is a judgement for an operator. Adds the error-code registry reported alongside the free-text description, so a controller can distinguish conditions that need different responses -- a shortage a repack would fix, one nothing would fix, and a request that can never be satisfied -- without pattern-matching prose. ErrorDescription carries the code and a retry condition through to the wire. Signed-off-by: Mikhail Malyshev <mike.malyshev@gmail.com>
Realizing a whole-core placement on QEMU/KVM has three parts. The guest is launched with an -smp topology computed from the assignment, so software inside it sees the real SMT structure and can place its own hot work on non-sibling cores. A poll-mode datapath deliberately runs a worker on each sibling; without a truthful topology it cannot tell which vCPUs share a core. Each vCPU thread is then pinned 1:1 to its assigned host CPU. QEMU is already started paused, so the vCPU threads exist while the guest has not executed and there is no pre-pin race. The guest-vCPU-to-host-thread mapping comes from QMP query-cpus-fast, which is the only place it exists: QEMU does not name its vCPU threads unless started with debug-threads=on, and a domain's thread group also holds vhost_task helpers that modern kernels create as user threads in that same group, indistinguishable from vCPU threads by name or by flags. The pin is applied after the cgroup cpuset has been written and before the guest is released, so it is not undone by the cpuset. Under io_placement=housekeeping the non-vCPU threads are pinned off the hot cores, so device emulation cannot steal cycles from a busy vCPU. A virtio-blk iothread keeps disk IO off the main loop. Kubevirt reports that it cannot bind individual vCPUs. The capability is separate from plain cpuset confinement, because a hypervisor that can confine a domain to a set of CPUs may still be unable to bind one vCPU to one CPU or to advertise the resulting topology -- and accepting a whole-core request it cannot apply would report the workload as optimally placed while nothing was pinned. Signed-off-by: Mikhail Malyshev <mike.malyshev@gmail.com>
The CPU inventory emitted one entry per physical core with the core id in the field meant for a logical CPU id, no frequency and no topology. That is worse than incomplete: the ids were not the ones CPU affinities are expressed in, and the SMT structure -- the thing a consumer reasoning about CPU placement needs most -- was absent entirely. It now reports one entry per logical CPU carrying its socket, physical core, NUMA node and L3 domain, taken from the same topology discovery the allocator uses so the report and the behaviour cannot drift, plus base and maximum frequency where the kernel exposes them. Cache domains are reported with the set of CPUs sharing each one, which is what tells a consumer which workloads would contend for the same cache. The per-CPU sysfs views are collapsed into one entry per real cache instance. Kernel-level CPU isolation is reported separately as a node fact rather than a CPU one, and read from sysfs rather than parsed out of the command line, so it describes what the kernel is actually doing. The two differ when a parameter is malformed or capped, which is exactly when a consumer needs the truth. Topology discovery failing degrades to the previous flat listing instead of failing the whole inventory, which is still useful on a platform whose sysfs layout we cannot read. Signed-off-by: Mikhail Malyshev <mike.malyshev@gmail.com>
Maps the VmConfig CPU placement fields onto the device-internal intent and derives CPUsPinned from it. Two properties matter. An unrecognised enum value from a newer controller degrades to "no preference" rather than being rejected, which is safe because a controller is expected to gate on the capability reports below. And a dedicated policy is self-sufficient: it implies pinning on its own, so a workload no longer has to set the legacy pin_cpu flag as well for its CPUs to actually be pinned. With no policy sent, pin_cpu decides exactly as before. Advertises API_CAPABILITY_CPU_PLACEMENT_POLICY. Until this is reported a controller has no way to know the device honours the placement fields at all, which is precisely the failure the existing enforced-network-interface-order capability guards against, and it is what makes the fail-open behaviour above sound. Reports the node's CPU pool utilization on device info, per pool, with both the CPU sets and the whole-core counts, so a controller can answer "will this fit?" before a deploy and explain a shortage after one. This is dynamic state, so it rides the change-driven message rather than the cached hardware inventory. Surfaces a sub-optimal placement per application as a non-fatal advisory. It is converted to an ErrorInfo only at the wire, and never placed in the status error fields, because those are read as fatal in several places and a workload whose placement is merely improvable must not be torn down for it. Signed-off-by: Mikhail Malyshev <mike.malyshev@gmail.com>
CPU placement has to be a function of the configured set of pinned workloads, but domainmgr only ever sees a DomainConfig, and a DomainConfig cannot exist before a workload's volumes are resolved -- it carries the disk list. So during boot, or while images download at different rates, whichever workload was ready first was placed as though it were alone and took cores the full plan would have assigned elsewhere. The same set of workloads landed differently on each boot. zedmanager knows the whole picture much earlier: it holds every AppInstanceConfig, it owns the profile resolution that decides what is meant to run, and it is the component that withholds the DomainConfig in the first place. It now publishes that demand set -- one aggregate object naming every workload intended to run with its CPU intent -- as soon as the config is resolved, with no dependence on volumes. The set is published as a single object rather than one item per workload on purpose. Per-workload items would leave the consumer planning over whatever had arrived so far, which is the same ordering bug on a faster topic. An empty set is published explicitly, so "no pinned workloads" is distinguishable from "zedmanager has not spoken yet". A workload that is configured but not activated is left out: its cores belong to the workloads that do run, exactly as an assigned PCI device returns to the pool when its workload stops. Also fixes a pre-existing bug this work depends on. The start moment of a delayed workload was computed from a base time set only when zedmanager processed a controller-status message, and the app config regularly won that race -- leaving a start moment derived from the zero time, which is always in the past, so the delay was silently dropped and never recomputed. The base time is now established on first use, so a workload created before that message arrives gets the same start moment as one created after. Signed-off-by: Mikhail Malyshev <mike.malyshev@gmail.com>
…told Placement now runs over the whole demand set published by zedmanager rather than over whichever DomainConfigs have arrived. The result for a given set of workloads is therefore the same regardless of the order they were configured, started or delayed in, and the same across a reboot -- properties an operator depends on when a workload's performance was validated against a specific placement. The plan is derived, never stored. Persisting an assignment would create a second source of truth that can disagree with the hardware after a CPU is offlined, a NUMA node changes or the config changes, and the failure mode of stale placement data is silent and hard to diagnose. Determinism comes from ordering the batch by how constrained each workload is and breaking ties on the workload's identity, so recomputation reproduces the same answer. A whole-core request consumes every thread of its cores. When only one thread per core is wanted, the sibling is parked -- held by that workload and offered to nobody. This is the point of asking for a whole core: a best-effort workload running on the parked sibling would evict the cache lines and contend for the execution units the request exists to protect. Parked threads are reported as consumed in the pool utilization rather than as spare capacity, so a controller sees the true remaining headroom. Placement failures are terminal. A workload that cannot be placed stops with an error naming the cause -- a shortage, a shortage a repack would fix, or a request nothing could satisfy -- and stays stopped until an operator changes the config, which is how EVE already treats an unavailable PCI device. Retrying would silently place the workload the moment some unrelated workload happened to release cores, at an arbitrary time, with no operator awareness that its performance envelope had changed. Two long-standing behaviours are corrected. The operator-editable override on /persist can now enable pinning for a workload the controller did not pin, not only disable it, which is what makes it usable for on-device diagnosis. And housekeeping IO placement no longer draws its CPUs from a set that could include cores already promised to another workload. Signed-off-by: Mikhail Malyshev <mike.malyshev@gmail.com>
…bserve it A test cannot assert anything about CPU placement on a node whose CPU topology it does not control: with a flat single-thread-per-core VM every placement policy looks alike, and whole-core, one-per-core and parked-sibling behaviour are indistinguishable. The device requirements gain a threads-per-core knob, so a test can ask for a node with real SMT siblings, and the QEMU and libvirt providers derive the -smp topology from the requested CPU count and that knob through one shared helper -- the two providers disagreeing about what a requirement means would make results depend on which one ran. On the observation side, tests get the node facts CPU placement work needs: the host's socket/core/sibling topology, which CPUs are online, which the kernel isolated, and the kernel command line, so an expectation can be stated in terms of what the node actually is rather than hard-coded numbers. Per-workload facts come from QMP over the existing SSH transport. The guest-vCPU-to-host-thread mapping is only available there -- QEMU does not name its vCPU threads, and a domain's thread group contains helper threads that cannot be told apart from vCPU threads by name. A QMP call is also a point-in-time question with a definite answer, which is what an assertion wants, where waiting for a log line is a race dressed up as a check. The call is bounded by closing the connection from a timer, since deadlines are not supported on SSH channels. The application config gains the CPU placement policy fields and a start delay. The delay exists to test the property that matters most here: that placement does not depend on the order workloads happen to start in. Signed-off-by: Mikhail Malyshev <mike.malyshev@gmail.com>
Covers the three placement shapes a controller can ask for, each asserted against the host's real topology rather than against expected CPU numbers. One-per-core: as many distinct physical cores as vCPUs, every vCPU pinned to exactly one host CPU, no two vCPUs sharing a core. Whole-core-SMT: both siblings of each core become vCPUs, and the guest's own view of its topology matches how it was actually pinned -- a guest told it has siblings that are not siblings will co-schedule work that then contends. The multi-app case is the one that catches interference: a whole-core-SMT app, a one-per-core app and a best-effort app deployed together must land on disjoint CPUs and disjoint physical cores, with housekeeping CPUs still available to the system. A test on any single app in isolation would pass while the allocator handed the same core to two workloads. Each reachable app needs its own forwarded edge-node port, since the port belongs to the node. Signed-off-by: Mikhail Malyshev <mike.malyshev@gmail.com>
The property an operator relies on is that a validated placement stays put. This test asserts it three ways on one unchanged set of applications: across a reboot, across a staggered start where one app is deliberately delayed, and across a restart in the reverse order. Every vCPU must land on the same host CPU each time. Order independence is what makes reboot stability real rather than incidental. Boot orders vary with image download times, network readiness and configured start delays, so a placement derived from arrival order would be reproducible only by luck. The delayed-start case exercises exactly the window in which a workload's config is known but its domain does not exist yet. Signed-off-by: Mikhail Malyshev <mike.malyshev@gmail.com>
When a workload asks for one thread per physical core, the other thread of each of its cores is deliberately left idle. That thread is consumed, not free: a best-effort workload placed on it would evict the cache lines and compete for the execution units the request exists to protect, which is the whole reason for asking for a whole core. Asserted three ways, because each alone is insufficient: no other workload gets the parked thread in its cpuset, the node does not advertise it as free capacity to a controller, and nothing is ever observed executing on it. The middle one is what stops a controller from confidently over-committing the node. Signed-off-by: Mikhail Malyshev <mike.malyshev@gmail.com>
A node can have enough free threads for a whole-core workload and still not have a single free whole core, because earlier thread-granular workloads left one thread busy on each. The two shortages call for opposite responses: nothing will help the first, while rearranging existing workloads would resolve the second, and only the workloads' owner can decide whether that disruption is acceptable. The test fragments the node deliberately, confirms the refusal carries cpu.placement.needs_repack rather than a plain shortage, and then repacks and confirms the workload really does run -- so the advice the code gives is demonstrated to be true, not merely plausible. Signed-off-by: Mikhail Malyshev <mike.malyshev@gmail.com>
A request that cannot be honoured must be refused, not approximated. Silently falling back to a weaker placement would hand back a running workload whose timing guarantees are gone, with nothing in the reported state to say so -- the worst outcome available, because it looks like success. Each class of unsatisfiable request is checked separately with its own error code, since a controller needs to distinguish a request that is malformed from one the node cannot support from one it merely has no room for. The workload must never boot, and the node's dedicated CPU pool must be unchanged afterwards: a refused request that leaked cores would shrink the node's capacity with every retry. The refusal also has to persist. A placement failure that healed itself as soon as some unrelated workload released cores would start the workload at an arbitrary moment with a placement nobody validated, so this asserts the workload stays down until its config changes -- the same way EVE already treats a workload whose PCI device is unavailable. Signed-off-by: Mikhail Malyshev <mike.malyshev@gmail.com>
The device-side code in this branch needs API that has not landed upstream yet: the VmConfig CPU placement fields, the CPU topology and capability reporting on device info, ZInfoDevice.cpu_pools, API_CAPABILITY_CPU_PLACEMENT_POLICY and ErrorInfo.error_code. Without them nothing here compiles, so this commit points pillar and evetest at the fork carrying the proposed API (lf-edge/eve-api#155) through a replace directive, and vendors it. This commit exists only so the branch can be built, run and reviewed while the API is under discussion. It must be dropped and replaced by an ordinary `make bump-eve-api` once the API lands in lf-edge/eve-api: a replace directive pointing at a personal fork breaks dependency tracking and SBOM/licensing, and is never acceptable on master. Signed-off-by: Mikhail Malyshev <mike.malyshev@gmail.com>
github-actions
Bot
requested review from
OhmSpectator,
eriknordmark,
jsfakian,
milan-zededa,
rene and
shjala
August 17, 2026 17:36
This was referenced Aug 17, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Places pinned application vCPUs on whole physical cores, deterministically, and
reports to the controller what the node can do and what each workload got.
Draft. It depends on API that is still under review (lf-edge/eve-api#155) and
on an adam bump so the new info fields survive ingest (lf-edge/adam#158). The last
commit here,
pillar, evetest: build against the CPU-placement eve-api [DO NOT MERGE], points pillar and evetest at the fork carrying that API so thebranch builds and runs today; it must be dropped and replaced by an ordinary
make bump-eve-apibefore this can merge.What a workload gets
A workload asking for
cpu_policy=dedicated, full_pcpus_onlyis given wholephysical cores. With
threads_per_core=2both SMT siblings become vCPUs and theguest is launched with an
-smptopology that tells it truthfully which vCPUsare siblings — software that places its own hot work cannot do so against a
fabricated topology. With
threads_per_core=1the sibling is parked and staysconsumed by that workload: a best-effort neighbour placed there would evict its
cache lines and contend for its execution units, which is the reason to ask for a
whole core in the first place.
Each vCPU is then pinned 1:1 to its host CPU. QEMU is already started paused, so
the pin lands while the guest has not executed. The guest-vCPU-to-host-thread
mapping comes from QMP
query-cpus-fast, the only place it exists: QEMU does notname its vCPU threads, and a domain's thread group also holds
vhost_taskhelpersthat modern kernels create as user threads indistinguishable from vCPU threads by
name. Under
io_placement=housekeepingthe non-vCPU threads are kept off the hotcores.
Why placement is planned for the whole set
Placement is computed over every workload the controller intends to run, not over
the workloads that happen to have activated. domainmgr only ever sees a
DomainConfig, which cannot exist before a workload's volumes are resolved, so
whichever workload was ready first used to be placed as though it were alone and
took cores the full plan would have assigned elsewhere — the same set of
applications landed differently on each boot, and after each image download race.
zedmanager knows the whole picture much earlier, so it publishes the demand set —
one aggregate object naming every workload intended to run with its CPU intent —
as soon as config is resolved. domainmgr plans over that. The result is that
placement depends only on the configured set: same set, same host CPUs, across a
reboot, across a staggered start, and in any restart order.
The plan is derived, never persisted. A stored assignment would become a second
source of truth that silently disagrees with the hardware after a CPU is
offlined, a NUMA node changes, or the config changes. Determinism comes from
ordering the batch by how constrained each workload is and breaking ties on
identity.
A workload that is configured but not activated releases its cores, exactly as an
assigned PCI device returns to the pool when its workload stops.
Failing closed
A placement that cannot be honoured is refused, never approximated: silently
handing back a running workload whose timing guarantees are gone is the worst
available outcome, because it looks like success. Each refusal carries a
machine-readable code (
types/errorcodes.go) and a retry condition saying whatwould change the answer, distinguishing a shortage a repack would fix from one
nothing would fix from a request that can never be satisfied.
Failures are terminal. The workload stays down until the config changes, which is
how EVE already treats a workload whose PCI device is unavailable; retrying would
start it at an arbitrary later moment on a placement nobody validated.
Reporting
api_capabilityadvertisesCPU_PLACEMENT_POLICY, without which a controllercannot know the device honours the config fields; the device info carries the real
CPU topology, cache domains, kernel CPU isolation and per-pool CPU utilization
(with parked threads counted as consumed, so the reported headroom is true); and a
sub-optimally placed workload is reported as a WARNING-severity advisory rather
than an error, since it runs normally and whether the improvement is worth a
restart is an operator's judgement.
Testing
make -C pkg/pillar testpasses (the CPU placement paths are covered by unittests in
cputopology,cpuallocator,types,cmd/domainmgr,cmd/zedagent,cmd/zedmanagerandhypervisor).real SMT topology: one-per-core and whole-core-SMT placement; three
differently-policied applications sharing a node on disjoint cores; stability
across a reboot, a delayed start and a reverse restart order; a parked sibling
being withheld from every other workload and from the reported free capacity; a
fragmented node reporting
needs_repackand a repack really letting theworkload run; and every class of unsatisfiable request being refused with its
own error code, never booting, and leaving the node's dedicated pool unchanged.
and to observe host topology, kernel isolation, cgroup cpusets and per-vCPU
affinities.
Two pre-existing bugs are fixed along the way: zedmanager dropped an application's
start_delay_in_secondswhenever its config arrived before the controller-statusmessage that set the delay base time, and the housekeeping IO placement could draw
CPUs from a set that included cores already promised to another workload.