Skip to content

Cluster self management - #23

Merged
9bany merged 6 commits into
masterfrom
cluster-self-management
Sep 5, 2026
Merged

9bany merged 6 commits into
masterfrom
cluster-self-management

Conversation

@9bany

@9bany 9bany commented Sep 4, 2026

Copy link
Copy Markdown
Member

No description provided.

The demo cluster is two nodes now, `a` over `0..64` and `b` over `64..`,
with no copy of either range. `config.rs:486` only requires three nodes
once a `replica` is declared, so the two-node file is valid.

What goes with it is every claim that is no longer true. The compose
header promised `stop a` and the queries keep working; docs/proxy.md used
`docker compose stop a` as the demonstration the proxy exists for, and
called it "the only demonstration that matters". Neither holds without a
copy: the proxy takes the node out of rotation exactly as designed, and
shards 0..64 are then held by nobody, so the query fails at `b` rather
than at the front door. The proxy removed one single point of failure; an
unreplicated range is the other one, and it is not a proxy's to remove.

docs/proxy.md now shows what adding a third node as `replica = "a"` buys
and points at deploy/readme.md for it. Two pre-existing staleness fixes
ride along, both in text the change touches: "Only `a` is published" in
two places, which stopped being true when the proxy took the port.
The freelist was ordered by generation, so `alloc` handed out the oldest
free page before the lowest one. Correct, and hopeless for the tail:
writes kept landing at the end of the file while holes near the front
stayed holes, so nothing was ever flush against EOF for `truncate_tail`
to release and a file that had churned only ever grew.

Ordering the runs by page number instead makes ordinary writes fill from
the front. Nothing about reclaim changes: `alloc` still skips any run
newer than the horizon, so this decides *which* reusable page is taken
and never *whether* a page is reusable. The merge is unaffected - two
runs that meet in page order can have nothing between them - and so is
the termination argument for `alloc_freelist_pages`, which rests on
`alloc` taking from the front of a run rather than on the order of runs.

The rule is pinned against `Freelist` directly rather than through a
commit, for the reason the module header of `reclamation.rs` already
gives: a commit takes pages for its own chains before the caller's
`alloc` sees the list, so a page number observed through one says as
much about the machinery as about the rule.
… own

Three things a cluster could not do for itself, each behind its own flag
and each off by default, plus the correctness work they turned out to
need first.

**The agreement had stopped failing anything over.** `promotion` guarded
itself with `commit_index() != log().len() - 1`, comparing a log index
with a vector position. Once anything had been compacted away the two
could never agree again, so every promotion after the first compaction
was refused - silently, on a cluster reporting itself healthy. Fixed with
`Raft::settled()`, and the pure rule pulled out as `promotion_for` so it
can be driven against a simulated agreement. Compaction also dropped the
newest map and membership rather than folding them into the sentinel the
follower snapshot path already builds, and the restart replay compared a
vector position with the commit index; `State` now carries the seed, so
a node handed a snapshot and then restarted no longer falls back to a
cluster file that describes a cluster which is gone. Format `BIGRAFT3`,
reading `BIGRAFT2`.

**A node that had just started believed it held its leases.** `lease_at`
began at zero and so does the clock it is compared against, so a fresh
process read its own silence as having just been in touch - exactly a
node that restarted and may already have been replaced. It is `NEVER`
now. A replicated node therefore refuses its range until the first
heartbeat, which is one election at boot and is what the lease always
meant.

**Two nodes could intern at once.** `peer_intern` and `peer_allocate`
checked nothing, on the reasoning that the config decided who led - true
when the leader was a name in a file, false since it became a field in
the map. A coordinator one decision behind sent keys to whoever led then,
and one string got two row ids, which nothing downstream can see. The
request now carries who it thinks leads, the receiver checks the map, a
step-down, and a lease, and the manual move stands the source down before
it reads its floor rather than racing it.

**Record ids are reserved before they are handed out.** The floor lived
in the leader's memory, so a successor started from what was on disk and
re-issued ids promised to writes that had not landed - the one failure
with no second chance. The map carries a ceiling per table, raised a
block of 65 536 at a time; a leader with no floor of its own starts at
the ceiling.

On top of that: `--balance` lets the agreement's leader take one
balancing step at a time, including admitting a learner that has answered
and caught up, which `add_node` had promised and nothing did.
`--elect-schema-leader` lets it give the row-key namespace away after
fifteen seconds of silence, with the handover rebuilt from the survivors
and blocked loudly if two of them disagree about a row id. `--reclaim`
hands trailing free pages back while serving. All three run on one
steward thread, which is neither the agreement's - it must not block -
nor a worker's.

`SIGTERM` now stands the daemon down instead of killing it: `serve_while`
did the right thing already and only tests called it.

Known and documented rather than hidden: an automatic handover can burn
row ids the dead leader assigned to writes that never landed, so a
deposed node is marked behind and must be repaired before it is trusted
again; and `--reclaim` releases only pages flush against the end of the
file, so a database that is emptied and then left alone stays large.
`--reclaim` gave back almost nothing on a churned file: 4594 pages, 4427
of them free, two returned. Cause found by accounting for every page by
hand and coming up one short - page 4591, the highest in the file.

**A page freed above the tail floor never reached the freelist.**
`WriteTxn::free` sends it to a scratch list instead, so `alloc` can hand
it straight back to the same transaction, and nothing drained that list
at commit: `commit` passes dirty, freelist, roots, snapshots and catalog
to `commit_txn` and drops `scratch` on the floor. A transaction whose
tree shrinks frees more than it allocates - a leaf merge, a level
removed, `free_tree` - so what is left over is counted in `page_count`,
materialised in the file, and named by nothing. `truncate_tail` walks
down from the end and stops at the first page that is not a reusable
free run, so one of those on top pinned all 4425 free pages under it.

`Freelist::absorb_scratch` runs at the top of `commit`, before the first
allocation `commit_txn` makes: a page flush against the end of the file
is given back by lowering `next_pgno` rather than recorded as free, and
the rest are pushed with the committing transaction's own id, which
keeps them pending for the whole of their own commit and out of reach of
the freelist's own allocation loop. Placing it in `commit` rather than
`commit_txn` is what makes that ordering true; lowering `next_pgno`
after a chain page had been taken from the tail would hand one page
number out twice. readme §7.3 carries the argument.

**And a counter, so the next one is found rather than deduced.**
`Db::audit_pages` accounts for every page in a file as reachable, free,
or neither, and `big leaks <file>` prints it. The scaffold is captured
by `Store::begin_audit` under a single `state.read()` - meta, freelist,
snapshots and roots are published together by `commit_txn`, and a mark
set stitched from two states is wrong in a way no test would catch - and
`Db` walks the trees, because `big-pager` cannot see `big-btree`. It
holds a `ReadTxn` for the walk and takes no write lock: the counts are
exact as of the transaction it opened at, which is all a report needs.

The mark covers the class both `scrub` and `copy_to` skip: the trees a
snapshot still names through its own old roots chain. `copy_to` skips
them because it is building a fresh history; an audit asserts what the
existing one needs, so what `copy_to` throws away is what the mark has
to add. A bitset rather than a counter, because snapshots share nearly
every page with the current trees under copy-on-write and `visit_tree`
does not deduplicate.

It reports one thing nothing else in the tree checks. A page that is
both reachable and in a run older than the horizon has been handed out
twice, which is corruption rather than waste - that and a reference past
the end of the file are the only two conditions that fail the command.

Read-only, and no repair: the mark is taken under a reader, which is
sound for reporting and not for acting on.

Two things still pin the tail, neither of them a leak, both now written
down where the next person will look: the freelist chain a large delete
must place at the top of the file, which clears on the next real write;
and after that a single live b-tree page, which nothing can move without
rewriting whoever points at it.
Three shipped behind their own flags and none of them reached the demo,
so the file described a cluster one release behind itself.

`--reclaim` is on for both nodes. It is the one that needs nothing from
the cluster - no peer asked, nothing proposed - and the comment says
what it does not do as well: it gives back the free run at the end of
the file, so a database that has churned in the middle keeps those holes
until an offline `big compact`. Its floor is 4096 pages with a quarter
reusable, which a demo this small never reaches, and saying so is
cheaper than someone wondering why the counter never moves.

`--balance` and `--elect-schema-leader` stay off, each with its reason
where an operator will look for it. The second reason is the one worth
writing down: **with two voters it could not fire even if it were on.**
Handing the row-key namespace on is a decision the agreement commits,
a majority of two is two, and the silence it waits for is the same node
being gone. It wants the third node that replication wants.
The obvious shape - one StatefulSet per node, because a node here is not
interchangeable with the next one - was written first and thrown away.
It is the correct argument and the wrong conclusion: what it produces is
a deployment where growing the cluster means writing YAML, and the
cluster has been able to admit a learner and give it work by itself
since `--balance`. A knob that looks like scaling and produces an idle
stranger is worse than no knob; a shape that has no knob at all is worse
still.

So: one StatefulSet, pod names as node names, and the four things a
joining node needs split by who can actually supply them. **Only two of
them are Kubernetes's.** A certificate is issued in advance - names are
cheap, because a certificate grants nothing until the roster accepts it
and the roster follows the agreement rather than a file. The cluster
file is written per pod by an initContainer. The other two are the
cluster's own: `add-node` registers the pod as a learner, and the
balancer on the leader admits it and cuts the tail of the shard space
over to it.

**`add-node` runs before the daemon, and that ordering is the whole
reason there is an initContainer rather than a sidecar.** A node that
starts while the cluster has never heard of it is a voter in a cluster
of its own imagining, and it stands for election once per timeout,
raising the term on every node that hears it. Registered first, it is
reached as a learner, and a learner does not stand.

Three lists, three rules, and confusing them is the mistake this
deployment makes easy, so they are in one file with the difference at
the top. `seed.toml` is a membership and must be exact: a node with no
log counts a majority out of it, so naming a pod that is not running is
naming a vote that never arrives. `upstreams.toml` is an address space
and deliberately names pods that do not exist - the proxy polls, finds
them unreachable and leaves them out of rotation, which is what lets a
scaled-up node reach a client with no proxy restart. The certificates
are the third. Three tests pin all of it, including the file `join.sh`
generates: a file the daemon refuses is a pod that crash-loops with no
clue pointing back at a shell script inside a ConfigMap.

Three nodes and not the demo's two, because two is the size at which
nothing works: a majority of two is two, so neither a failover nor a
schema-leader handover can be committed at the moment one node is the
one that went away. `--balance` and `--elect-schema-leader` are on here
for the same reason they are off in `cluster/` - there, the two ranges
are the thing being explained; here the shape is meant to change.

`certs.sh` and `users.sh` grow `NODES=` and `SECRETS=` overrides, so
this deployment gets a CA and a users file of its own rather than a
second copy of either script: two deployments sharing one CA are two
deployments that can impersonate each other.

Scaling in is not the mirror image and is not automated: drain, wait,
remove, then scale. `cluster remove` refuses a node that still holds a
range, and that refusal is the point.
@9bany
9bany merged commit 76d57bd into master Sep 5, 2026
14 of 20 checks passed
@9bany
9bany deleted the cluster-self-management branch September 7, 2026 02:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant