Repository navigation
Cluster self management - #23
Merged
Merged
Conversation
The demo cluster is two nodes now, `a` over `0..64` and `b` over `64..`, with no copy of either range. `config.rs:486` only requires three nodes once a `replica` is declared, so the two-node file is valid. What goes with it is every claim that is no longer true. The compose header promised `stop a` and the queries keep working; docs/proxy.md used `docker compose stop a` as the demonstration the proxy exists for, and called it "the only demonstration that matters". Neither holds without a copy: the proxy takes the node out of rotation exactly as designed, and shards 0..64 are then held by nobody, so the query fails at `b` rather than at the front door. The proxy removed one single point of failure; an unreplicated range is the other one, and it is not a proxy's to remove. docs/proxy.md now shows what adding a third node as `replica = "a"` buys and points at deploy/readme.md for it. Two pre-existing staleness fixes ride along, both in text the change touches: "Only `a` is published" in two places, which stopped being true when the proxy took the port.
The freelist was ordered by generation, so `alloc` handed out the oldest free page before the lowest one. Correct, and hopeless for the tail: writes kept landing at the end of the file while holes near the front stayed holes, so nothing was ever flush against EOF for `truncate_tail` to release and a file that had churned only ever grew. Ordering the runs by page number instead makes ordinary writes fill from the front. Nothing about reclaim changes: `alloc` still skips any run newer than the horizon, so this decides *which* reusable page is taken and never *whether* a page is reusable. The merge is unaffected - two runs that meet in page order can have nothing between them - and so is the termination argument for `alloc_freelist_pages`, which rests on `alloc` taking from the front of a run rather than on the order of runs. The rule is pinned against `Freelist` directly rather than through a commit, for the reason the module header of `reclamation.rs` already gives: a commit takes pages for its own chains before the caller's `alloc` sees the list, so a page number observed through one says as much about the machinery as about the rule.
… own Three things a cluster could not do for itself, each behind its own flag and each off by default, plus the correctness work they turned out to need first. **The agreement had stopped failing anything over.** `promotion` guarded itself with `commit_index() != log().len() - 1`, comparing a log index with a vector position. Once anything had been compacted away the two could never agree again, so every promotion after the first compaction was refused - silently, on a cluster reporting itself healthy. Fixed with `Raft::settled()`, and the pure rule pulled out as `promotion_for` so it can be driven against a simulated agreement. Compaction also dropped the newest map and membership rather than folding them into the sentinel the follower snapshot path already builds, and the restart replay compared a vector position with the commit index; `State` now carries the seed, so a node handed a snapshot and then restarted no longer falls back to a cluster file that describes a cluster which is gone. Format `BIGRAFT3`, reading `BIGRAFT2`. **A node that had just started believed it held its leases.** `lease_at` began at zero and so does the clock it is compared against, so a fresh process read its own silence as having just been in touch - exactly a node that restarted and may already have been replaced. It is `NEVER` now. A replicated node therefore refuses its range until the first heartbeat, which is one election at boot and is what the lease always meant. **Two nodes could intern at once.** `peer_intern` and `peer_allocate` checked nothing, on the reasoning that the config decided who led - true when the leader was a name in a file, false since it became a field in the map. A coordinator one decision behind sent keys to whoever led then, and one string got two row ids, which nothing downstream can see. The request now carries who it thinks leads, the receiver checks the map, a step-down, and a lease, and the manual move stands the source down before it reads its floor rather than racing it. **Record ids are reserved before they are handed out.** The floor lived in the leader's memory, so a successor started from what was on disk and re-issued ids promised to writes that had not landed - the one failure with no second chance. The map carries a ceiling per table, raised a block of 65 536 at a time; a leader with no floor of its own starts at the ceiling. On top of that: `--balance` lets the agreement's leader take one balancing step at a time, including admitting a learner that has answered and caught up, which `add_node` had promised and nothing did. `--elect-schema-leader` lets it give the row-key namespace away after fifteen seconds of silence, with the handover rebuilt from the survivors and blocked loudly if two of them disagree about a row id. `--reclaim` hands trailing free pages back while serving. All three run on one steward thread, which is neither the agreement's - it must not block - nor a worker's. `SIGTERM` now stands the daemon down instead of killing it: `serve_while` did the right thing already and only tests called it. Known and documented rather than hidden: an automatic handover can burn row ids the dead leader assigned to writes that never landed, so a deposed node is marked behind and must be repaired before it is trusted again; and `--reclaim` releases only pages flush against the end of the file, so a database that is emptied and then left alone stays large.
`--reclaim` gave back almost nothing on a churned file: 4594 pages, 4427 of them free, two returned. Cause found by accounting for every page by hand and coming up one short - page 4591, the highest in the file. **A page freed above the tail floor never reached the freelist.** `WriteTxn::free` sends it to a scratch list instead, so `alloc` can hand it straight back to the same transaction, and nothing drained that list at commit: `commit` passes dirty, freelist, roots, snapshots and catalog to `commit_txn` and drops `scratch` on the floor. A transaction whose tree shrinks frees more than it allocates - a leaf merge, a level removed, `free_tree` - so what is left over is counted in `page_count`, materialised in the file, and named by nothing. `truncate_tail` walks down from the end and stops at the first page that is not a reusable free run, so one of those on top pinned all 4425 free pages under it. `Freelist::absorb_scratch` runs at the top of `commit`, before the first allocation `commit_txn` makes: a page flush against the end of the file is given back by lowering `next_pgno` rather than recorded as free, and the rest are pushed with the committing transaction's own id, which keeps them pending for the whole of their own commit and out of reach of the freelist's own allocation loop. Placing it in `commit` rather than `commit_txn` is what makes that ordering true; lowering `next_pgno` after a chain page had been taken from the tail would hand one page number out twice. readme §7.3 carries the argument. **And a counter, so the next one is found rather than deduced.** `Db::audit_pages` accounts for every page in a file as reachable, free, or neither, and `big leaks <file>` prints it. The scaffold is captured by `Store::begin_audit` under a single `state.read()` - meta, freelist, snapshots and roots are published together by `commit_txn`, and a mark set stitched from two states is wrong in a way no test would catch - and `Db` walks the trees, because `big-pager` cannot see `big-btree`. It holds a `ReadTxn` for the walk and takes no write lock: the counts are exact as of the transaction it opened at, which is all a report needs. The mark covers the class both `scrub` and `copy_to` skip: the trees a snapshot still names through its own old roots chain. `copy_to` skips them because it is building a fresh history; an audit asserts what the existing one needs, so what `copy_to` throws away is what the mark has to add. A bitset rather than a counter, because snapshots share nearly every page with the current trees under copy-on-write and `visit_tree` does not deduplicate. It reports one thing nothing else in the tree checks. A page that is both reachable and in a run older than the horizon has been handed out twice, which is corruption rather than waste - that and a reference past the end of the file are the only two conditions that fail the command. Read-only, and no repair: the mark is taken under a reader, which is sound for reporting and not for acting on. Two things still pin the tail, neither of them a leak, both now written down where the next person will look: the freelist chain a large delete must place at the top of the file, which clears on the next real write; and after that a single live b-tree page, which nothing can move without rewriting whoever points at it.
Three shipped behind their own flags and none of them reached the demo, so the file described a cluster one release behind itself. `--reclaim` is on for both nodes. It is the one that needs nothing from the cluster - no peer asked, nothing proposed - and the comment says what it does not do as well: it gives back the free run at the end of the file, so a database that has churned in the middle keeps those holes until an offline `big compact`. Its floor is 4096 pages with a quarter reusable, which a demo this small never reaches, and saying so is cheaper than someone wondering why the counter never moves. `--balance` and `--elect-schema-leader` stay off, each with its reason where an operator will look for it. The second reason is the one worth writing down: **with two voters it could not fire even if it were on.** Handing the row-key namespace on is a decision the agreement commits, a majority of two is two, and the silence it waits for is the same node being gone. It wants the third node that replication wants.
The obvious shape - one StatefulSet per node, because a node here is not interchangeable with the next one - was written first and thrown away. It is the correct argument and the wrong conclusion: what it produces is a deployment where growing the cluster means writing YAML, and the cluster has been able to admit a learner and give it work by itself since `--balance`. A knob that looks like scaling and produces an idle stranger is worse than no knob; a shape that has no knob at all is worse still. So: one StatefulSet, pod names as node names, and the four things a joining node needs split by who can actually supply them. **Only two of them are Kubernetes's.** A certificate is issued in advance - names are cheap, because a certificate grants nothing until the roster accepts it and the roster follows the agreement rather than a file. The cluster file is written per pod by an initContainer. The other two are the cluster's own: `add-node` registers the pod as a learner, and the balancer on the leader admits it and cuts the tail of the shard space over to it. **`add-node` runs before the daemon, and that ordering is the whole reason there is an initContainer rather than a sidecar.** A node that starts while the cluster has never heard of it is a voter in a cluster of its own imagining, and it stands for election once per timeout, raising the term on every node that hears it. Registered first, it is reached as a learner, and a learner does not stand. Three lists, three rules, and confusing them is the mistake this deployment makes easy, so they are in one file with the difference at the top. `seed.toml` is a membership and must be exact: a node with no log counts a majority out of it, so naming a pod that is not running is naming a vote that never arrives. `upstreams.toml` is an address space and deliberately names pods that do not exist - the proxy polls, finds them unreachable and leaves them out of rotation, which is what lets a scaled-up node reach a client with no proxy restart. The certificates are the third. Three tests pin all of it, including the file `join.sh` generates: a file the daemon refuses is a pod that crash-loops with no clue pointing back at a shell script inside a ConfigMap. Three nodes and not the demo's two, because two is the size at which nothing works: a majority of two is two, so neither a failover nor a schema-leader handover can be committed at the moment one node is the one that went away. `--balance` and `--elect-schema-leader` are on here for the same reason they are off in `cluster/` - there, the two ranges are the thing being explained; here the shape is meant to change. `certs.sh` and `users.sh` grow `NODES=` and `SECRETS=` overrides, so this deployment gets a CA and a users file of its own rather than a second copy of either script: two deployments sharing one CA are two deployments that can impersonate each other. Scaling in is not the mirror image and is not automated: drain, wait, remove, then scale. `cluster remove` refuses a node that still holds a range, and that refusal is the point.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.