Repository navigation
Feat/join by address - #28
Merged
Merged
Conversation
A schema leader that restarted assigned row ids from one past the highest record anybody held, instead of from the block its predecessor had committed through the agreement. Every id in between had been promised to a write that may not have landed, so those ids were handed out twice - which nothing downstream can see. `Raft` marks an entry applied while its lock is held, and `Controller::run` folds the decision it carries into the shared range map after releasing it. In that window a node has applied everything and is still reading the map it started from, whose reservations are empty - so the ceiling read zero and the allocation fell back to the highest record. Two signals, and the first alone was not enough: - `Raft::behind`, from the commit index the leader advertises. A follower clamps its own to the length of its log, so a node missing half the log records a commit at the end of what it has and looks caught up. Recorded at the top of the `Append` arm, beside the lease it renews, because the appends carrying it are rejected for log mismatch until the log is repaired. - `Controller.published`, the index up to which applied decisions have reached the shared state. This is the one that closed it. `guard_schema` refuses with `schema_catching_up`, a 503 that clears in one round trip, when either says this node is behind. Found by an existing test that had been treated as flaky. It was not: it failed whenever the window was hit. Put the bug back and it fails 2 runs in 10; with the fix, 20 runs and four `make check`s in a row are clean. Two unit tests cover `behind` without depending on timing.
A node that joined needed a `cluster.toml` naming the whole cluster, copied to
its machine and kept current. It now needs one address:
bigctl cluster join d 10.0.0.4:7654 # against a node already in it
big serve /data/d.big 0.0.0.0:7654 \
--join 10.0.0.1:7654 --cluster-id big-demo --node d
`GET /internal/join` answers the membership the agreement holds, so a node
admitted last week is in it and no file anywhere had to be edited when it was.
`ClusterConfig::joining` builds the same kind of configuration a file produces,
so everything downstream of startup is the path a clustered node always took.
Three things are still given rather than discovered, and none of them can be
asked for: the name, because the answer is a list this node has to find itself
in; the cluster id, because every peer request carries a stamp derived from it;
and the peer CA, because trusting a peer is what makes its answer worth
reading. `--cluster-id` is the rule `cluster_id` always had, made the whole of
the configuration rather than one line of it.
Being added comes before starting, and the order is enforced rather than
advised: a certificate naming somebody the cluster has never heard of is
refused at the handshake. Startup says so by name and names the command that
was skipped. `bigctl cluster join` does the adding and prints the starting
command so the two halves cannot be run the wrong way round - which needed the
cluster's name in `/cluster/topology`, where an operator can read it.
Two ways this could have failed silently, both refused at the route instead:
- A cluster with no `cluster_id` is identified by the shape of its file, so a
node with no file would start cleanly and have every later request refused as
a mismatch. `409 cluster_unnamed`.
- A cluster whose ranges have no copies commits no decisions, and admission is
a decision - so the answer would send an operator to `add-node`, which is
itself refused. `409 no_agreement`, naming what is missing rather than what
is next.
Also fixes a solo node reporting `127.0.0.1:7654` in `/cluster/topology`
whatever address it was serving: it is the same code path, and the two could
not be separated. `Server::bind_with` now binds before it builds the
coordinator, so the report is the port actually bound rather than the one
asked for.
The file path is untouched and stays the documented way to run a fixed cluster.
`bigproxy` read its upstreams once at startup, so a node admitted afterwards was one no request could reach - a front door in front of a cluster that grows had to be restarted to see the growth. `--discover` reads `/cluster/topology` from a node already in rotation, on the health interval, and adopts what it finds. Off by default, like everything else here that changes a shape by itself: a deployment that grew an upstream nobody wrote down is one whose shape an operator cannot predict from what they wrote. `--discover-credentials` names a file holding one `user:password` line, mode 600, because reading the membership demands `Operate`; a cluster with no users file needs nothing. Two properties that had to hold, and each has a test: - **A discovered node starts out of rotation** and earns its way in with `--health-pass` probes, where a node named on the command line starts in. The asymmetry is the argument for the optimistic start, inverted: at startup the alternative is answering 503 while every node is fine, but a node appearing mid-flight invents no outage by waiting - and the moment a cluster announces a node is the moment it is least likely to have caught up. - **A node already known keeps its health record.** Rebuilding it every couple of seconds would restart that record just as often, so a node that had been failing would return to rotation on the next poll and stay there - discovery quietly switching the health check off. `ranges` in the same answer is deliberately not read: which node holds which shard is the agreement's business, and this process is not in it. The pool's node list is behind a lock and held as `Arc`s, so a request already in flight keeps working on its node after the list has moved on.
…d list `bigctl` matched the answer to `/cluster/topology` against an exact list of keys, so adding `cluster_id` to that route broke every client built before it: the answer fell through to the flat renderer, which cannot render an array, and the operator got the raw JSON and `members is a shape this client does not know how to show`. The daemon is allowed to grow a field. What identifies a shape therefore has to be what the shape is *for* - here the two arrays it is about - and not the complete set of keys it happened to have on the day the client was written. The proxy's reader of `cluster.toml` already follows that rule, and says so: unknown top-level keys are ignored rather than refused.
A range's replication was fixed when the cluster started and could not be changed afterwards. `RangeMap::assign` is the only thing that sets a group and had one caller - `split_range`, with a group of one - so a node could join, be admitted, be given a new range and have one moved to it, and could never become a second holder of a range somebody already serves. For a system whose whole failure story is "another node holds the same data", going from one copy to two on a running deployment had no answer at all. bigctl cluster add-replica <range> to <node> bigctl cluster drop-replica <range> from <node> **Nothing is refused while it copies**, which is the one way this differs from a move. A move ends by dropping its source, so the new holder must be provably exact before the handover - that is what the cutover is for. Nothing is dropped here, so an incomplete copy cannot lose a record; it can only be behind, and being behind is a state the agreement already models and already refuses to read from. So the copy joins the group **and is marked behind in one decision**. Writes reach it from that moment, so it stops falling further behind, and no read is answered from it until a repair has proved it agrees. Two decisions would leave a window in which the map names a holder believed current and missing records, and a promotion there answers a smaller count with no symptom. The catch-up, the digest and the clearing of the mark are `move`'s and `repair`'s own, unchanged. A catch-up that fails leaves a copy in the group and behind. That is a resting state rather than a half-finished change to unwind: it is what a node that missed a write looks like, and `/repair` is what finishes it. Two guards that were missing, both found by running the thing by hand: - **A range is never handed to a node that does not vote.** `node_named` excluded only `Gone`, so `cluster split <at> to <learner>` produced a holder the agreement does not count towards a majority - the state the console describes as impossible. `split`, `move` and `add-replica` now take `voter_named` and name `admit` in the refusal. The balancer already ordered `Admit` first for this reason; the manual verbs never learned it. - **The two-node rule, at runtime.** A cluster file with a replica and two nodes is refused because a majority of two is two. Without the same check here, this verb would be the way to build that shape after startup instead. Neither verb runs by itself. How many copies a range should have is a policy with teeth - a loop that adds them can fill a disk, one that removes them can take a cluster below its majority - so the balancer is deliberately not given it.
`bigproxy` refused to start without an upstream, while a proxy whose upstreams are all down keeps running and answers `503`. Those are the same state one entry apart, and the inconsistency made "bring the front door up first" impossible for no reason a running proxy could name. An empty start is allowed now - and still refused when nothing could ever fill it, which is a process that will answer `503` until somebody restarts it. What fills it is `--admin-addr`, a second listener carrying three routes: POST /admin/upstream?name=a&addr=10.0.0.1:7654 DELETE /admin/upstream?name=a GET /admin/upstream **Loopback only, refused anywhere else, and the bind is the whole of the access control.** This proxy authenticates nothing: it forwards the client's `Authorization` byte for byte and has never seen a password. A route that can add an upstream can point every client's request at a machine of the caller's choosing, so on the public port it would be the one unauthenticated way to take the deployment over. Rather than give this process a first credential to hold, reaching the port has to mean already being on the machine. Every address the name resolves to is checked, so a name with one loopback record among several cannot open it on the others. A node telling the proxy it exists would not do: nothing here could check the claim. An operator's command carries the same trust `--upstream` does. `--cluster-id` adds a second question - a seeded address is asked which cluster it is in before it is adopted - which is not authentication and is not offered as any. It is what stops a typo pointing this proxy at the wrong cluster. A seeded upstream starts out of rotation like a discovered one, and seeding a name that is already here moves it rather than adding a second entry.
./scripts/local-cluster up | status | join <name> | logs | down | clean Three nodes rather than one, and not out of caution: a cluster whose ranges have no copies runs no agreement, and admission is a decision - so a cluster of one or two can never be joined, split or reshaped. Three with a copy is the smallest that can change its own shape. The proxy is given one seed and `--discover`, so the other two are found rather than listed - which is the thing worth demonstrating. `join` asks whoever leads the agreement, because membership is a decision and any other node refuses it, and starts the new node with an address and no cluster file at all. `clean` leaves cluster.toml alone: it is the one file here somebody edited on purpose. Everything else lives under local/, which .gitignore already covers.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.