Skip to content

Feat/join by address - #28

Merged
9bany merged 7 commits into
masterfrom
feat/join-by-address
Sep 5, 2026
Merged

9bany merged 7 commits into
masterfrom
feat/join-by-address

Conversation

@9bany

@9bany 9bany commented Sep 5, 2026

Copy link
Copy Markdown
Member

No description provided.

A schema leader that restarted assigned row ids from one past the highest
record anybody held, instead of from the block its predecessor had committed
through the agreement. Every id in between had been promised to a write that
may not have landed, so those ids were handed out twice - which nothing
downstream can see.

`Raft` marks an entry applied while its lock is held, and `Controller::run`
folds the decision it carries into the shared range map after releasing it. In
that window a node has applied everything and is still reading the map it
started from, whose reservations are empty - so the ceiling read zero and the
allocation fell back to the highest record.

Two signals, and the first alone was not enough:

- `Raft::behind`, from the commit index the leader advertises. A follower
  clamps its own to the length of its log, so a node missing half the log
  records a commit at the end of what it has and looks caught up. Recorded at
  the top of the `Append` arm, beside the lease it renews, because the appends
  carrying it are rejected for log mismatch until the log is repaired.
- `Controller.published`, the index up to which applied decisions have reached
  the shared state. This is the one that closed it.

`guard_schema` refuses with `schema_catching_up`, a 503 that clears in one
round trip, when either says this node is behind.

Found by an existing test that had been treated as flaky. It was not: it failed
whenever the window was hit. Put the bug back and it fails 2 runs in 10; with
the fix, 20 runs and four `make check`s in a row are clean. Two unit tests
cover `behind` without depending on timing.
A node that joined needed a `cluster.toml` naming the whole cluster, copied to
its machine and kept current. It now needs one address:

  bigctl cluster join d 10.0.0.4:7654       # against a node already in it
  big serve /data/d.big 0.0.0.0:7654 \
    --join 10.0.0.1:7654 --cluster-id big-demo --node d

`GET /internal/join` answers the membership the agreement holds, so a node
admitted last week is in it and no file anywhere had to be edited when it was.
`ClusterConfig::joining` builds the same kind of configuration a file produces,
so everything downstream of startup is the path a clustered node always took.

Three things are still given rather than discovered, and none of them can be
asked for: the name, because the answer is a list this node has to find itself
in; the cluster id, because every peer request carries a stamp derived from it;
and the peer CA, because trusting a peer is what makes its answer worth
reading. `--cluster-id` is the rule `cluster_id` always had, made the whole of
the configuration rather than one line of it.

Being added comes before starting, and the order is enforced rather than
advised: a certificate naming somebody the cluster has never heard of is
refused at the handshake. Startup says so by name and names the command that
was skipped. `bigctl cluster join` does the adding and prints the starting
command so the two halves cannot be run the wrong way round - which needed the
cluster's name in `/cluster/topology`, where an operator can read it.

Two ways this could have failed silently, both refused at the route instead:

- A cluster with no `cluster_id` is identified by the shape of its file, so a
  node with no file would start cleanly and have every later request refused as
  a mismatch. `409 cluster_unnamed`.
- A cluster whose ranges have no copies commits no decisions, and admission is
  a decision - so the answer would send an operator to `add-node`, which is
  itself refused. `409 no_agreement`, naming what is missing rather than what
  is next.

Also fixes a solo node reporting `127.0.0.1:7654` in `/cluster/topology`
whatever address it was serving: it is the same code path, and the two could
not be separated. `Server::bind_with` now binds before it builds the
coordinator, so the report is the port actually bound rather than the one
asked for.

The file path is untouched and stays the documented way to run a fixed cluster.
`bigproxy` read its upstreams once at startup, so a node admitted afterwards
was one no request could reach - a front door in front of a cluster that grows
had to be restarted to see the growth. `--discover` reads `/cluster/topology`
from a node already in rotation, on the health interval, and adopts what it
finds.

Off by default, like everything else here that changes a shape by itself: a
deployment that grew an upstream nobody wrote down is one whose shape an
operator cannot predict from what they wrote. `--discover-credentials` names a
file holding one `user:password` line, mode 600, because reading the membership
demands `Operate`; a cluster with no users file needs nothing.

Two properties that had to hold, and each has a test:

- **A discovered node starts out of rotation** and earns its way in with
  `--health-pass` probes, where a node named on the command line starts in. The
  asymmetry is the argument for the optimistic start, inverted: at startup the
  alternative is answering 503 while every node is fine, but a node appearing
  mid-flight invents no outage by waiting - and the moment a cluster announces a
  node is the moment it is least likely to have caught up.
- **A node already known keeps its health record.** Rebuilding it every couple
  of seconds would restart that record just as often, so a node that had been
  failing would return to rotation on the next poll and stay there - discovery
  quietly switching the health check off.

`ranges` in the same answer is deliberately not read: which node holds which
shard is the agreement's business, and this process is not in it.

The pool's node list is behind a lock and held as `Arc`s, so a request already
in flight keeps working on its node after the list has moved on.
…d list

`bigctl` matched the answer to `/cluster/topology` against an exact list of
keys, so adding `cluster_id` to that route broke every client built before it:
the answer fell through to the flat renderer, which cannot render an array, and
the operator got the raw JSON and `members is a shape this client does not know
how to show`.

The daemon is allowed to grow a field. What identifies a shape therefore has to
be what the shape is *for* - here the two arrays it is about - and not the
complete set of keys it happened to have on the day the client was written. The
proxy's reader of `cluster.toml` already follows that rule, and says so:
unknown top-level keys are ignored rather than refused.
A range's replication was fixed when the cluster started and could not be
changed afterwards. `RangeMap::assign` is the only thing that sets a group and
had one caller - `split_range`, with a group of one - so a node could join, be
admitted, be given a new range and have one moved to it, and could never become
a second holder of a range somebody already serves. For a system whose whole
failure story is "another node holds the same data", going from one copy to two
on a running deployment had no answer at all.

  bigctl cluster add-replica <range> to <node>
  bigctl cluster drop-replica <range> from <node>

**Nothing is refused while it copies**, which is the one way this differs from a
move. A move ends by dropping its source, so the new holder must be provably
exact before the handover - that is what the cutover is for. Nothing is dropped
here, so an incomplete copy cannot lose a record; it can only be behind, and
being behind is a state the agreement already models and already refuses to
read from.

So the copy joins the group **and is marked behind in one decision**. Writes
reach it from that moment, so it stops falling further behind, and no read is
answered from it until a repair has proved it agrees. Two decisions would leave
a window in which the map names a holder believed current and missing records,
and a promotion there answers a smaller count with no symptom. The catch-up,
the digest and the clearing of the mark are `move`'s and `repair`'s own, unchanged.

A catch-up that fails leaves a copy in the group and behind. That is a resting
state rather than a half-finished change to unwind: it is what a node that
missed a write looks like, and `/repair` is what finishes it.

Two guards that were missing, both found by running the thing by hand:

- **A range is never handed to a node that does not vote.** `node_named`
  excluded only `Gone`, so `cluster split <at> to <learner>` produced a holder
  the agreement does not count towards a majority - the state the console
  describes as impossible. `split`, `move` and `add-replica` now take
  `voter_named` and name `admit` in the refusal. The balancer already ordered
  `Admit` first for this reason; the manual verbs never learned it.
- **The two-node rule, at runtime.** A cluster file with a replica and two nodes
  is refused because a majority of two is two. Without the same check here, this
  verb would be the way to build that shape after startup instead.

Neither verb runs by itself. How many copies a range should have is a policy
with teeth - a loop that adds them can fill a disk, one that removes them can
take a cluster below its majority - so the balancer is deliberately not given it.
`bigproxy` refused to start without an upstream, while a proxy whose upstreams
are all down keeps running and answers `503`. Those are the same state one entry
apart, and the inconsistency made "bring the front door up first" impossible for
no reason a running proxy could name. An empty start is allowed now - and still
refused when nothing could ever fill it, which is a process that will answer
`503` until somebody restarts it.

What fills it is `--admin-addr`, a second listener carrying three routes:

  POST   /admin/upstream?name=a&addr=10.0.0.1:7654
  DELETE /admin/upstream?name=a
  GET    /admin/upstream

**Loopback only, refused anywhere else, and the bind is the whole of the access
control.** This proxy authenticates nothing: it forwards the client's
`Authorization` byte for byte and has never seen a password. A route that can
add an upstream can point every client's request at a machine of the caller's
choosing, so on the public port it would be the one unauthenticated way to take
the deployment over. Rather than give this process a first credential to hold,
reaching the port has to mean already being on the machine. Every address the
name resolves to is checked, so a name with one loopback record among several
cannot open it on the others.

A node telling the proxy it exists would not do: nothing here could check the
claim. An operator's command carries the same trust `--upstream` does.

`--cluster-id` adds a second question - a seeded address is asked which cluster
it is in before it is adopted - which is not authentication and is not offered
as any. It is what stops a typo pointing this proxy at the wrong cluster.

A seeded upstream starts out of rotation like a discovered one, and seeding a
name that is already here moves it rather than adding a second entry.
  ./scripts/local-cluster up | status | join <name> | logs | down | clean

Three nodes rather than one, and not out of caution: a cluster whose ranges have
no copies runs no agreement, and admission is a decision - so a cluster of one
or two can never be joined, split or reshaped. Three with a copy is the smallest
that can change its own shape.

The proxy is given one seed and `--discover`, so the other two are found rather
than listed - which is the thing worth demonstrating. `join` asks whoever leads
the agreement, because membership is a decision and any other node refuses it,
and starts the new node with an address and no cluster file at all.

`clean` leaves cluster.toml alone: it is the one file here somebody edited on
purpose. Everything else lives under local/, which .gitignore already covers.
@9bany
9bany merged commit 766f744 into master Sep 5, 2026
20 checks passed
@9bany
9bany deleted the feat/join-by-address branch September 7, 2026 02:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant