Skip to content

Make cluster consensus, crash recovery, and restore crash-safe - #403

Merged
farhan-syah merged 12 commits into
mainfrom
fix/cluster-recovery
Oct 3, 2026
Merged

farhan-syah merged 12 commits into
mainfrom
fix/cluster-recovery

Conversation

@farhan-syah

Copy link
Copy Markdown
Member

A restart or partition no longer loses acknowledged state, and a restore no longer drops committed rows. Backups cut consistently per database, and point-in-time restore runs from restore points and WAL time anchors.

What changed

Area Change Paths
Raft Queue each committed index once per batch. Serve linearizable reads from a leader lease. Refuse competing votes while the lease holds. Track per-peer contact for check-quorum. nodedb-raft/src/node/leader_lease.rs, nodedb-raft/src/node/quorum_contact.rs
Raft persistence Persist votes, snapshot installs, and membership changes before acknowledging them. Make group mount and unmount explicit steps. nodedb-cluster/src/raft_loop/, nodedb-cluster/src/group_disk/
Routing Follow elections with leader hints, leader balancing, and routing persistence. Name a hosted group in redirects. nodedb-cluster/src/raft_loop/
Calvin sequencer Collect, complete, and garbage collect multi-part transactions. Bound the entry size. nodedb-cluster/src/calvin/, nodedb-cluster/src/calvin/sequencer/
Calvin snapshots Capture sequencer state in snapshots. Restore it on a joiner, including after log compaction. Gate schedulers on snapshot install and metadata catch-up. nodedb-cluster/src/calvin/sequencer/state_machine/, nodedb/src/control/cluster/calvin_snapshot/, nodedb/src/control/state/calvin_bases.rs
Compaction Floor sequencer log compaction at the lowest voter match index. nodedb-cluster/src/multi_raft/proposals.rs, nodedb-raft/src/node/peer_contact.rs
Metadata replay Stamp every delete with the modification_hlc of the row it targeted. Refuse the delete once the row has moved past that clock. Restart replay from a durable floor. nodedb/src/control/catalog_entry/incarnation/, nodedb/src/control/cluster/metadata_applier/
Crash replay Replay redo images, write-set journals, and hash chains after a crash. Stage snapshot install so an interrupted install recovers. nodedb/src/wal/redo/, nodedb/src/control/cluster/snapshot_install/
CDC Carry transaction boundaries in change events. Replay events from the WAL. Resume consumers across leader changes. Give each sink one owner. nodedb/src/event/cdc/, nodedb/src/control/change_stream/
Triggers Fire triggers in a dedicated lane that survives failover. nodedb/src/event/trigger/lane/
Placement Reconcile placement against live peers. Reclaim descriptor leases from dead holders. Route graph, array, and timeseries reads to the owning home. nodedb/src/control/lease/, nodedb/src/control/cluster/
Backup Cut backups consistently per database. Schedule them. Check them. Store them locally or remotely. nodedb/src/control/backup/cut_capture/, nodedb/src/control/backup/schedule/, nodedb/src/control/backup/verify/
Restore Reissue arrays and graph edges. Protect the surrogate floor. Resolve bind conflicts. Retry safely. nodedb/src/control/backup/restore/, nodedb/src/ctl/restore/
PITR Restore to a point in time from restore points and WAL time anchors. Replay stops at the target. Archive WAL segments. Base snapshots upload only changed chunks. nodedb/src/control/pitr/, nodedb/src/wal/archiver/, nodedb-wal/src/record/restore_point.rs
Docs and CI Describe recovery, backup, and PITR. Adjust the dispatch and reconstructed-SQL checks. scripts/ci/check_authorized_dispatch.py, scripts/ci/check_reconstructed_sql.py
Tests Cover each area in cluster, in-process, native, and wire suites. nodedb-cluster-tests/, nodedb/tests/

Why

Symptom Cause
A replayed delete removed a recreated object The delete did not name the incarnation it targeted.
A joiner lost Calvin inputs after log compaction Snapshots did not carry sequencer state.
A crash during snapshot install left a half-installed group Install was not staged.
Restore missed rows and lost graph edges and arrays Restore did not reissue them or protect the surrogate floor.

Linked issues

Issue Status
#165 Closes. The rolling-upgrade wire floor stays pinned until 1.0, as wire_version.rs documents.
#166 Closes

Closes #165
Closes #166

Known limits

  • Autocommit writes with no Calvin lock representation skip the lock table. These are batch, INSERT ... SELECT, upsert, CRDT, columnar, timeseries, spatial, array, and cross-home edge writes. A vShard with no running scheduler also skips it. Point writes and single-shard predicate writes take the locks.
  • A single-replica group that loses its Calvin base stops Calvin for that vShard. The node logs one error for the loss. Only a snapshot restores the base, and a sole replica has no peer to supply one.
  • A voter that stops responding holds back sequencer log compaction. The floor is the lowest voter match index.
  • Point-in-time restore runs offline through nodedb restore. SQL RESTORE DATABASE restores a logical backup and has no time target.
  • RESTORE DATABASE checks row counts and digests against the backup. nodedb restore checks segment CRCs only, with no row tally.
  • BACKUP DATABASE writes a full logical envelope each time. Incremental storage comes from PITR base snapshots.
  • A committed index can reach the applier twice across Ready batches. The applier drops any index at or below last_applied.

How to check it

Command Result
cargo nextest run --workspace --exclude nodedb-cluster-tests --all-features --no-fail-fast Must pass
cargo nextest run -p nodedb-cluster-tests --all-features --no-fail-fast Must pass
cargo clippy --workspace --all-targets --all-features -- -D warnings Must be clean

@farhan-syah farhan-syah added the run-ci Opt this PR into the full test suite; re-add to force a re-run label Oct 1, 2026
Comment thread nodedb/src/control/server/dispatch_utils/dispatch.rs Fixed
@farhan-syah
farhan-syah force-pushed the fix/cluster-recovery branch from 8da09ef to cde126f Compare October 1, 2026 11:40
The metadata log replays from its start on every boot. A delete proposed
against one incarnation of a name must not remove a later incarnation
of that name after a create-drop-recreate sequence. Stamp every delete
with the modification_hlc (and descriptor version, where applicable) of
the row it targeted at propose time, and refuse to apply it once the
row has moved past that clock.

Carries the fence through purge, sequence, trigger, function, procedure,
materialized view, continuous aggregate, synonym group, topic, and
vector index params deletes, across apply, post-apply dispatch, DDL
compensation/reversal, and cluster replication. Adds a committed-only
topic lookup and a compensation module that reverses finalized DDL
under the same fencing rule when buffered DML fails to dispatch.
Raft and consensus
- Leaders serve linearizable reads from a lease, and refuse competing
  votes while the lease holds; check-quorum tracks per-peer contact.
- Voting, snapshot install, and membership changes persist before they
  are acknowledged, and group mount and unmount are explicit steps.
- Leader hints, leader balancing, and routing persistence follow
  elections, and redirects name a hosted group.

Calvin sequencing and snapshots
- Multi-part transactions are collected, completed, and garbage
  collected by the sequencer, with a bounded entry size.
- Sequencer state is captured in snapshots and restored on a joiner,
  including after the log has been compacted.
- Schedulers gate on snapshot install and metadata catch-up.

Crash recovery and replay
- Metadata replay fences deletes to the incarnation they targeted and
  restarts from a durable floor.
- Redo images, write-set journals, and hash chains replay after a
  crash, and snapshot install is staged so an interrupted install
  recovers.

CDC and the trigger lane
- Change events carry transaction boundaries and replay from the WAL.
- Consumers resume across leader changes, and sinks have a single owner.
- Trigger firing runs in a dedicated lane that survives failover.

Placement and routing
- Placement reconciliation, leader preference, and membership-aware
  routing track live peers; descriptor leases reclaim from dead holders.
- Graph, array, and timeseries reads route to the owning home.

Backup, restore, and PITR
- Backups are cut consistently per database, scheduled, verified, and
  stored locally or remotely.
- Restore reissues arrays and graph edges, protects the surrogate
  floor, resolves bind conflicts, and retries safely.
- Point-in-time restore uses restore points and WAL time anchors.

Tests
- Cluster, in-process, native, and wire suites cover each area above,
  with matching harness support and nextest config.
Update the architecture, query language, real-time, and security docs
and the changelog, and adjust the CI dispatch and reconstructed-SQL
checks to match the code changes.
An empty data directory previously resolved the incarnation file
against the process working directory, minting an incarnation in an
unrelated location. Return a storage error instead.
The presence check and the old-expiry read now share a single entry
metadata lookup.
retry_through_drain re-runs an operation refused with a retryable
schema change until it succeeds, fails terminally, or a caller-supplied
budget elapses. Each wait ends early when a drain ends on this node.
An ILP batch now projects its inferred time column apart from the tag
and field columns. The catalog merge drops it for a timeseries
collection whose declared time key names one of its fields, so a
steady ingest stream no longer proposes a new descriptor version on
every flush.

The batch flush also takes its write leases through retry_through_drain,
so a flush that meets a descriptor drain waits it out instead of
dropping the connection.
@farhan-syah
farhan-syah force-pushed the fix/cluster-recovery branch from 9fc96f8 to 8583388 Compare October 1, 2026 18:08
Removing a vshard's sender now drops its armed catch-up, and a new
scheduler registering a sender starts with none. The compaction floor
counts only vshards with a registered sender, so an arm made by an
exiting scheduler can no longer hold sequencer log compaction down.
…aker

A receiver whose handler fails now answers with a typed refusal frame
instead of dropping the stream. The sender treats any answer, a
refusal included, as proof the link is up: it closes the peer's circuit
breaker and does not resend the request. Only a failed connect,
handshake, stream open, write, read timeout or lost connection counts
as a failure.

The breaker gains explicit admission, so an open circuit lets a single
recovery probe through. Health pings use that probe path, and a ping
refused by this node's own open circuit is no longer recorded as a ping
failure. Per-peer Raft batches end on a link failure but move past a
refusal, so one group a peer does not host yet no longer holds back
another group's heartbeat.
Replacing a vshard's output sender no longer drops the catch-up armed
for it. Only a vshard that leaves this node drops its catch-up, so the
sequencer compaction floor keeps covering the replay a new scheduler
still needs.
Raft: a successful AppendEntries response now reports the last entry
the follower shares with the leader and holds durably, and the leader
uses it as the follower's match index. Quorum contact tracking and the
leader lease are reworked around this, and the contested-election
marker is removed.

Transport: concurrent sends to one peer share a single dial, and a
connection is evicted only when it failed.

Control plane: the auth lease, leased reads, the calvin scheduler,
gateway routing, and graph dispatch read leaders through a new
live-leaders snapshot that takes the Raft status before the routing
guard. The lock order between MultiRaft and the routing table is
documented to prevent a nested-read deadlock.
@farhan-syah
farhan-syah merged commit 2ce633a into main Oct 3, 2026
12 of 13 checks passed
@farhan-syah
farhan-syah deleted the fix/cluster-recovery branch October 3, 2026 05:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-ci Opt this PR into the full test suite; re-add to force a re-run

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Epic] P3 — Backup, Restore & PITR (v0.7) [Epic] P2 — Cluster Consensus Safety (v0.6)

2 participants