You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Goal: an operator can take scheduled backups and restore to an exact point in time, verified.
High
PITR is inert scaffolding — execute_restore, create_base_snapshot, WalArchiver, SnapshotCatalog have zero production callers; no base-snapshot task is scheduled; resolve_pitr always returns None. Reads as shipped but cannot be invoked. Replay also ignores PitrTarget.replay_lsn (no upper bound). storage/snapshot_executor.rs:62, storage/snapshot_writer.rs:167, storage/snapshot.rs:147, wal/archiver.rs:99. Every caller of each is a test module. resolve_pitr is None end-to-end because nothing populates SnapshotCatalog in production. Replay bounds only the LOW side — dry_run_restore (storage/snapshot_restore.rs:61) checks replay_lsn < base_snapshot.begin_lsn and never an upper bound. The filename parser is broken independently: parse_segment_filename (wal/archiver.rs:233) splits on three hyphen-separated parts, expecting wal-{first}-{last}.seg; no writer produces that name. Done in Make cluster consensus, crash recovery, and restore crash-safe #403. PITR runs from nodedb restore (ctl/restore/) with --target-time, --target-lsn, or --cluster --restore-point. Replay stops at the target (ctl/restore/segment.rs, cut_rule.rs). The WAL archiver and base-snapshot cycle run in production (control/pitr/cycle.rs). Segment names round-trip (nodedb-wal/src/segment/meta.rs).
Medium
WAL-archive upload failures only warn!ed and truncation proceeded, leaving a hole in the archived stream. A failed upload now keeps the segment on local disk and holds truncation back at the first unarchived segment (control/checkpoint_archival.rs, failed_first_lsns), and files the wal_archival_failed_truncation_held diagnostic. 1744dec.
Hand-rolled UTC/leap-year math for PITR targets. parse_utc_timestamp now parses with chrono (DateTime::parse_from_rfc3339, then NaiveDateTime for the space and T forms) in storage/snapshot_restore.rs. 98dd82e.
Restore staleness gate bypassed on lock poisoning → stale envelope overwrites newer writes. Both directions failed open: the reader took .lock().ok()...unwrap_or(0), and advance_tenant_write_hlc skipped the write entirely under if let Ok. A poisoned mutex therefore stopped recording writes AND reported a high-water mark of 0, which every watermark clears. Both now recover the lock (unwrap_or_else(|p| p.into_inner())) — the map is a plain HashMap a panic elsewhere cannot corrupt — and the read is behind a single SharedState::tenant_write_hlc accessor. 36296e0.
No restore verification (post-restore row-count/checksum reconciliation). RestoreStats counts only what the envelope itself carried (orchestrate/restore.rs:96-105), never a comparison against the source. Done in Make cluster consensus, crash recovery, and restore crash-safe #403.RESTORE DATABASE checks per-collection row counts and digests against the backup (control/backup/verify/). nodedb restore checks segment CRCs only.
Backup is full-snapshot only — no incremental/differential, no scheduling/retention. control/backup/ exposes one entry point, backup_tenant (orchestrator.rs:37). Done in Make cluster consensus, crash recovery, and restore crash-safe #403. Scheduled backups with keep retention (control/backup/schedule/), and PITR base-snapshot retention. PITR base snapshots are incremental: only changed chunks upload. BACKUP DATABASE stays a full logical copy.
Success: scheduled base snapshots + WAL archiver wired; restore replays to a target LSN/timestamp; end-to-end "base + WAL → restore to T, assert state == state@T" test passes.
Re-verified against main @ 1ff3551 (2026-09-28): the archive-upload and UTC-parser items are fixed and checked. PITR wiring, the replay upper bound, restore verification, and incremental backups remain open.
Epic — v0.7 production-hardening, phase P3: Backup, Restore & PITR
Goal: an operator can take scheduled backups and restore to an exact point in time, verified.
High
execute_restore,create_base_snapshot,WalArchiver,SnapshotCataloghave zero production callers; no base-snapshot task is scheduled;resolve_pitralways returnsNone. Reads as shipped but cannot be invoked. Replay also ignoresPitrTarget.replay_lsn(no upper bound).storage/snapshot_executor.rs:62,storage/snapshot_writer.rs:167,storage/snapshot.rs:147,wal/archiver.rs:99. Every caller of each is a test module.resolve_pitrisNoneend-to-end because nothing populatesSnapshotCatalogin production. Replay bounds only the LOW side —dry_run_restore(storage/snapshot_restore.rs:61) checksreplay_lsn < base_snapshot.begin_lsnand never an upper bound. The filename parser is broken independently:parse_segment_filename(wal/archiver.rs:233) splits on three hyphen-separated parts, expectingwal-{first}-{last}.seg; no writer produces that name. Done in Make cluster consensus, crash recovery, and restore crash-safe #403. PITR runs fromnodedb restore(ctl/restore/) with--target-time,--target-lsn, or--cluster --restore-point. Replay stops at the target (ctl/restore/segment.rs,cut_rule.rs). The WAL archiver and base-snapshot cycle run in production (control/pitr/cycle.rs). Segment names round-trip (nodedb-wal/src/segment/meta.rs).Medium
warn!ed and truncation proceeded, leaving a hole in the archived stream. A failed upload now keeps the segment on local disk and holds truncation back at the first unarchived segment (control/checkpoint_archival.rs,failed_first_lsns), and files thewal_archival_failed_truncation_helddiagnostic. 1744dec.parse_utc_timestampnow parses with chrono (DateTime::parse_from_rfc3339, thenNaiveDateTimefor the space andTforms) instorage/snapshot_restore.rs. 98dd82e..lock().ok()...unwrap_or(0), andadvance_tenant_write_hlcskipped the write entirely underif let Ok. A poisoned mutex therefore stopped recording writes AND reported a high-water mark of0, which every watermark clears. Both now recover the lock (unwrap_or_else(|p| p.into_inner())) — the map is a plainHashMapa panic elsewhere cannot corrupt — and the read is behind a singleSharedState::tenant_write_hlcaccessor. 36296e0.RestoreStatscounts only what the envelope itself carried (orchestrate/restore.rs:96-105), never a comparison against the source. Done in Make cluster consensus, crash recovery, and restore crash-safe #403.RESTORE DATABASEchecks per-collection row counts and digests against the backup (control/backup/verify/).nodedb restorechecks segment CRCs only.control/backup/exposes one entry point,backup_tenant(orchestrator.rs:37). Done in Make cluster consensus, crash recovery, and restore crash-safe #403. Scheduled backups withkeepretention (control/backup/schedule/), and PITR base-snapshot retention. PITR base snapshots are incremental: only changed chunks upload.BACKUP DATABASEstays a full logical copy.Success: scheduled base snapshots + WAL archiver wired; restore replays to a target LSN/timestamp; end-to-end "base + WAL → restore to T, assert state == state@T" test passes.
Re-verified against
main@ 1ff3551 (2026-09-28): the archive-upload and UTC-parser items are fixed and checked. PITR wiring, the replay upper bound, restore verification, and incremental backups remain open.