Skip to content

[Epic] P3 — Backup, Restore & PITR (v0.7) #166

Description

@habibtalib

Epic — v0.7 production-hardening, phase P3: Backup, Restore & PITR

Goal: an operator can take scheduled backups and restore to an exact point in time, verified.

High

  • PITR is inert scaffolding — execute_restore, create_base_snapshot, WalArchiver, SnapshotCatalog have zero production callers; no base-snapshot task is scheduled; resolve_pitr always returns None. Reads as shipped but cannot be invoked. Replay also ignores PitrTarget.replay_lsn (no upper bound). storage/snapshot_executor.rs:62, storage/snapshot_writer.rs:167, storage/snapshot.rs:147, wal/archiver.rs:99. Every caller of each is a test module. resolve_pitr is None end-to-end because nothing populates SnapshotCatalog in production. Replay bounds only the LOW side — dry_run_restore (storage/snapshot_restore.rs:61) checks replay_lsn < base_snapshot.begin_lsn and never an upper bound. The filename parser is broken independently: parse_segment_filename (wal/archiver.rs:233) splits on three hyphen-separated parts, expecting wal-{first}-{last}.seg; no writer produces that name. Done in Make cluster consensus, crash recovery, and restore crash-safe #403. PITR runs from nodedb restore (ctl/restore/) with --target-time, --target-lsn, or --cluster --restore-point. Replay stops at the target (ctl/restore/segment.rs, cut_rule.rs). The WAL archiver and base-snapshot cycle run in production (control/pitr/cycle.rs). Segment names round-trip (nodedb-wal/src/segment/meta.rs).

Medium

  • WAL-archive upload failures only warn!ed and truncation proceeded, leaving a hole in the archived stream. A failed upload now keeps the segment on local disk and holds truncation back at the first unarchived segment (control/checkpoint_archival.rs, failed_first_lsns), and files the wal_archival_failed_truncation_held diagnostic. 1744dec.
  • Hand-rolled UTC/leap-year math for PITR targets. parse_utc_timestamp now parses with chrono (DateTime::parse_from_rfc3339, then NaiveDateTime for the space and T forms) in storage/snapshot_restore.rs. 98dd82e.
  • Restore staleness gate bypassed on lock poisoning → stale envelope overwrites newer writes. Both directions failed open: the reader took .lock().ok()...unwrap_or(0), and advance_tenant_write_hlc skipped the write entirely under if let Ok. A poisoned mutex therefore stopped recording writes AND reported a high-water mark of 0, which every watermark clears. Both now recover the lock (unwrap_or_else(|p| p.into_inner())) — the map is a plain HashMap a panic elsewhere cannot corrupt — and the read is behind a single SharedState::tenant_write_hlc accessor. 36296e0.
  • No restore verification (post-restore row-count/checksum reconciliation). RestoreStats counts only what the envelope itself carried (orchestrate/restore.rs:96-105), never a comparison against the source. Done in Make cluster consensus, crash recovery, and restore crash-safe #403. RESTORE DATABASE checks per-collection row counts and digests against the backup (control/backup/verify/). nodedb restore checks segment CRCs only.
  • Backup is full-snapshot only — no incremental/differential, no scheduling/retention. control/backup/ exposes one entry point, backup_tenant (orchestrator.rs:37). Done in Make cluster consensus, crash recovery, and restore crash-safe #403. Scheduled backups with keep retention (control/backup/schedule/), and PITR base-snapshot retention. PITR base snapshots are incremental: only changed chunks upload. BACKUP DATABASE stays a full logical copy.

Success: scheduled base snapshots + WAL archiver wired; restore replays to a target LSN/timestamp; end-to-end "base + WAL → restore to T, assert state == state@T" test passes.

Re-verified against main @ 1ff3551 (2026-09-28): the archive-upload and UTC-parser items are fixed and checked. PITR wiring, the replay upper bound, restore verification, and incremental backups remain open.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:storage-walWAL, durability, flush/replaytype:epicTracking issue spanning multiple sub-issues

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions