You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
HlcClock::update is unwired — no cross-node HLC merge and no clock-skew bound anywhere #264
HlcClock never folds in a remote timestamp, so it is not hybrid: every node's clock is an isolated monotone wall clock with no skew bound anywhere in the workspace.
What is wired
HlcClock::update (nodedb-types/src/hlc.rs:108) is the only path that folds a remote observation into the local clock. Its production callers:
Folds the prior catalog entry's modification_hlc for the same descriptor, only when prior_hlc >= hlc, only during DDL stamping
That is the complete list. The remaining update calls are in hlc.rs unit tests and one test in control/lease/drain_propose.rs:575.
No Raft path, cluster RPC path, replication path, or lease path calls it. Inbound messages carry HLCs that are read and never merged.
Consequences
No cross-node frame.HlcClock::now() returns wall.max(st.wall_ns) (hlc.rs:96) — monotone against this node's own history and nothing else. Two nodes' wall_ns values are independent local clocks.
Every cross-node HLC comparison is a raw wall-clock comparison. Descriptor-lease expires_at is stamped by the holder via hlc_clock.now() (control/lease/propose.rs:192) and judged by a different node against its own clock (control/lease/drain_propose.rs:256). [Epic] P2 — Cluster Consensus Safety (v0.6) #165 records this under "raw cross-node wall clock, no skew bound"; the cause is that update is unwired, not that the lease path chose wall time.
Happens-before is not established. An HLC that a node never observes cannot order that node's later events after it. Any consumer treating these values as causally ordered across nodes is reading a bare timestamp.
Backup watermarks inherit it.control/backup/orchestrator.rs:75 takes hlc_clock.now().wall_ns as snapshot_watermark, with the comment that it "advances past any previously observed HLC" — true only for HLCs this node observed, which is one descriptor-stamp path.
No skew bound exists
grep -rn "MAX_SKEW\|MAX_DRIFT\|max_offset\|MAX_CLOCK_OFFSET\|clock_offset" --include='*.rs' across the workspace returns nothing. There is no max-offset rejection and no drift alarm.
Both directions are unguarded, and wiring update without a bound trades one defect for another:
Unwired today: a peer's clock skew is invisible, and cross-node deadlines are wrong by the full offset.
Wired without a bound: update takes an unconditional max (hlc.rs:112), so one peer sending a far-future HLC drags the receiver's clock forward permanently, and every node it then talks to. Nothing brings it back — now() clamps to st.wall_ns forever after.
What lands
Fold update in at every inbound cluster boundary that carries an HLC — Raft RPC decode, metadata apply, lease grant.
Reject or alarm on a remote HLC more than a bounded offset ahead of local wall time, at that same boundary. The bound belongs at ingestion, once, not at each comparison site.
Derive any downstream skew allowance from that bound rather than declaring an independent constant. See PR fix(lease): clamp lease expiry against MAX_SKEW clock skew #250 for what an independent one costs: a 300 s allowance against a 35 s DEFAULT_DRAIN_TIMEOUT turns a crashed holder's lease into five minutes of failing DDL.
Relation to existing work
#165 tracks the lease-level symptom: "Descriptor-lease expiry uses raw cross-node wall clock, no skew bound / fencing token". This issue is the cause underneath it, and it reaches past leases into backup watermarks and every cross-node HLC comparison. The clamp #165 lists as remaining work needs the bound established here to derive from.
HlcClocknever folds in a remote timestamp, so it is not hybrid: every node's clock is an isolated monotone wall clock with no skew bound anywhere in the workspace.What is wired
HlcClock::update(nodedb-types/src/hlc.rs:108) is the only path that folds a remote observation into the local clock. Its production callers:nodedb/src/control/catalog_entry/descriptor_stamp.rs:65modification_hlcfor the same descriptor, only whenprior_hlc >= hlc, only during DDL stampingThat is the complete list. The remaining
updatecalls are inhlc.rsunit tests and one test incontrol/lease/drain_propose.rs:575.No Raft path, cluster RPC path, replication path, or lease path calls it. Inbound messages carry HLCs that are read and never merged.
Consequences
HlcClock::now()returnswall.max(st.wall_ns)(hlc.rs:96) — monotone against this node's own history and nothing else. Two nodes'wall_nsvalues are independent local clocks.expires_atis stamped by the holder viahlc_clock.now()(control/lease/propose.rs:192) and judged by a different node against its own clock (control/lease/drain_propose.rs:256). [Epic] P2 — Cluster Consensus Safety (v0.6) #165 records this under "raw cross-node wall clock, no skew bound"; the cause is thatupdateis unwired, not that the lease path chose wall time.control/backup/orchestrator.rs:75takeshlc_clock.now().wall_nsassnapshot_watermark, with the comment that it "advances past any previously observed HLC" — true only for HLCs this node observed, which is one descriptor-stamp path.No skew bound exists
grep -rn "MAX_SKEW\|MAX_DRIFT\|max_offset\|MAX_CLOCK_OFFSET\|clock_offset" --include='*.rs'across the workspace returns nothing. There is no max-offset rejection and no drift alarm.Both directions are unguarded, and wiring
updatewithout a bound trades one defect for another:updatetakes an unconditionalmax(hlc.rs:112), so one peer sending a far-future HLC drags the receiver's clock forward permanently, and every node it then talks to. Nothing brings it back —now()clamps tost.wall_nsforever after.What lands
updatein at every inbound cluster boundary that carries an HLC — Raft RPC decode, metadata apply, lease grant.DEFAULT_DRAIN_TIMEOUTturns a crashed holder's lease into five minutes of failing DDL.Relation to existing work
#165 tracks the lease-level symptom: "Descriptor-lease expiry uses raw cross-node wall clock, no skew bound / fencing token". This issue is the cause underneath it, and it reaches past leases into backup watermarks and every cross-node HLC comparison. The clamp #165 lists as remaining work needs the bound established here to derive from.