Skip to content

feat(metrics): count graph edge writes and deletes - #409

Closed
EnRaiha wants to merge 2 commits into
NodeDB-Lab:mainfrom
EnRaiha:pr/374-372-edge-write-counters
Closed

EnRaiha wants to merge 2 commits into
NodeDB-Lab:mainfrom
EnRaiha:pr/374-372-edge-write-counters

Conversation

@EnRaiha

@EnRaiha EnRaiha commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

What

The graph engine exposed a gauge of live edges and nothing else.

An edge put that rewrote a live edge was therefore invisible: the gauge stayed flat and no counter moved. An operator could not tell a hot rewrite path from an idle one, and an edge delete that removed nothing looked the same as one that removed an edge. A loader pacing against graph writes had no arrival-versus-apply signal and had to estimate from query counts.

The counters exist so that a loader can pace on arrival versus apply rate instead of estimating it, and so that an operator sees a hot rewrite path. SystemMetrics gains graph_edges_written_total and graph_edges_deleted_total, recorded from the four edge write handlers and rendered on /metrics and as the graph_edges_written_total / graph_edges_deleted_total rows of SHOW STATS.

nodedb_graph_edges_written_total counts applied edge versions, so a put that rewrites a live edge increments it while the nodedb_graph_edges gauge stays flat. nodedb_graph_edges_deleted_total counts live edges removed, so a delete of an edge that was already absent writes a tombstone and increments neither counter, matching the affected count the statement reports.

Notes

  • WAL replay applies edges again on restart. The counters are gated so re-applied edges do not count as new writes; without that, a restart would inflate the write rate and mislead pacing.
  • The four handlers share one recording point per direction, so a single put and a batch put cannot drift apart.

Evidence

The tests fail on main without this change. Proof: the new test file was copied onto a clean origin/main (bd8da7dc2) worktree and run there first.

  • cargo nextest run -p nodedb --test wire -E 'test(~graph_edge_write_counters)' on a clean main: FAIL, 3 tests (graph_insert_edge_advances_the_write_counter, graph_delete_edge_counts_only_a_live_edge, the_counters_on_a_multi_core_server).
  • On this branch: graph_edge_write_counters 3/3 pass.
  • Crash recovery: crash_recovery_overlays 7/7 pass, including replay_does_not_count_graph_edges_as_new_writes.
  • cargo check -p nodedb, cargo fmt --all -- --check, and a lib clippy run are clean.
  • Mutation arm: with the two counter increments neutered, graph_insert_edge_advances_the_write_counter and graph_delete_edge_counts_only_a_live_edge both fail; restored, both pass.

What CI does not cover locally

  • 32-bit and ASan fuzz targets, and the full nightly matrix.

Closes #372

Copilot AI balanced review requested due to automatic review settings October 1, 2026 22:04

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

The graph engine exposed a gauge of live edges and nothing else, so an edge
put that rewrote a live edge was invisible: the gauge stayed flat and no
counter moved. An operator could not tell a hot rewrite path from an idle
one, and an edge delete that removed nothing looked the same as one that
removed an edge.

Add `nodedb_graph_edges_written_total` and
`nodedb_graph_edges_deleted_total` to SystemMetrics, record them from the
four edge write handlers, and render them on `/metrics` and as
`graph_edges_written_total` / `graph_edges_deleted_total` rows of
`SHOW STATS`.

Semantics: the write counter counts applied edge versions, so a rewrite
counts; the delete counter counts live edges tombstoned, so a delete of an
absent edge counts zero, matching the affected count that same statement
reports. Batch handlers count per edge as it is applied, never once for the
batch, so a batch that fails midway still reports the edges it already
wrote.
Boot replay re-enters the ordinary put and delete handlers to rebuild
engine state, and the metrics are attached before recovery runs. Every edge
replay re-applied was written by a client before the restart, so the new
counters reported it as fresh activity: a restart with no client write
increased the write counter by the size of the replayed edge redo set.

Hold a boot-replay flag across the boot pass and skip the counters while it
is set. The flag is raised in `replay_all_wal`, the boot entry, and cleared
before the core serves, so an online committed-redo apply — which reaches
the same handlers — stays a real write and still counts.

The delete counter ships as WIP and says so in the places an operator
reads: its HELP text, `docs/architecture.md`, and the doc comment on
`counts_logical_edge_delete`. It counts each endpoint home that removes a
live row, so on a multi-core cluster one cross-shard delete reads as two.
Counted on the source home instead it reads zero on one core, because the
destination home performs the removal there: the deciding fact is the
Control Plane's `single_home`, which the plan does not carry. The predicate
is the seam the fix needs, and `the_counters_on_a_multi_core_server` pins
the current behaviour so the fix has to change that test on purpose.

The write counter has no such caveat: the source home owns the logical edge
in both shapes, so it is correct on one core and on many.
@EnRaiha
EnRaiha force-pushed the pr/374-372-edge-write-counters branch from f88975f to 6d35de0 Compare October 2, 2026 07:24
@farhan-syah farhan-syah closed this Oct 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

obsv: add loader-facing write-pressure counters (graph edge writes, dispatch capacity busy)

3 participants