Skip to content

perf: coalesce object access updates - #371

Open
Traviis wants to merge 2 commits into
zhaofengli:mainfrom
Traviis:push-woqzwxupuplo
Open

Traviis wants to merge 2 commits into
zhaofengli:mainfrom
Traviis:push-woqzwxupuplo

Conversation

@Traviis

@Traviis Traviis commented Sep 16, 2026 •

Copy link
Copy Markdown

Move NAR access timestamp updates out of the download response path and coalesce repeated
accesses.

Previously, every successful NAR request synchronously updated object.last_accessed_at before
returning the response. With remote PostgreSQL, this adds a database round trip to every
download and generates unnecessary writes when an object is requested repeatedly.

This change adds a bounded background access tracker:

  • NAR requests enqueue the object ID without waiting for PostgreSQL.
  • A background worker collects IDs for up to one second.
  • Duplicate object IDs within each collection window are collapsed.
  • A batch is bounded to 4,096 unique IDs.
  • Each unique object retains its access timestamp update.
  • Queue saturation drops the tracking update with a warning rather than delaying the download.
  • Database update failures are logged.

The worker currently performs one database update per unique object ID. This change reduces
duplicate writes and removes access tracking from response latency; it does not introduce a
bulk SQL update. (EDIT: No longer true, see comment)

Also fixes the PostgreSQL integration-test connection URL to explicitly use the atticd
database user.

last_accessed_at supports retention and observability. It does not need to block a NAR
download or record every individual access precisely.

In the benchmark workload, 8,000 requests to one object previously caused 8,000 synchronous
PostgreSQL updates. Coalescing reduced this to approximately 8 updates while preserving access
tracking.

Benchmark results

NixOS VM benchmark using Attic, PostgreSQL, and Garage S3:

  • NAR throughput: 618–832 → 1,028–1,257 requests/second
  • Concurrency-1 p50: 1.600 → 0.897 ms
  • Concurrency-128 p50: 160.894 → 110.814 ms
  • Attic CPU: 9.907 → 8.557 seconds
  • Timestamp writes for 8,000 accesses to one object: 8,000 → 8
  • All measured requests succeeded

Failure behavior

Access tracking is intentionally best-effort:

  • A full queue logs a warning and drops the timestamp update.
  • A database failure logs an error.
  • Neither condition delays or fails the NAR download.

Testing

  • Unit test verifies duplicate IDs are coalesced.
  • Unit test verifies batch size is bounded.
  • Server unit tests and doctests pass.
  • Clippy passes with warnings denied.
  • PostgreSQL/Garage NixOS VM integration test passes.

@Traviis

Traviis commented Sep 16, 2026

Copy link
Copy Markdown
Author

The initial version moved timestamp writes off the request path and deduplicated repeated
object IDs, but still issued one SQL UPDATE per unique object in each batch. This meant
high-cardinality batches could still generate up to 4,096 PostgreSQL round trips.

The worker now updates the entire batch with one statement:

  UPDATE object
  SET last_accessed_at = NOW()
  WHERE id IN (...);

A dedicated PostgreSQL benchmark compared the previous per-object writes with the bulk
statement:

Unique objects Per-object updates Bulk update Speedup
1 0.594 ms 0.614 ms effectively neutral
16 7.473 ms 0.704 ms 10.6×
256 124.287 ms 2.256 ms 55.1×
1,024 492.219 ms 6.402 ms 76.9×

This keeps the single-object case effectively unchanged while substantially reducing database
round trips and batch processing time when many different objects are accessed during the same
flush window.

The benchmark also verifies that every expected object receives an updated timestamp.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant