Skip to content

Combine NAR blocks into fixed-size chunks and bound the channel - #373

Open
emmakloe wants to merge 1 commit into
zhaofengli:mainfrom
emmakloe:nar-fixed-size-chunks
Open

emmakloe wants to merge 1 commit into
zhaofengli:mainfrom
emmakloe:nar-fixed-size-chunks

Conversation

@emmakloe

@emmakloe emmakloe commented Sep 21, 2026 •

Copy link
Copy Markdown

The C++ serializer emits the NAR incrementally and calls back into Rust for each piece. Every callback does a Vec::from(data), so each one becomes its own heap allocation. The stream contains a lot of metadata along with the ~64 KiB content chunks, which means we end up doing a huge number of allocations of all sorts of different sizes.

Attic is not asking the kernel for that memory directly, it goes through glibc, which splits the heap into arenas to reduce lock contention between threads (up to 8 x nproc of them). Freeing a Vec does not mean the memory is returned to the OS. With lots of differently sized allocations, the arenas become fragmented, which makes it harder for glibc to hand large chunks back. So from Attic's point of view the objects have been freed, but from the kernel's point of view the process can still own a large amount of memory. On a machine with a high core count and a lot of store paths going through watch-store, that adds up to tens of GiB.

The other half of the problem is that the channel between the serializer and the uploader is unbounded, so there is no back pressure. The producer runs at disk speed while the consumer is limited by network speed, and any difference between the two just gets queued.

The patch fixes this where the NAR crosses from the serializer into the uploader. Instead of turning every callback into its own Vec, the sender combines the incoming pieces into uniformly sized buffers of NAR_CHUNK_SIZE and only sends when one is full, flushing the remainder at EOF. The allocator then mostly sees buffers of the same size, which is a much more regular pattern and avoids the fragmentation. The channel is also made bounded, so the serializer has to wait once NAR_CHANNEL_CAPACITY buffers are queued instead of running ahead indefinitely.

That gives an actual upper bound on what can sit between serialization and upload:

NAR_CHANNEL_CAPACITY x NAR_CHUNK_SIZE = 64 x 64 KiB = 4 MiB

So the two parts are:

  • fixed-size buffers reduce allocator fragmentation and the amount of memory glibc retains after uploads complete
  • a bounded channel stops the producer accumulating an unlimited amount of NAR data when serialization is faster than the upload

Note that #362 proposes the bounded channel half of this. This includes that and adds the coalescing, which covers the fragmentation side that bounding alone does not: with a Vec per callback the queue is capped but the churn remains.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant