Skip to content

Latest commit

 

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..
 
 
 
 
 
 

README.md

ch1 - Seeing the page cache: resident, cowpfn

The question: which pages of a file are in RAM right now? And can you find out without changing the answer by asking?

That second half is the entire design constraint. Reading the file to check whether it is cached would cache it. The measurement has to not touch the data.

The programs

resident <path> prints one character per page of a file: # if that page is in the page cache, . if it is not, plus a summary line. It maps the file and asks mincore(2), which reports residency without faulting anything in. mmap alone moves no data.

cowpfn resolves a virtual address to the physical frame behind it via /proc/self/pagemap, then forks and shows what copy-on-write actually does to that frame. Needs sudo.

Predict first

  1. Drop the caches, then read() exactly 4 KiB from offset 0 of a cold file. How many pages are resident afterwards?
  2. Same read, but starting at offset 4096 instead of 0. Same answer?
  3. fork() a process holding a 4 KiB anonymous page. How many physical frames exist immediately after the fork? After the child writes one byte?

Write the numbers down. Two of the three are probably wrong.

Run it

make ch1
dd if=/dev/urandom of=/mnt/work/small.dat bs=1M count=1

# cold, then warm. One command line: never leave a gap between load and measure.
sync; echo 3 | sudo tee /proc/sys/vm/drop_caches; ./ch1-cache/resident /mnt/work/small.dat
sync; echo 3 | sudo tee /proc/sys/vm/drop_caches; cat /mnt/work/small.dat >/dev/null; ./ch1-cache/resident /mnt/work/small.dat

What you get

Cold, right after dropping the caches:

page        0  ........ ........ ........ ........ ........ ........ ........ ........    0/64
page       64  ........ ........ ........ ........ ........ ........ ........ ........    0/64
page      128  ........ ........ ........ ........ ........ ........ ........ ........    0/64
page      192  ........ ........ ........ ........ ........ ........ ........ ........    0/64

resident: 0 / 256 pages (0.0%)  0.0 B / 1.0 MiB

Warm, after one cat:

page        0  ######## ######## ######## ######## ######## ######## ######## ########   64/64
page       64  ######## ######## ######## ######## ######## ######## ######## ########   64/64
page      128  ######## ######## ######## ######## ######## ######## ######## ########   64/64
page      192  ######## ######## ######## ######## ######## ######## ######## ########   64/64

resident: 256 / 256 pages (100.0%)  1.0 MiB / 1.0 MiB

That is the sanity check. Cross-check it against vmtouch small.dat if you want a second opinion before trusting your own instrument.

The experiment worth doing: one read, four pages

Drop the caches, read exactly one 4 KiB page, and look at what arrived.

sync; echo 3 | sudo tee /proc/sys/vm/drop_caches; \
  dd if=/mnt/work/small.dat bs=4096 count=1 of=/dev/null; \
  ./ch1-cache/resident /mnt/work/small.dat
### read at page 0
page        0  ####.... ........ ...
resident: 4 / 256 pages (1.6%)  16.0 KiB / 1.0 MiB

### the same read, one page in (dd ... skip=1)
page        0  .#...... ........ ...
resident: 1 / 256 pages (0.4%)  4.0 KiB / 1.0 MiB

### the same read, at page 40 (dd ... skip=40)
page        0  ........ ........ ........ ........ ........ #....... ...
resident: 1 / 256 pages (0.4%)  4.0 KiB / 1.0 MiB

### no read at all
resident: 0 / 256 pages (0.0%)  0.0 B / 1.0 MiB

Same file, same syscall, same size, same kernel. The only thing that changed is where the read started, and the amount of I/O the kernel performed changed by 4x.

An offset of 0 is treated as a strong hint that you are about to stream the whole file, so the kernel seeds a readahead window immediately. Any other starting offset is treated as random access until a pattern proves otherwise. Keep reading sequentially and the window ramps up; on this bench, asking for 16 sequential pages leaves 24 resident:

### 16 sequential 4 KiB reads (64 KiB asked for)
page        0  ######## ######## ######## ........ ...   24/64
resident: 24 / 256 pages (9.4%)  96.0 KiB / 1.0 MiB

The ceiling is read_ahead_kb, 128 KiB on this bench, which is 32 pages:

losetup -a                                       # find your loop device first
cat /sys/block/loop1/queue/read_ahead_kb         # 128

Why this matters outside kernel work: a cold-cache benchmark that opens a file and reads the first block is not measuring one block of I/O. Move the same benchmark one page in and the number changes with nothing else different. The kernel is guessing at your access pattern constantly, it is usually right, and the time it guesses differently than you assumed tends to be the time you are holding a stopwatch.

The other experiment: read() misses are not page faults

Read two cold files of very different sizes and compare the fault counters against the blocks actually read:

for f in small.dat big.dat; do          # 1 MiB and 64 MiB
  sync; echo 3 | sudo tee /proc/sys/vm/drop_caches >/dev/null
  /usr/bin/time -v cat /mnt/work/$f 2>&1 >/dev/null \
    | grep -E "Major \(|File system inputs"
done
### cold cat of small.dat (1048576 bytes)
	Major (requiring I/O) page faults: 5
	File system inputs: 7920
### cold cat of big.dat (67108864 bytes)
	Major (requiring I/O) page faults: 5
	File system inputs: 136944

64x the file, 17x the blocks pulled off the device, and the same five major faults. Those five are cat and libc being paged back in, because drop_caches evicted the executable too. The 64 MiB of file data contributed exactly zero.

A read() that misses the page cache blocks in the syscall and issues I/O without ever taking a fault, because nothing was mapped into your address space to fault on. handle_mm_fault(), where maj_flt is incremented, never runs. Only mmap and swap move that counter.

If you have ever used maj_flt as a proxy for "how much disk did this touch", this is where that stops being true. Use /proc/<pid>/io or iostat -x 1 instead.

The mirror image, mmap the same cold file and touch every page, should push major faults up instead. I could not show that cleanly on this bench: the loop image's backing file is cached by the host, so the pages arrive without a real device trip and land as minor faults. On hardware with a genuinely cold device it climbs. Treat that half as read, not measured, until you run it yourself.

cowpfn: fork copies page tables, not pages

sudo ./ch1-cache/cowpfn
  stage                              VA         PFN       physical
  ---------------------------------------------------------------------
  parent (before fork)   0xffffa4769000  0x00172ea4  0x00172ea4000
  child  (after fork)    0xffffa4769000  0x00172ea4  0x00172ea4000
  child  (after write)   0xffffa4769000  0x0022d568  0x0022d568000
  parent (after wait)    0xffffa4769000  0x00172ea4  0x00172ea4000

  COW confirmed: the child's frame changed on write, the parent's did not.
flowchart LR
  subgraph after["after fork(), before any write"]
    pv1["parent VA<br/>0xffffa4769000"] --> f1["frame 0x172ea4<br/>read-only in both"]
    cv1["child VA<br/>0xffffa4769000"] --> f1
  end
  subgraph write["after the child writes one byte"]
    pv2["parent VA<br/>0xffffa4769000"] --> f2["frame 0x172ea4<br/>unchanged"]
    cv2["child VA<br/>0xffffa4769000"] --> f3["frame 0x22d568<br/>fresh copy"]
  end
  after --> write
Loading

Same virtual address in both processes, same physical frame, until one byte is written. The write traps, the kernel allocates a new frame, copies 4 KiB into it, and repoints only that process's page table entry. The parent never notices.

Two details that make or break this program:

  • It writes to the page (memset(p, 'A', ps)) before forking. An anonymous page that has only ever been read maps the system-wide shared zero page, so both processes would appear to share a frame for reasons that have nothing to do with fork or COW. You would get the right output for the wrong reason.
  • Without CAP_SYS_ADMIN every PFN reads back as 0, so the program checks for that and tells you to rerun under sudo instead of printing a table of zeroes:
$ ./ch1-cache/cowpfn
  parent (before fork)   0xffff9a5ce000  0x00000000  0x00000000000

  PFN is 0 -- rerun under sudo (needs CAP_SYS_ADMIN).

Traps

  • Never separate loading the cache from measuring it. Chain drop, load and measure on one command line. Anything else on the machine can evict your pages in the gap, and "the page cache is empty again" almost always means the environment, not the behaviour you were studying.
  • mmap with length 0 fails EINVAL, so an empty file looks like a bug in your code. Guard st_size == 0 before mapping.
  • Dropping caches needs a sync first, or dirty pages simply will not drop.
  • Generate test files with dd, not truncate or fallocate. The latter two create sparse files with no real data blocks, and a read from a hole takes a different kernel path than a read from real data. Your readahead experiment would quietly measure nothing.
  • Find your loop device with losetup -a before reading /sys/block/loop<N>/queue/read_ahead_kb. Guessing loop0 is a coin flip.
  • mincore's address argument must be page-aligned. Whatever mmap returned always is, so this only bites if you do pointer arithmetic on it first.
  • Only bit 0 of each vec byte is defined. The other seven are reserved, so compare with & 1, never != 0.

Read

  • man 2 mincore. The paragraph describing vec is the whole contract and is easy to misread.
  • man 2 mmap, in particular everything that does not happen when you call it.
  • man 3 sysconf for _SC_PAGESIZE. Do not hardcode 4096; arm64 ships 16 KiB pages on some configurations.
  • Kernel docs: admin-guide/mm/pagemap.rst for the pagemap entry layout, and mm/readahead.c if you want the offset-0 special case in the source.