The question: which pages of a file are in RAM right now? And can you find out without changing the answer by asking?
That second half is the entire design constraint. Reading the file to check whether it is cached would cache it. The measurement has to not touch the data.
resident <path> prints one character per page of a file: # if that page is
in the page cache, . if it is not, plus a summary line. It maps the file and
asks mincore(2), which
reports residency without faulting anything in. mmap alone moves no data.
cowpfn resolves a virtual address to the physical frame behind it via
/proc/self/pagemap, then forks and shows what copy-on-write actually does to
that frame. Needs sudo.
- Drop the caches, then
read()exactly 4 KiB from offset 0 of a cold file. How many pages are resident afterwards? - Same read, but starting at offset 4096 instead of 0. Same answer?
fork()a process holding a 4 KiB anonymous page. How many physical frames exist immediately after the fork? After the child writes one byte?
Write the numbers down. Two of the three are probably wrong.
make ch1
dd if=/dev/urandom of=/mnt/work/small.dat bs=1M count=1
# cold, then warm. One command line: never leave a gap between load and measure.
sync; echo 3 | sudo tee /proc/sys/vm/drop_caches; ./ch1-cache/resident /mnt/work/small.dat
sync; echo 3 | sudo tee /proc/sys/vm/drop_caches; cat /mnt/work/small.dat >/dev/null; ./ch1-cache/resident /mnt/work/small.datCold, right after dropping the caches:
page 0 ........ ........ ........ ........ ........ ........ ........ ........ 0/64
page 64 ........ ........ ........ ........ ........ ........ ........ ........ 0/64
page 128 ........ ........ ........ ........ ........ ........ ........ ........ 0/64
page 192 ........ ........ ........ ........ ........ ........ ........ ........ 0/64
resident: 0 / 256 pages (0.0%) 0.0 B / 1.0 MiB
Warm, after one cat:
page 0 ######## ######## ######## ######## ######## ######## ######## ######## 64/64
page 64 ######## ######## ######## ######## ######## ######## ######## ######## 64/64
page 128 ######## ######## ######## ######## ######## ######## ######## ######## 64/64
page 192 ######## ######## ######## ######## ######## ######## ######## ######## 64/64
resident: 256 / 256 pages (100.0%) 1.0 MiB / 1.0 MiB
That is the sanity check. Cross-check it against vmtouch small.dat if you want
a second opinion before trusting your own instrument.
Drop the caches, read exactly one 4 KiB page, and look at what arrived.
sync; echo 3 | sudo tee /proc/sys/vm/drop_caches; \
dd if=/mnt/work/small.dat bs=4096 count=1 of=/dev/null; \
./ch1-cache/resident /mnt/work/small.dat### read at page 0
page 0 ####.... ........ ...
resident: 4 / 256 pages (1.6%) 16.0 KiB / 1.0 MiB
### the same read, one page in (dd ... skip=1)
page 0 .#...... ........ ...
resident: 1 / 256 pages (0.4%) 4.0 KiB / 1.0 MiB
### the same read, at page 40 (dd ... skip=40)
page 0 ........ ........ ........ ........ ........ #....... ...
resident: 1 / 256 pages (0.4%) 4.0 KiB / 1.0 MiB
### no read at all
resident: 0 / 256 pages (0.0%) 0.0 B / 1.0 MiB
Same file, same syscall, same size, same kernel. The only thing that changed is where the read started, and the amount of I/O the kernel performed changed by 4x.
An offset of 0 is treated as a strong hint that you are about to stream the whole file, so the kernel seeds a readahead window immediately. Any other starting offset is treated as random access until a pattern proves otherwise. Keep reading sequentially and the window ramps up; on this bench, asking for 16 sequential pages leaves 24 resident:
### 16 sequential 4 KiB reads (64 KiB asked for)
page 0 ######## ######## ######## ........ ... 24/64
resident: 24 / 256 pages (9.4%) 96.0 KiB / 1.0 MiB
The ceiling is read_ahead_kb, 128 KiB on this bench, which is 32 pages:
losetup -a # find your loop device first
cat /sys/block/loop1/queue/read_ahead_kb # 128Why this matters outside kernel work: a cold-cache benchmark that opens a file and reads the first block is not measuring one block of I/O. Move the same benchmark one page in and the number changes with nothing else different. The kernel is guessing at your access pattern constantly, it is usually right, and the time it guesses differently than you assumed tends to be the time you are holding a stopwatch.
Read two cold files of very different sizes and compare the fault counters against the blocks actually read:
for f in small.dat big.dat; do # 1 MiB and 64 MiB
sync; echo 3 | sudo tee /proc/sys/vm/drop_caches >/dev/null
/usr/bin/time -v cat /mnt/work/$f 2>&1 >/dev/null \
| grep -E "Major \(|File system inputs"
done### cold cat of small.dat (1048576 bytes)
Major (requiring I/O) page faults: 5
File system inputs: 7920
### cold cat of big.dat (67108864 bytes)
Major (requiring I/O) page faults: 5
File system inputs: 136944
64x the file, 17x the blocks pulled off the device, and the same five major
faults. Those five are cat and libc being paged back in, because
drop_caches evicted the executable too. The 64 MiB of file data contributed
exactly zero.
A read() that misses the page cache blocks in the syscall and issues I/O
without ever taking a fault, because nothing was mapped into your address space
to fault on. handle_mm_fault(), where maj_flt is incremented, never runs.
Only mmap and swap move that counter.
If you have ever used maj_flt as a proxy for "how much disk did this touch",
this is where that stops being true. Use /proc/<pid>/io or iostat -x 1
instead.
The mirror image, mmap the same cold file and touch every page, should push
major faults up instead. I could not show that cleanly on this bench: the loop
image's backing file is cached by the host, so the pages arrive without a real
device trip and land as minor faults. On hardware with a genuinely cold device
it climbs. Treat that half as read, not measured, until you run it yourself.
sudo ./ch1-cache/cowpfn stage VA PFN physical
---------------------------------------------------------------------
parent (before fork) 0xffffa4769000 0x00172ea4 0x00172ea4000
child (after fork) 0xffffa4769000 0x00172ea4 0x00172ea4000
child (after write) 0xffffa4769000 0x0022d568 0x0022d568000
parent (after wait) 0xffffa4769000 0x00172ea4 0x00172ea4000
COW confirmed: the child's frame changed on write, the parent's did not.
flowchart LR
subgraph after["after fork(), before any write"]
pv1["parent VA<br/>0xffffa4769000"] --> f1["frame 0x172ea4<br/>read-only in both"]
cv1["child VA<br/>0xffffa4769000"] --> f1
end
subgraph write["after the child writes one byte"]
pv2["parent VA<br/>0xffffa4769000"] --> f2["frame 0x172ea4<br/>unchanged"]
cv2["child VA<br/>0xffffa4769000"] --> f3["frame 0x22d568<br/>fresh copy"]
end
after --> write
Same virtual address in both processes, same physical frame, until one byte is written. The write traps, the kernel allocates a new frame, copies 4 KiB into it, and repoints only that process's page table entry. The parent never notices.
Two details that make or break this program:
- It writes to the page (
memset(p, 'A', ps)) before forking. An anonymous page that has only ever been read maps the system-wide shared zero page, so both processes would appear to share a frame for reasons that have nothing to do with fork or COW. You would get the right output for the wrong reason. - Without
CAP_SYS_ADMINevery PFN reads back as 0, so the program checks for that and tells you to rerun under sudo instead of printing a table of zeroes:
$ ./ch1-cache/cowpfn
parent (before fork) 0xffff9a5ce000 0x00000000 0x00000000000
PFN is 0 -- rerun under sudo (needs CAP_SYS_ADMIN).
- Never separate loading the cache from measuring it. Chain drop, load and measure on one command line. Anything else on the machine can evict your pages in the gap, and "the page cache is empty again" almost always means the environment, not the behaviour you were studying.
mmapwith length 0 failsEINVAL, so an empty file looks like a bug in your code. Guardst_size == 0before mapping.- Dropping caches needs a
syncfirst, or dirty pages simply will not drop. - Generate test files with
dd, nottruncateorfallocate. The latter two create sparse files with no real data blocks, and a read from a hole takes a different kernel path than a read from real data. Your readahead experiment would quietly measure nothing. - Find your loop device with
losetup -abefore reading/sys/block/loop<N>/queue/read_ahead_kb. Guessingloop0is a coin flip. mincore's address argument must be page-aligned. Whatevermmapreturned always is, so this only bites if you do pointer arithmetic on it first.- Only bit 0 of each
vecbyte is defined. The other seven are reserved, so compare with& 1, never!= 0.
man 2 mincore. The paragraph describingvecis the whole contract and is easy to misread.man 2 mmap, in particular everything that does not happen when you call it.man 3 sysconffor_SC_PAGESIZE. Do not hardcode 4096; arm64 ships 16 KiB pages on some configurations.- Kernel docs:
admin-guide/mm/pagemap.rstfor the pagemap entry layout, andmm/readahead.cif you want the offset-0 special case in the source.