The question: you type cat a/b.txt. Something turns that string into a
position on a disk. Where exactly are the bytes, and can you find them yourself
with arithmetic instead of a filesystem driver?
You can. This chapter walks the chain twice: once through the kernel's directory
API, once through the raw image with dd, and checks that both arrive at the
same bytes.
flowchart LR
n["name<br/>hello.txt"] -->|"dirent in the parent's data block"| i["inode 17"]
i -->|"i_block[0]"| b["block 4608"]
b --> d["'find me on disk'"]
r["root inode 2"] -->|"dirent 'demo'"| p["inode 15<br/>directory"]
p -->|"i_block[0]"| pb["block 1039"]
pb --> n
A directory is a file. Its contents are records mapping names to inode numbers. That is the entire trick, and once you have seen the bytes it stops being a metaphor.
dirdump <directory> prints a directory's raw name-to-inode records: inode
number, record length, type, and the offset cookie the kernel hands back.
It uses getdents64(2)
and parses the byte stream itself. No opendir, no readdir: those exist
precisely to hide what this chapter is about.
- A directory with a hard link in it. How many records, how many distinct inode numbers?
- Records are variable length. What decides the length, and is it just the name?
d_offis called an offset. Offset into what?
sudo mkdir -p /mnt/dissect/demo && sudo chown $USER /mnt/dissect/demo
cd /mnt/dissect/demo
echo "find me on disk" > hello.txt
echo "second" > a-longer-name.txt
mkdir sub
ln hello.txt hardlink.txtmake ch2
./ch2-ext2/dirdump /mnt/dissect/demo ino reclen type d_off name
--- gdents64 -> 176 bytes ---
15 24 d( 4) 3764832237084642961 .
17 32 -( 8) 6113707203434424980 hardlink.txt
2 24 d( 4) 6725371938080703233 ..
17 32 -( 8) 8432168408353557006 hello.txt
20 24 d( 4) 8678366608238310076 sub
18 40 -( 8) 9223372036854775807 a-longer-name.txt
Cross-check with ls -lai:
15 drwxr-xr-x+ 3 vkit root 4096 .
2 drwxr-xr-x+ 5 vkit root 4096 ..
18 -rw-r--r-- 1 vkit vkit 7 a-longer-name.txt
17 -rw-r--r-- 2 vkit vkit 16 hardlink.txt
17 -rw-r--r-- 2 vkit vkit 16 hello.txt
20 drwxr-xr-x 2 vkit vkit 4096 sub
hello.txt and hardlink.txt are both inode 17. Two names, one inode, one
copy of the bytes. ls -l shows link count 2 on both. A hard link is not a
pointer to a file, it is a second name at the same rank as the first: delete
either one and the other is untouched, because what rm removes is a directory
record, not a file. The bytes go away when the last name does.
reclen grows with the name: 24, 32, 40. The kernel's record is a 19-byte
header (d_ino 8, d_off 8, d_reclen 2, d_type 1) followed by a
NUL-terminated name, padded up to an 8-byte boundary. hello.txt: 19 + 9 + 1 =
29, rounded to 32. a-longer-name.txt: 19 + 17 + 1 = 37, rounded to 40. That is
why the buffer you pass is a packed byte stream and not an array, and why walking
it with buf[i] indexing desyncs you mid-buffer and prints garbage names.
d_off is not an offset into anything you can compute with. The values here
are enormous and the last one is exactly INT64_MAX. It is an opaque cookie: the
only legal thing to do with it is hand it back to lseek to resume the scan. The
struct field's name is a historical lie the man page is honest about.
The order is not the order on disk, and neither matches creation order. Keep that in mind next time something depends on directory listing order.
debugfs reads the image file directly, so unmount first. Reading a mounted
filesystem's raw bytes gives you a torn view, because the kernel is holding state
in RAM that has not been written yet.
sync && sudo umount /mnt/dissect
cd ~/iolab
debugfs -R "stat /demo/hello.txt" ext2.imgInode: 17 Type: regular Mode: 0644 Flags: 0x0
User: 501 Group: 501 Size: 16
Links: 2 Blockcount: 8
BLOCKS:
(0):4608
TOTAL: 1
Inode 17, the same number dirdump printed, and its data lives in block 4608.
Block size is 4096, so the bytes start at byte 4608 × 4096 of the image. Go get
them:
dd if=ext2.img bs=4096 skip=4608 count=1 status=none | xxd | head -300000000: 6669 6e64 206d 6520 6f6e 2064 6973 6b0a find me on disk.
00000010: 0000 0000 0000 0000 0000 0000 0000 0000 ................
00000020: 0000 0000 0000 0000 0000 0000 0000 0000 ................
Your string, located by arithmetic, with no filesystem mounted.
The parent directory is also just a file. debugfs -R "stat /demo" ext2.img says
inode 15, block 1039. Dump that block and you are looking at the records
dirdump was showing you, in their on-disk form:
dd if=ext2.img bs=4096 skip=1039 count=1 status=none | xxd | head -800000000: 0f00 0000 0c00 0102 2e00 0000 0200 0000 ................
00000010: 0c00 0202 2e2e 0000 1100 0000 1400 0901 ................
00000020: 6865 6c6c 6f2e 7478 7400 0000 1200 0000 hello.txt.......
00000030: 1c00 1101 612d 6c6f 6e67 6572 2d6e 616d ....a-longer-nam
00000040: 652e 7478 7400 0000 1400 0000 0c00 0302 e.txt...........
00000050: 7375 6200 1100 0000 ac0f 0c01 6861 7264 sub.........hard
00000060: 6c69 6e6b 2e74 7874 0000 0000 0000 0000 link.txt........
00000070: 0000 0000 0000 0000 0000 0000 0000 0000 ................
The on-disk record is __le32 inode · __le16 rec_len · __u8 name_len ·
__u8 file_type · char name[], so little-endian, eight bytes of header:
| bytes | inode | rec_len | name_len | type | name |
|---|---|---|---|---|---|
0f000000 0c00 01 02 |
15 | 12 | 1 | 2 (dir) | . |
02000000 0c00 02 02 |
2 | 12 | 2 | 2 (dir) | .. |
11000000 1400 09 01 |
17 | 20 | 9 | 1 (file) | hello.txt |
12000000 1c00 11 01 |
18 | 28 | 17 | 1 (file) | a-longer-name.txt |
14000000 0c00 03 02 |
20 | 12 | 3 | 2 (dir) | sub |
11000000 ac0f 0c 01 |
17 | 4012 | 12 | 1 (file) | hardlink.txt |
Three things fall out of that table.
There are two different dirent formats and you just saw both. On disk,
hello.txt has rec_len 20. Through getdents64, the same entry has d_reclen
32. The on-disk record is an 8-byte header and a name padded to 4; the kernel's
is a 19-byte header and a NUL-terminated name padded to 8. struct linux_dirent64 is an API shape, not a disk shape, and no filesystem stores it.
The last entry's rec_len is 4012, which is not its size, it is everything
left in the 4096-byte block. 12 + 12 + 20 + 28 + 12 + 4012 = 4096 exactly. The
final record always swallows the remainder, so the walk terminates at the block
edge rather than on a count.
Inode 17 appears twice, which is the hard link again, now visible as two records with the same four leading bytes.
Delete one name and dump the same block again:
sudo mount /mnt/dissect && rm /mnt/dissect/demo/hello.txt
sync && sudo umount /mnt/dissect
dd if=ext2.img bs=4096 skip=1039 count=1 status=none | xxd | head -300000000: 0f00 0000 0c00 0102 2e00 0000 0200 0000 ................
00000010: 2000 0202 2e2e 0000 0000 0000 0000 0000 ...............
00000020: 0000 0000 0000 0000 0000 0000 1200 0000 ................
One field moved. The .. entry's rec_len went from 0c00 (12) to 2000
(32): it swallowed the 20 bytes that used to be hello.txt. Nothing was
rearranged and no record was shifted. The name is now unreachable because the
walk steps over it, not because it was removed from a list.
On this kernel the vacated bytes were also zeroed, so the name is gone from the hex too. That zeroing is not something the on-disk format requires, and it is exactly the detail that decides whether an undelete tool can find anything. Worth re-running on your own filesystem rather than assuming either way.
The file's data block is still intact at this point, and inode 17 still has a
link count of 1 because hardlink.txt still names it.
- Unmount before
debugfs. A mounted filesystem's raw bytes lag what the kernel thinks is true. - Records are variable length. Step through the buffer by adding
d_reclen, and nothing else. d_nameis NUL-terminated with padding after it. Deriving the name length arithmetically fromd_reclenwill include that padding and bite you.getdents64returns as many records as fit in your buffer, and may need several calls. One call is not the whole directory.struct linux_dirent64is not exported by any header. The definition in the man page is the definition; you copy it into your own source, which is whydirdump.ccarries astatic_asserton the offset ofd_nameto catch a typo in the copy.debugfsblock numbers are in filesystem blocks.ddneeds a matchingbs. Mismatch these and you will read a plausible-looking wrong block.
man 2 getdents64. Its example program is worth diffing against yours after yours works.man 8 debugfs:stat,blocks,ls,imap,icheck.- Poirier, The Second Extended File System, for the on-disk structures.