The previous lesson ended with a deliberate loose end. We know that inode 1180934 is the readings file for 31 August: it holds its size, its owner, its dates and its blocks. And we know it does not hold its name. But then, when you type cat /var/lib/meteora/readings/2026-08-31.dat, who turns that 43-character string into the number 1180934?
The answer is the directory, and it deserves a whole lesson for three reasons. The first is that its design explains behaviors that baffle everybody: why deleting a file is called "unlinking", why a deleted file still takes up 17 GB, why a symbolic link breaks and a hard one does not, and why you cannot hard-link a directory. The second is that path resolution is one of the most frequently executed operations in the entire system — millions of times per second on a loaded server — and understanding its cost explains real design decisions. And the third is that organizing names is a classic data-structures problem with measurable consequences: a directory with 100,000 readings files behaves radically differently depending on how it is organized inside.
We are going to open a directory and see what is inside, walk through the historical evolution of the organizations up to the UNIX tree, follow step by step how a path is turned into an inode while counting the disk accesses, distinguish precisely between the two kinds of link, measure the cost of searching enormous directories, and place Meteora's paths within the FHS standard.
Contents
- A directory is a file containing (name, inode) pairs
- Evolution of directory organizations
- Absolute and relative paths,
.and.. - The working directory:
chdir,getcwdand why it is per process - Path resolution step by step, with the cost counted
- The dentry cache: why that does not cost what it looks like
- Hard links: one inode, several names
- Symbolic links: a file containing a path
- Internal organization and the cost of enormous directories
- The FHS standard and why Meteora is where it is
- Deletion: what
unlinkreally does
A directory is a file containing (name, inode) pairs
Let us start with the central claim, which usually sounds odd the first time:
A directory is an ordinary file whose content is a table of (name, inode number) pairs. Its only peculiarity is that its inode's type is
d, and that the kernel reserves the right to write it itself.
It has its inode, its size, its permissions, its dates and its data blocks, exactly like 2026-08-31.dat. What changes is what is inside those blocks.
Let us prove it. First, the directory has a size and blocks like any file:
$ stat /var/lib/meteora/readings File: /var/lib/meteora/readings Size: 4096 Blocks: 8 IO Block: 4096 directory Device: 9,0 Inode: 1180929 Links: 3 Access: (0750/drwxr-x---) Uid: (990/meteora) Gid: (990/meteora)
4,096 bytes, 8 blocks of 512 B, inode 1180929. It is a file. Now let us try to read it as such:
It fails, and not because there is no content, but because since 1988 Linux has forbidden read() on a directory: if every program parsed the binary format by hand, changing ext4's internal structure would break half the system. In its place there is a specific system call, getdents64(), which returns the entries in a format independent of the file system. It is the same idea of indirection we will meet as the VFS in Partitions, Mounting and the Virtual File System.
With strace you can see that ls uses it:
$ strace -e trace=openat,getdents64 ls /var/lib/meteora/readings openat(AT_FDCWD, "/var/lib/meteora/readings", O_RDONLY|O_NONBLOCK|O_DIRECTORY|O_CLOEXEC) = 3 getdents64(3, 0x5586a2f1b0d0 /* 8 entries */, 32768) = 240 getdents64(3, 0x5586a2f1b0d0 /* 0 entries */, 32768) = 0
ls opens the directory with O_DIRECTORY, asks for its entries with getdents64 — eight at once, 240 bytes — and calls again until it gets 0, which means "no more". Every ls, every find and every os.listdir() ends up here.
To see the raw content you need debugfs:
$ sudo debugfs -R "ls -l /var/lib/meteora/readings" /dev/md0 1180929 40750 (2) 990 990 4096 1-Sep-2026 00:00 . 1180928 40750 (2) 990 990 4096 1-Sep-2026 00:00 .. 1180934 100640 (1) 990 990 17280000 31-Aug-2026 23:59 2026-08-31.dat 1180935 100640 (1) 990 990 17280000 1-Sep-2026 12:40 2026-09-01.dat 1180936 120777 (7) 990 990 14 1-Sep-2026 00:00 today.dat 1180940 40750 (2) 990 990 4096 1-Sep-2026 00:00 archive
There is the table: six (name, inode) pairs. The first column is the inode number and the last one is the name; in between, the mode in octal with the type in front (40 directory, 100 regular, 120 symbolic link).
The real format of an ext4 entry, which is almost disappointingly simple:
struct ext4_dir_entry_2 {
__le32 inode; /* 4 B: inode number, 0 = deleted entry */
__le16 rec_len; /* 2 B: TOTAL length of this entry */
__u8 name_len; /* 1 B: length of the name (max. 255) */
__u8 file_type; /* 1 B: 1=regular 2=dir 7=symlink... */
char name[]; /* the name, with no null terminator */
};Eight bytes of header plus the name, rounded up to a multiple of 4. The entry for 2026-08-31.dat (14 characters) takes 8 + 14 = 22 → 24 bytes. Two details with consequences:
rec_lenis the total length, not the name's. Walking the directory means advancingrec_lenbytes from each entry. When an entry is deleted, the system moves nothing: it simply adds itsrec_lento the previous entry's, which "swallows" it. That is why a directory that once had a million files and now has three still takes up megabytes: the holes exist but the file does not shrink.file_typeduplicates information that is already in the inode. It is a deliberate optimization: it letsfind -type ffilter without reading each entry's inode, saving one access per file.
And the detail with the most consequences, already visible above: the same inode number can appear in several entries. Those are the hard links of section 7.
Evolution of directory organizations
The UNIX tree did not fall from the sky. It is the fourth attempt at solving the same problem, and knowing the three earlier ones explains why the current one is the way it is.
Single level. All files in one flat namespace; it is what the first batch systems had and what many embedded devices still have. Trivial to implement, but with a fatal problem: name collisions. If two users both want a data.dat, one of them loses.
Two levels. One directory per user hanging off a master directory. It solves the collisions and provides isolation, but it is rigid: a user cannot organize their 400 files into subgroups, and sharing with another user is awkward. Meteora could not separate readings/ from archive/.
Hierarchical tree. A directory can contain directories, with no depth limit. It is the structure of UNIX and of Windows: natural organization by topic, unique names guaranteed by the full path, and exactly one route from the root to each file. Its limitation is that it does not allow real sharing: if 2026-08-31.dat must also appear in /srv/publishing/, the only option is to copy it — 34 MB instead of 17, and divergence as soon as one of them changes.
Acyclic graph. The tree plus the possibility that several names point to the same file, as long as no cycles are formed. It is what UNIX achieves with links. It brings real sharing, but it brings two new problems: a traversal may visit a file twice (find, du and tar keep track of the inodes they have already seen), and deletion stops being obvious — if two paths are the same inode, does deleting one destroy the content? The solution is the link count of section 7.
General graph. If cycles are also allowed — a directory that contains itself by an indirect route — two serious problems appear:
- Infinite traversals. An unprotected
findwould enter an eternal loop. Cycles would have to be detected on every traversal, at a cost in time and memory. - Garbage collection. The link count stops working. If
A/contains a link toB/andB/contains a link toA/, and we delete both names from outside, both counters still read 1 — they hold each other up — but nobody can reach them. That space would be lost forever unless a reachability-based garbage collector were run, walking the whole disk from the root.
And that is why UNIX makes a blunt decision that you now understand:
Hard links to directories are not allowed. Only
.and.., which the kernel itself creates and which are controlled, known cycles.
With that single prohibition, the directory structure is guaranteed to be acyclic, and the link count is enough as a garbage collector: there can never be an unreachable island with a positive count. It is a small restriction in exchange for eliminating two entire classes of problems.
| Organization | Collisions | Sharing | Cycles | Safe deletion | Used today |
|---|---|---|---|---|---|
| Single level | Yes | No | No | Trivial | Simple embedded systems |
| Two levels | No | Hard | No | Trivial | Historical systems |
| Tree | No | No | No | Trivial | FAT (no links) |
| Acyclic graph | No | Yes | No | Link count | UNIX, Linux, NTFS |
| General graph | No | Yes | Yes | Needs a collector | None serious |
Absolute and relative paths, . and ..
A path is the recipe for reaching a file by walking the graph.
- Absolute: starts with
/and departs from the root./var/lib/meteora/readings/2026-08-31.dat. It means the same thing whoever runs it and wherever they run it from. - Relative: does not start with
/and departs from the process's current working directory.readings/2026-08-31.datmeans different things depending on where you are.
Every directory always contains two entries that the kernel creates when it creates the directory:
| Entry | Points to | What it is for |
|---|---|---|
. |
Its own inode | Referring to the current directory: ./program, cp file . |
.. |
Its parent's inode | Going up one level: cd .., ../config/ |
At the root there is an elegant exception: /.. points to /. The root is its own parent, so cd /../../../.. leaves you in / with no error. That is what stops you from walking off the top of the tree.
These two entries explain directories' link count, which confuses a lot of people:
Why 3? Because three names point to that inode:
readings, the entry in the parent directory/var/lib/meteora/.., insidereadings/itself..., inside its only subdirectory,archive/.
Hence the rule: a directory's link count is 2 + the number of subdirectories. A practical trick comes out of this: to find out how many subdirectories a directory has without walking it, stat -c %h and subtract 2. With a directory of 100,000 entries, that is instantaneous compared with the seconds ls -l | grep ^d would take.
The working directory: chdir, getcwd and why it is per process
Every process has its current working directory (cwd), stored in its task_struct — the structure we dissected in Process Management. It is a pointer to a directory inode, not a text string.
You can see it in /proc, like everything else:
$ ls -l /proc/2841/cwd lrwxrwxrwx 1 meteora meteora 0 Sep 1 12:44 /proc/2841/cwd -> /var/lib/meteora
And it is manipulated with two system calls:
#include <unistd.h>
char buf[PATH_MAX];
if (chdir("/var/lib/meteora/readings") == -1) /* change the process's cwd */
perror("chdir");
if (getcwd(buf, sizeof buf) == NULL) /* rebuild the current path */
perror("getcwd");
printf("Working in %s\n", buf); /* → /var/lib/meteora/readings */chdir() changes the process's inode pointer — it requires execute permission on the target, as we will see in File Security and Permissions — and getcwd() does the opposite: it starts from the cwd's inode and walks up through .. to the root, searching each parent for the name matching the child's inode, and assembles the string backwards. That is why getcwd() is surprisingly expensive and why it can fail with ENOENT if somebody deleted a directory along the path while you were inside it.
Why the cwd is per process and not global deserves an explicit answer. If it were global, two processes doing cd would step on each other: meteo-api would change directory and the aggregator would write in the wrong place, a textbook race condition (03-01) over shared state. Being per process, it is inherited across fork() — the child starts where the parent was, which is why cd /tmp && ls works as you expect — and it explains the classic bewilderment: a script that does cd /var/lib/meteora does not change your shell's directory, because it runs in a child process whose cwd dies with it. For it to affect the shell you have to run it with source or ..
An important practical detail for services: a daemon that keeps its cwd inside a directory prevents that file system from being unmounted (the "target is busy" of 04-03). That is why well-written daemons do chdir("/") at startup, and why systemd offers WorkingDirectory=.
Path resolution step by step, with the cost counted
This is the central operation of the lesson. When meteo-api runs:
the kernel has to turn 43 characters into inode 1180934. The algorithm, called path resolution, is a loop:
graph TB
A["Path: /var/lib/meteora/readings/2026-08-31.dat"] --> B{"Does it start with /?"}
B -->|Yes| C["current inode = root inode (no. 2)"]
B -->|No| D["current inode = the process's cwd inode"]
C --> E["Take the next component"]
D --> E
E --> F{"Is it a directory?<br/>Is there x permission?"}
F -->|No| G["Error: ENOTDIR or EACCES"]
F -->|Yes| H["Look up the name in its entries"]
H --> I{"Found?"}
I -->|No| J["Error: ENOENT"]
I -->|Yes| K["current inode = the inode found"]
K --> L{"Is it a symbolic link?"}
L -->|Yes| M["Replace it with its target<br/>and restart (max. 40 times)"]
M --> E
L -->|No| N{"Any components left?"}
N -->|Yes| E
N -->|No| O["Return the final inode"]
Run over our path, with the access count assuming empty caches:
| Step | Component | What the kernel does | Disk accesses |
|---|---|---|---|
| 0 | / |
Takes inode 2 (the root is always inode 2 in ext4) | 1 (inode) |
| 1 | var |
Reads the data blocks of /, looks up var → inode 131073 |
1 (data) + 1 (inode) |
| 2 | lib |
Reads the data of /var, looks up lib → inode 262145 |
1 + 1 |
| 3 | meteora |
Reads the data of /var/lib, looks up meteora → inode 1180928 |
1 + 1 |
| 4 | readings |
Reads the data of /var/lib/meteora → inode 1180929 |
1 + 1 |
| 5 | 2026-08-31.dat |
Reads the data of readings, looks up the name → inode 1180934 |
1 + 1 |
| 6 | — | Reads the final inode to check permissions and type | (already counted) |
| Total | 11 accesses |
Eleven disk accesses to open a file. On an NVMe, at 80 µs each, that is about 0.9 ms; on the HDD from 02-05, with 8 ms average latency, it would be 88 milliseconds. And all of that before reading a single byte of data.
Three direct consequences worth internalizing:
- Every path component costs.
/a/b/c/d/e/f/g/file.datis more expensive to open than/data/file.dat. Absurdly deep paths have a real, if small, cost. - Execute permission is checked on every directory along the path, not only on the final file. That is the subject of 04-06, but the mechanism is this loop.
- Opening once and reusing the descriptor is much cheaper than opening the file on every operation. It is the reason the
ingestoropens2026-08-31.datat the start of the day and keeps it open, instead of opening and closing it on every reading received.
If any intermediate component is a symbolic link, it is replaced with its target and the loop restarts. Linux limits chained resolutions to 40; beyond that it returns ELOOP ("Too many levels of symbolic links"), which is the protection against a link that points to itself.
The dentry cache: why that does not cost what it looks like
Eleven accesses per open() would be unacceptable: a server opens thousands of files per second and most of them share the first components of the path. The kernel's solution is the dentry cache (dcache).
A dentry (directory entry) is an in-memory structure associating (parent directory inode, name) → child inode. The dcache is a hash table of dentries, keyed by the partial path. Its effect is devastating: if /var/lib/meteora/readings is cached — and it will be, because it is used constantly — resolving our path drops from 11 disk accesses to one, that of the final inode, and often to zero if that is cached too.
$ sudo slabtop -o | grep -E 'dentry|inode_cache' OBJS ACTIVE USE OBJ SIZE SLABS OBJ/SLAB CACHE SIZE NAME 418320 411927 98% 0.19K 19920 21 79680K dentry 96432 94215 97% 0.58K 3444 28 55104K ext4_inode_cache
On this meteo-01 there are 418,320 cached dentries taking up 78 MiB, with 98 % of them active. That RAM consumption is what makes the file system feel fast, and it is also part of what free counts as "available" memory: they are reclaimable caches, exactly like the pages of 02-04.
Two details you see in practice:
Negative dentries. The cache also stores the "does not exist" answers: if an interpreter looks for a module across a list of twenty paths, the first lookup on each costs disk accesses and the subsequent ones return ENOENT from memory.
It can be flushed for measurement. To compare the real cost with and without cache:
sync # flush pending writes
echo 3 | sudo tee /proc/sys/vm/drop_caches # 3 = dentries + inodes + pages
time cat /var/lib/meteora/readings/2026-08-31.dat > /dev/null # cold
time cat /var/lib/meteora/readings/2026-08-31.dat > /dev/null # warmThe values for drop_caches are 1 (page cache), 2 (dentries and inodes) and 3 (both). Testing only: in production, flushing the caches causes a latency spike while everything is read again. The difference between the cold and the warm run is usually one to two orders of magnitude, and it is the direct measure of the dcache's value.
Hard links: one inode, several names
A hard link is simply another directory entry pointing to the same inode number. It is not a copy nor a special reference: it is one more name, with exactly the same standing as the original.
$ cd /var/lib/meteora/readings $ ls -li 2026-08-31.dat 1180934 -rw-r----- 1 meteora meteora 17280000 Aug 31 23:59 2026-08-31.dat $ ln 2026-08-31.dat /var/lib/meteora/archive/august-final.dat $ ls -li 2026-08-31.dat /var/lib/meteora/archive/august-final.dat 1180934 -rw-r----- 2 meteora meteora 17280000 Aug 31 23:59 2026-08-31.dat 1180934 -rw-r----- 2 meteora meteora 17280000 Aug 31 23:59 .../august-final.dat
Notice the two things that have changed: the same inode 1180934 in both, and the link count has gone from 1 to 2. And what has not changed: the disk space.
$ df -h /var/lib/meteora | tail -1 /dev/md0 196G 41G 146G 22% /var/lib/meteora ← identical to before
The 17 MB have not been duplicated because there are not two files: there is one file with two names. If you write through one, you see it through the other, because it is the same inode, the same blocks and the same bytes.
Now, delete one:
$ rm 2026-08-31.dat $ ls -li /var/lib/meteora/archive/august-final.dat 1180934 -rw-r----- 1 meteora meteora 17280000 Aug 31 23:59 august-final.dat
The file is still there, with the count back to 1. rm destroys nothing: it removes a name and decrements the count. Only when the count reaches 0 does the system free the blocks.
The two restrictions on hard links, with their reasons:
They cannot cross file systems. A hard link is an inode number, and inode numbers are unique only within their file system (04-01). Inode 1180934 exists on /dev/md0 and probably on /dev/sda1 too, being different things. A directory entry has no room to say "from which device":
$ ln /var/lib/meteora/readings/2026-09-01.dat /tmp/test.dat ln: failed to create hard link '/tmp/test.dat' => ...: Invalid cross-device link
The EXDEV error ("Invalid cross-device link") is the same one that forces mv between different disks to copy and delete instead of renaming, with the performance difference that implies: moving 17 GB within the same file system is instantaneous; between two, it takes minutes.
They cannot link directories. It is the prohibition from section 2: it guarantees that the graph is acyclic and that the link count is enough as a garbage collector.
$ sudo ln /var/lib/meteora/readings /var/lib/meteora/shortcut ln: /var/lib/meteora/readings: Operation not permitted
Symbolic links: a file containing a path
A symbolic link is a file of type l whose content is a text string holding a path. Nothing more. When the kernel finds it during resolution, it replaces the link with that string and carries on.
$ ln -s 2026-09-01.dat today.dat $ ls -li today.dat 1180936 lrwxrwxrwx 1 meteora meteora 14 Sep 1 00:00 today.dat -> 2026-09-01.dat
Three observations packed with information: it has its own inode (1180936), so it is a different file; its size is 14, exactly the characters of 2026-09-01.dat, which are its entire content; and its permissions lrwxrwxrwx are always 777 and mean nothing, because the ones that apply are the target's.
A pretty implementation detail: ext4 takes advantage of the 60 bytes the inode reserves for block pointers and, if the path takes fewer than 60 characters, stores it right there. It is called a fast symlink and it consumes no data block at all: resolving it does not cost even one extra access. Almost every real symbolic link fits.
The crucial difference from a hard link is that the symbolic one points to a name, not to an inode. Hence its two characteristic properties:
It can cross file systems and link directories, because a text string has no such limitations:
It can end up broken. If the target disappears or is renamed, the link still exists, pointing at a path that no longer resolves:
$ rm 2026-09-01.dat $ ls -l today.dat lrwxrwxrwx 1 meteora meteora 14 Sep 1 00:00 today.dat -> 2026-09-01.dat $ cat today.dat cat: today.dat: No such file or directory
The link is perfectly healthy; what is missing is the target. Broken links are found with find /var/lib/meteora -xtype l.
And a trap that bites everybody: relative links are resolved relative to the link's directory, not relative to whoever uses it.
$ ln -s 2026-09-01.dat /tmp/today.dat # WRONG! $ cat /tmp/today.dat cat: /tmp/today.dat: No such file or directory
/tmp/today.dat points to 2026-09-01.dat, which the kernel looks for in /tmp/. Practical rule: use absolute paths in links unless link and target are in the same directory or you want the whole set to be movable.
The full comparison table:
| Aspect | Hard link | Symbolic link |
|---|---|---|
| What it is | Another entry with the same inode | A file with a path inside |
| Inode | The same as the original | Its own, and different |
Type in ls -l |
- (indistinguishable) |
l |
| Original's link count | Increases | Unchanged |
| Crosses file systems | No (EXDEV) |
Yes |
| Links directories | No | Yes |
| Survives deletion of the original | Yes (the content remains) | No (left broken) |
| Survives renaming of the original | Yes | No |
| Space used | 24 bytes in the directory | One inode (+0 blocks if < 60 B) |
| Cost at resolution time | None | One extra resolution |
| Which one is the "original"? | Neither: they are identical | The target, clearly |
| Command | ln source new |
ln -s source new |
| Typical use | Incremental backups, deduplication | Versions (today.dat), /usr/bin, libraries |
The "which one is the original?" row is the hardest to accept: after ln a b, there is no way to tell which was created first. Both entries are equally legitimate; the inode records no preference.
Meteora uses both, each where it belongs: a symbolic today.dat that is repointed every midnight at the day's file (clients always ask for today.dat and the link is remade with ln -sfn, which is atomic), and hard links in the daily backups, where files that have not changed are linked instead of copied — that is exactly what rsync --link-dest does and why thirty daily copies of 17 GB can take up 18 GB in total.
Internal organization and the cost of enormous directories
So far we have said that a directory "contains a table". How is that table organized? It is a data-structures decision with measurable effects.
Linear list. The entries, one after another, in the order they were created. Looking up a name means walking from the beginning comparing strings.
- Lookup: O(n). Create: O(n), because you have to check that the name does not already exist.
- It is what ext2 did and what ext4 still does in small directories.
Let us do the arithmetic with a realistic case. Suppose Meteora, instead of one file per day, had stored one file every ten minutes for two years: 105,120 files in readings/. With entries of 24 bytes on average:
- Directory size: 105,120 × 24 ≈ 2.5 MB, that is 616 blocks of 4 KiB.
- Looking up a name: reading an average of 308 blocks and comparing 52,560 strings.
- And worse: creating a new file forces a walk through all 616 blocks to verify that the name does not exist.
The result is the classic pathology: ls takes seconds, and creating the files gets slower and slower as the directory grows. It is quadratic behavior in aggregate: creating n files costs O(n²).
Hash table. A hash function is applied to the name and you jump straight to the corresponding bucket.
- Lookup: O(1) in the average case. It is what ext4 (
dir_index) and NTFS use, with variations. - Drawback: a full traversal returns the entries in hash order, which looks random.
B / B+ tree. The entries are kept sorted in a balanced tree.
- Lookup: O(log n), and moreover the traversal comes out sorted and ranges are efficient.
- It is what XFS, Btrfs and NTFS use.
In ext4, the concrete solution is called an htree (hashed tree) and is enabled by the dir_index feature. It is a hybrid: it applies a variant of MD4 to the name and uses the hash as the key of a one- or two-level tree, each level with 4 KiB of index. With two levels it addresses on the order of millions of entries with at most three accesses: root node, leaf node and data block.
The comparison, over the hypothetical directory of 105,120 files:
| Organization | Look up a name | Create a file | List everything | Output order |
|---|---|---|---|---|
| Linear list | 308 blocks (average) | 616 blocks | 616 blocks | Creation |
| htree (ext4) | 3 blocks | 3 blocks | 616 blocks | Hash (apparently random) |
| B+ tree (XFS) | ~3 blocks | ~3 blocks | 616 blocks | Alphabetical |
From 308 accesses to 3: a hundred times fewer. You check whether it is enabled with:
$ sudo tune2fs -l /dev/md0 | grep features Filesystem features: has_journal ext_attr dir_index extent 64bit metadata_csum
dir_index is there, and it will be on any ext4 created in the last fifteen years. If for some reason it were not, you enable it with tune2fs -O dir_index /dev/md0 followed by e2fsck -D to rebuild the indexes.
Two practical consequences of the htree are worth knowing. The first is that ls comes out unsorted and sorting it costs: the traversal returns the entries in hash order and ls sorts them in memory before printing anything, so with 105,120 entries ls -f or find . -maxdepth 1 are much faster. The second is that listing in hash order destroys inode locality: processing them in the order readdir() returns means reading them in random order, which on an HDD means continuous seeks. The classic solution — and what tar and several backup tools do — is to read all the names, sort them by inode number and process them that way.
And the design conclusion, which is the one that matters: even with an htree, a directory with a hundred thousand files is a bad idea. Listing is still O(n), backups slow down, ls with * can overflow the command line, and any tool that sorts consumes memory. The correct practice is to spread across subdirectories, which is exactly what Meteora does with its historical archive:
/var/lib/meteora/readings/2026-08-31.dat ← the current day and the recent ones /var/lib/meteora/archive/2026/08/2026-08-01.dat ← history by year/month
With that hierarchy, no directory goes beyond 31 entries. It is the same pattern Git uses for its objects (ab/cdef...) and that web caches everywhere use.
The FHS standard and why Meteora is where it is
The Filesystem Hierarchy Standard (FHS) is the convention that says what goes in each directory of the root on Linux. The kernel does not impose it: it is an agreement that lets an administrator know where to look on any distribution.
| Directory | What it contains | Persistent? | Shareable? |
|---|---|---|---|
/bin, /sbin |
Essential system binaries (today, links into /usr) |
Yes | Yes |
/boot |
Kernel, initramfs and boot loader | Yes | No |
/dev |
Device files (devtmpfs, 04-03) | No | No |
/etc |
System configuration, specific to this machine | Yes | No |
/home |
Users' home directories | Yes | Yes |
/lib |
Essential shared libraries | Yes | Yes |
/mnt, /media |
Temporary and removable-media mount points | — | — |
/opt |
Self-contained third-party software | Yes | Yes |
/proc |
Process and kernel information (procfs, 04-03) | No | No |
/root |
The superuser's home directory | Yes | No |
/run |
Runtime state: PIDs, sockets, FIFOs (tmpfs) | No | No |
/srv |
Data served by this system (web, ftp) | Yes | Yes |
/sys |
Interface to the device model (sysfs, 04-03) | No | No |
/tmp |
Any user's temporary files, emptied at boot | No | No |
/usr |
Read-only, non-essential programs and data | Yes | Yes |
/var |
Variable data: logs, queues, caches, databases | Yes | No |
The underlying logic has two axes: variable versus static (/var changes constantly, /usr only when updating) and shareable versus specific (/usr could be mounted over the network and shared between machines, /etc could not).
With that, Meteora's paths stop being arbitrary:
| Meteora path | Why there |
|---|---|
/var/lib/meteora/readings/ |
/var/lib is the canonical place for an application's state data: information the app creates, modifies and needs to keep across restarts. Not /srv because it is not content served as-is, nor /opt because that is for the software, not for its data |
/etc/meteora/meteora.conf |
/etc is configuration specific to this machine, editable by the administrator, which must be kept in the backups and tracked in version control |
/var/log/meteora/meteo-api.log |
/var/log is the standard destination for logs, with rotation managed by logrotate |
/run/meteora/readings.fifo |
/run is tmpfs: it is emptied at boot. Perfect for a FIFO and a socket, which must not survive a reboot: an orphan socket from a previous run would stop the service from starting |
/dev/shm/meteora-cache |
tmpfs for POSIX shared memory (03-03). Volatile by definition: it is a cache |
/tmp/ |
Only for short-lived temporary files, with mkstemp (04-04). Never for data that has to last |
The reason /run and /dev/shm are tmpfs and /var/lib is not sums up the whole chapter: a file's location declares its persistence guarantees. Putting the socket in /var/run (which today is a symbolic link to /run) or the readings in /tmp are not matters of taste: they change what happens after a reboot.
Deletion: what unlink really does
We reach the behavior that baffles people most, and which now has a complete explanation. The system call behind rm is not called delete:
Its exact steps:
- Resolve the path down to the parent directory and check write permission on the directory (not on the file: 04-06).
- Remove the (name, inode) entry from the directory, absorbing its
rec_leninto the previous one. - Decrement the inode's link count.
- If the count reaches 0 and no process has the file open: mark the inode as free in its bitmap and release all its blocks.
- If the count reaches 0 but some process has it open: release nothing. The inode moves to an orphan list and will be released when the last descriptor is closed.
Step 5 is the one that produces the classic incident. Somebody notices that /var/log/meteora/meteo-api.log has grown to 17 GB, deletes it with rm, and the space does not come back:
$ ls -l /var/log/meteora/ total 0 ← the file is gone $ df -h /var/log Filesystem Size Used Avail Use% Mounted on /dev/sda3 20G 19G 340M 99% /var/log ← still full!
The file no longer has a name, but meteo-api has it open. The link count is 0, but the open-descriptor count is 1, so the blocks are still reserved and — worse still — the process keeps writing to a file nobody can open any more.
The diagnosis:
$ sudo lsof +L1 COMMAND PID USER FD TYPE DEVICE SIZE/OFF NLINK NODE NAME meteo-api 2841 meteora 5w REG 8,3 18253611008 0 4457 /var/log/meteora/meteo-api.log (deleted)
+L1 filters the open files with fewer than one link, that is, the deleted ones. The NLINK column reads 0 and the name carries (deleted). There is the culprit.
The three fixes, from best to worst:
# 1. Reload the service: it closes and reopens its files (03-03, SIGHUP)
sudo systemctl reload meteo-api # or kill -HUP 2841
# 2. If it does not support reloading, restart it
sudo systemctl restart meteo-api
# 3. Truncate the file WITHOUT deleting it (through /proc, restarting nothing)
sudo truncate -s 0 /proc/2841/fd/5The third one is an excellent trick: /proc/<pid>/fd/5 is a link to the orphan inode, so you can get at it even though it has no name, and truncating it to zero frees the blocks instantly with the service running. It is the same /proc/<pid>/fd we will explore in File Management.
And the underlying lesson: the right way to empty a log is truncate -s 0 or : > file, never rm. It is exactly what logrotate does with the copytruncate option.
A happy corollary of the same mechanism: deleting an open file is a legitimate and widely used technique. A program can create a temporary file, open it, delete it immediately and go on using it through the descriptor. The file has no name, so nobody else can touch it, and the system frees it by itself when the process dies, even if it dies from a kill -9. We will see it with mkstemp in 04-04.
Common Mistakes and Tips
Believing that rm deletes the file. It deletes a name. If there is another hard link, the content remains; if a process has it open, the space stays occupied. rm is unlink, and the name of the call is literal.
Looking for free space with du when df says it is full. du walks names, so it cannot see deleted files that are still open. If df and du disagree by gigabytes, run lsof +L1.
Using relative paths in symbolic links without thinking. ln -s data.dat /other/place/link creates a link that looks for data.dat in /other/place/. Use absolute paths unless you know exactly what you are doing.
Confusing a directory's link count with "how many files it has". It is 2 + the number of subdirectories; regular files do not count.
Putting a hundred thousand files in one directory. Even though dir_index saves the lookup, listing, sorting, backups and shell wildcards are still O(n) or worse. And a directory does not shrink when you delete: one that reached 2.5 MB will still take just as long to list even if it ends up empty, and the only cure is to recreate it. Spread across subdirectories from the start: migrating later is painful.
Leaving a daemon's cwd inside a volume you want to unmount. It is a common cause of "target is busy" (04-03). Services ought to do chdir("/").
Tip: use ls -f or find -maxdepth 1 in large directories, which avoid the in-memory sort, and stat -c '%i %h %n' to see inode, link count and name at a glance.
Tip: when in doubt whether two paths are the same file, compare st_dev and st_ino, with stat -c '%d %i' or directly find / -samefile path. Comparing contents is slow and comparing names says nothing.
Exercises
Exercise 1: proving that the name is not in the inode
Create a file, make a hard link and a symbolic link to it, and design a sequence of commands that proves the following five statements, showing the output that establishes each one: (a) the hard link and the original share an inode and the symbolic one does not; (b) creating the hard link consumes no data space; (c) deleting the original does not destroy the content if there is a hard link; (d) deleting the original leaves the symbolic link broken; (e) the link count reflects exactly the number of names. Explain in each point which inode field justifies it.
Exercise 2: the cost of path resolution
Calculate the number of disk accesses needed to open /var/lib/meteora/archive/2026/08/2026-08-15.dat with empty caches, detailing the steps. Then, assuming that /var/lib/meteora/archive/2026/08/ is already in the dentry cache, recalculate. With a latency of 80 µs per access on NVMe and 8 ms on HDD, give the four times. Finally, explain what changes if archive is a symbolic link to /mnt/historical/meteora and how many accesses that adds.
Exercise 3: the deleted file that does not free space
Reproduce the whole incident: write a program (in shell or in C) that opens a file, writes to it continuously and does not close it; in another terminal, delete the file and check that df does not go down. Diagnose it with lsof and with du versus df, and fix it without killing the process. Then explain why logrotate with copytruncate exists, and what would have happened if truncate -s 0 had been used instead of rm.
Solutions
Solution 1
cd /tmp && mkdir demo && cd demo
dd if=/dev/urandom of=original.dat bs=1M count=10 status=none
df --output=avail /tmp | tail -1 # free space BEFORE: e.g. 8123456
ln original.dat hard.dat
ln -s original.dat soft.dat
df --output=avail /tmp | tail -1 # free space AFTER: identical
ls -liExpected output:
264531 -rw-r--r-- 2 joan joan 10485760 Sep 1 13:02 hard.dat 264531 -rw-r--r-- 2 joan joan 10485760 Sep 1 13:02 original.dat 264532 lrwxrwxrwx 1 joan joan 12 Sep 1 13:02 soft.dat -> original.dat
(a) original.dat and hard.dat share inode 264531; soft.dat has its own, 264532. The hard link is one more directory entry pointing at the same inode; the symbolic one is a different file whose content is the 12 characters original.dat.
(b) The free space has not changed after ln: a hard link adds 24 bytes to a directory block that was already reserved, without touching the data area. The inode field that justifies it is that the block pointers have not been duplicated: they are still the same ones.
(c) and (e):
rm original.dat
ls -li hard.dat # → 264531 -rw-r--r-- 1 ... hard.dat
md5sum hard.dat # the full content is still thereThe link count went from 2 to 1. unlink decremented the count; since it did not reach 0, the blocks were not freed. The field is i_nlink.
(d):
ls -l soft.dat # the link still exists, with its size of 12
cat soft.dat # cat: soft.dat: No such file or directory
find . -xtype l # ./soft.datThe symbolic link stores the name original.dat, not the inode. Once that name disappears from the directory, resolution fails with ENOENT on the last component. Curiously, ln -s hard.dat soft2.dat would work, because the content is still reachable under that other name: proof that the symbolic link depends on the name and the hard one on the inode.
Solution 2
The path has 6 components: var, lib, meteora, archive, 2026, 08, 2026-08-15.dat. Actually there are 7.
Empty caches. One access for the root's inode, and for each component one to read the directory's data and another to read the inode found:
1 (root inode) + 7 × 2 = 15 accesses.
With /var/lib/meteora/archive/2026/08/ in the dcache, resolution starts directly from the inode of 08/ (which the cache holds), and only the last component remains to be looked up: 1 (data of 08/) + 1 (final inode) = 2 accesses. And if the final inode were cached too, 0.
| Scenario | Accesses | NVMe (80 µs) | HDD (8 ms) |
|---|---|---|---|
| Empty caches | 15 | 1.2 ms | 120 ms |
| Cached dentries | 2 | 0.16 ms | 16 ms |
The 120 ms of the cold HDD are the reason the first start of a service with many files is so slow and the second almost instantaneous, and a numerical justification for the dcache.
With archive as a symbolic link to /mnt/historical/meteora: on reaching archive, the kernel reads its inode, detects type l, and — if the path takes fewer than 60 bytes, as here (24) — reads it from the inode itself, with no extra access. Then it restarts resolution with the new absolute path: root, mnt, historical, meteora, and then 2026, 08 and the file.
Cold count: 1 (root) + 2×3 (var, lib, meteora) + 2 (archive, whose inode already carries the target) + 1 (root again, most likely cached) + 2×3 (mnt, historical, meteora) + 2×3 (2026, 08, file) = ~22 accesses, some 7 more. If the link were 60 bytes or longer, its data block would have to be read as well: one additional access per link traversed. It is little, but it explains why long chains of symbolic links on critical paths are noticeable.
Solution 3
Reproduction:
# Terminal 1
( while :; do dd if=/dev/zero bs=1M count=10 status=none; sleep 1; done ) > /tmp/fat.log &
echo $! # note the PID, e.g. 5512
# Terminal 2
sleep 30
df -h /tmp | tail -1 # usage goes up
rm /tmp/fat.log
ls -l /tmp/fat.log # No such file
df -h /tmp | tail -1 # usage does NOT go down, and keeps rising!
du -sh /tmp # du shows far less than df: the discrepancyDiagnosis:
sudo lsof +L1 /tmp
# COMMAND PID USER FD TYPE DEVICE SIZE/OFF NLINK NODE NAME
# dd 5513 joan 1w REG 0,25 943718400 0 312 /tmp/fat.log (deleted)NLINK 0 and (deleted): inode 312 has no name at all but is still alive because descriptor 1 of process 5513 keeps it open. unlink completed steps 1-3 and stopped at step 5.
Fixing it without killing the process:
/proc/<pid>/fd/1 gives access to the orphan inode through the open descriptor. Truncating it to 0 frees all its blocks instantly. The process keeps writing — to a file that now starts from zero — without having noticed a thing.
Why copytruncate exists. When logrotate renames a log and creates a new one, a service holding the descriptor open will keep writing to the renamed file, because the descriptor points to the inode, not to the name (that is the whole lesson). There are two solutions: send SIGHUP to the service so it reopens the file by name — the elegant one, and what Meteora has done since 03-03 — or use copytruncate, which copies the content and truncates the original to zero without changing the inode, so the service's descriptor stays valid. It is the option for programs that do not know how to reopen their log, and it carries the risk of losing the lines written between the copy and the truncation.
With truncate -s 0 instead of rm there would have been no incident at any point: the name still exists, the inode is still the same, the process's descriptor is still valid, and the blocks are freed immediately. The only quirk is that a process with O_APPEND carries on at the end (which is now 0) while one without O_APPEND keeps its old offset and creates a sparse file (04-04). Hence the rule: to empty a log in use, truncate -s 0 or : > file, never rm.
Conclusion
A directory is a file whose content is a table of (name, inode number) pairs. We have opened one with debugfs and seen ext4's real structure: 8 bytes of header with inode, rec_len, name_len and file_type, plus the name. Out of that format come two behaviors you see every day: directories do not shrink when entries are deleted, and the same inode can appear in several entries, which is the definition of a hard link.
UNIX's acyclic graph organization is the fourth historical attempt, and each earlier one fell for a concrete reason: the single level for name collisions, the two levels for their rigidity, the pure tree for not allowing sharing. The general graph is ruled out by two serious problems — infinite traversals and unreachable garbage that the link count does not detect — and that is why UNIX forbids hard links to directories: a single restriction that guarantees acyclicity and makes the link count a sufficient garbage collector.
Path resolution is a component-by-component loop that starts at the root or at the process's cwd, checks type and permission at every step, and expands symbolic links with a cap of 40 to avoid ELOOP. With empty caches it costs 11 disk accesses for /var/lib/meteora/readings/2026-08-31.dat — 0.9 ms on NVMe, 88 ms on HDD — and with a warm dentry cache it drops to one or to none. Those 78 MiB of dentries on meteo-01 are the reason the file system feels instantaneous.
The two kinds of link differ in one single thing, from which everything else follows: the hard one points to an inode and the symbolic one to a name. That is why the hard one does not cross file systems or link directories, survives deletion and renaming, and has no "original"; and why the symbolic one crosses whatever it likes, breaks easily and costs one extra resolution — although in ext4, if it is under 60 characters, it does not even consume a block.
The cost of enormous directories is a classic data-structures problem: a linear list makes lookup O(n) and the creation of n files O(n²), while ext4's htree (dir_index) resolves in three accesses what used to cost three hundred. But even so, a hundred thousand files in one directory is still a bad idea, and the right answer is to spread across subdirectories, as Meteora's historical archive does in archive/2026/08/.
The FHS explains why each of Meteora's paths is where it is, and its logic boils down to one sentence worth memorizing: a file's location declares its persistence guarantees. /var/lib for data that must last, /etc for this machine's configuration, /var/log for the logs, and /run and /dev/shm — which are tmpfs — for what must disappear on reboot.
And unlink closes the circle of the nameless inode: it removes a directory entry, decrements the link count, and frees the blocks only if the count reaches 0 and nobody has it open. Hence the incident of the 17 GB that are not released, which is diagnosed with lsof +L1 and cured with truncate -s 0 /proc/<pid>/fd/N without restarting anything.
One last link is missing. We have said "the file system of /var/lib/meteora", "that of /run", "that of /proc", as if the single tree we walk with paths were made of different pieces. And it is: / is an ext4 on a partition, /var/lib/meteora is another ext4 on the RAID /dev/md0, /run is tmpfs in RAM, /proc has no device behind it. How are all those pieces glued into a continuous tree that path resolution walks without noticing the change? What exactly happens when you run mount? And how can the same open() work on an SSD, on RAM and on a server at the other end of the network?
That is what we will see in Partitions, Mounting and the Virtual File System.
Operating Systems Fundamentals
Module 1: Introduction to Operating Systems
- Basic Concepts of Operating Systems
- History and Evolution of Operating Systems
- Types of Operating Systems
- Main Functions of an Operating System
- Kernel Architecture: Monolithic, Microkernel and Hybrid
- User Mode, Kernel Mode and System Calls
Module 2: Resource Management
- Process Management
- CPU Scheduling
- Memory Management
- Virtual Memory and Paging
- Storage Management
- Device Management
- Drivers, Interrupts and I/O Operations
Module 3: Concurrency
- Concurrency Concepts
- Threads and Processes
- Inter-Process Communication (IPC)
- Synchronization and Mutual Exclusion
- Classic Concurrency Problems
- Deadlocks: Prevention, Detection and Recovery
Module 4: File Structures
- File Systems
- Directory Structures
- Partitions, Mounting and the Virtual File System
- File Management
- Space Allocation, Journaling and Integrity
- File Security and Permissions
Module 5: System Protection and Security
- Protection Principles and Access Control
- Users, Authentication and Privilege Escalation
- Common Threats and System Hardening
- Auditing, Logging and Incident Response
Module 6: Virtualization and Containers
- Virtualization: Hypervisors and Virtual Machines
- Containers: Namespaces and cgroups
- The Operating System in the Cloud
- Mobile and Real-Time Operating Systems
