Linux: How does hard-linking to a directory work

hard linkinodelinux

I'm aware that Linux does not allow hard-linking to a directory. I read somewhere,

that this is to prevent unintentional loops (or graphs, instead of the more desirable tree structure) in the file-system.
that some *nix systems do allow the root user to hard-link to directories.

So, if we are on one such system (that does allow hard-linking to a directory) and if we are the root user, then how is the parent directory entry, .., handled following the deletion of the (hard-link's) target and its parent?

a (200)
\-- .  (200)
\-- .. (100)
\-- b  (300)
|   \-- .  (300)
|   \-- .. (200)
|   \-- c  (400)
|       \-- .  (400)
|       \-- .. (300)
|       \-- d  (500)

 <snip>

|
\-- H (400)

(In the above figure, the numbers in the parentheses are the inode addresses.)

If a/H is an (attempted) hard-link to the directory a/b/c, then

What should be the reference count stored in the inode 400: 2, 3, or 4? In other words, does hard-linking to a directory increases the reference count of the target directory's inode by 1 or by 2?
If we delete a/b/c, the . and .. entries in inode 400 continue to point to valid inodes 400 and 300, respectively. But what happens to the reference count stored in inode 400 if the directory tree a/b is recursively deleted?

Even if the inode 400 could be kept intact via a non-zero reference count (of either 1 or 2 – see the preceding question) in it, the inode address corresponding to .. inside inode 400 would still become invalid!

Thus, after the directory tree b stands deleted, if the user changes into the a/H directory and then does a cd .. from there, what is supposed to happen?

Note: If the default file-system on Linux (ext4) does not allow hard-linking to directories even by a root user, then I'd still be interested in knowing the answer to the above question for an inode-based file-system that does allow this feature.

Best Answer

Hard links to directories aren't fundamentally different to hard links for files. In fact, many filesystems do have hard links on directories, but only in a very disciplined way.

In a filesystem that doesn't allow users to create hard links to directories, a directory's links are exactly

the . entry in the directory itself;
the .. entries in all the directories that have this directory as their parent;
one entry in the directory that .. points to.

An additional constraint in such filesystems is that from any directory, following .. nodes must eventually lead to the root. This ensures that the filesystem is presented as a single tree. This constraint is violated on filesystems that allow hard links to directories.

Filesystems that allow hard links to directories allow more cases than the three above. However they maintain the constraint that these cases do exist: a directory's . always exists and points to itself; a directory's .. always points to a directory that has it as an entry. Unlinking a directory entry that is a directory only removes it if it contains no entry other than . and ...

Thus a dangling .. cannot happen. What can go wrong is that a part of the filesystem can become detached. If a directory's .. pointing to one of its descendants, so that ../../../.. eventually forms a loop. (As seen above, filesystems that don't allow hard link manipulations prevent this.) If all the paths from the root to such a directory are unlinked, the part of the filesystem containing this directory cannot be reached anymore, unless there are processes that still have their current directory on it. That part can't even be deleted since there's no way to get at it.

GCFS allows directory hard links and runs a garbage collector to delete such detached parts of the filesystem. You should read its specification, which addresses your concerns in details. This is an interesting intellectual exercise, but I don't know of any filesystem that's used in practice that provides garbage collection.

Related Solutions

Shell – Forcibly create directory hard link(s)

Don't do this. If you want to have a backup system using hard links to save space, better to use rsync with --link-dest, which will hard link files appropriately to save space, without causing the problems that this causes (that is, hard linking between directories is a corruption of the filesystem, and will cause it to report wrong inode counts + fail fsck + generally have unknown semantics due to not being a DAG).

Linux – Why do hard links seem to take the same space as the originals

A file is an inode with meta data among which a list of pointers to where to find the data.

In order to be able to access a file, you have to link it to a directory (think of directories as phone directories, not folders), that is add one or more entries to one of more directories to associate a name with that file.

All those links, those file names point to the same file. There's not one that is the original and the other ones that are links. They are all access points to the same file (same inode) in the directory tree. When you get the size of the file (lstat system call), you're retrieving information (that metadata referred to above) stored in the inode, it doesn't matter which file name, which link you're using to refer to that file.

By contrast symlinks are another file (another inode) whose content is a path to the target file. Like any other file, those symlinks have to be linked to a directory (must have a name) so you can access them. You can also have several links to a symlinks, or in other words, symlinks can be given several names (in one or more directories).

$ touch a
$ ln a b
$ ln -s a c
$ ln c d
$ ls -li [a-d]
10486707 -rw-r--r-- 2 stephane stephane 0 Aug 27 17:05 a
10486707 -rw-r--r-- 2 stephane stephane 0 Aug 27 17:05 b
10502404 lrwxrwxrwx 2 stephane stephane 1 Aug 27 17:05 c -> a
10502404 lrwxrwxrwx 2 stephane stephane 1 Aug 27 17:05 d -> a

Above the file number 10486707 is a regular file. Two entries in the current directory (one with name a, one with name b) link to it. Because the link count is 2, we know there's no other name of that file in the current directory or any other directory. File number 10502404 is another file, this time of type symlink linked twice to the current directory. Its content (target) is the relative path "a".

Note that if 10502404 was linked to another directory than the current one, it would typically point to a different file depending on how it was accessed.

$ mkdir 1 2
$ echo foo > 1/a
$ echo bar > 2/a
$ ln -s a 1/b
$ ln 1/b 2/b
$ ls -lia 1 2
1:
total 92
10608644 drwxr-xr-x   2 stephane stephane  4096 Aug 27 17:26 ./
10485761 drwxrwxr-x 443 stephane stephane 81920 Aug 27 17:26 ../
10504186 -rw-r--r--   1 stephane stephane     4 Aug 27 17:24 a
10539259 lrwxrwxrwx   2 stephane stephane     1 Aug 27 17:26 b -> a

2:
total 92
10608674 drwxr-xr-x   2 stephane stephane  4096 Aug 27 17:26 ./
10485761 drwxrwxr-x 443 stephane stephane 81920 Aug 27 17:26 ../
10539044 -rw-r--r--   1 stephane stephane     4 Aug 27 17:24 a
10539259 lrwxrwxrwx   2 stephane stephane     1 Aug 27 17:26 b -> a
$ cat 1/b
foo
$ cat 2/b
bar

Files have no names associated with them other than in the directories that link them. The space taken by their names is the entries in those directories, it's accounted for in the file size/disk usage of the directories.

You'll notice that the system call to remove a file is unlink. That is, you don't remove files, you unlink them from the directories they're referenced in. Once unlinked from the last directory that had an entry to a given file, that file is then destroyed (as long as no process has it opened).

Best Answer

Related Solutions

Shell – Forcibly create directory hard link(s)

Linux – Why do hard links seem to take the same space as the originals

Related Question