Skip to content
academia.sh

Lesson 03 / 20

The Directory Hierarchy

The single-rooted file tree, mount points, the standard directory layout, and step-by-step path resolution; building the project tree used throughout the course.

Contents

Commands operate on files, and files do not sit in random places. On Unix systems, every file’s address is a path starting from the root of a single tree, and every branch of that tree has a specific meaning.

This lesson establishes the logic of that layout, explains how the kernel resolves a path, and builds the project/ tree the course will work on throughout.

A Single Root

The file system is a single tree, and its root is shown as /. There are no drive letters; a second disk does not open a second tree. Instead, disks are mounted onto a branch of the tree.

Mounting makes a storage device’s contents visible over an existing directory. Once mounted, a program entering that directory does not notice it is reaching a different device; the path is still part of the same tree. Where the device boundary lies matters only when an operation crosses devices — the Hard Links and Move Operations lesson will show what this boundary prevents.

The root has no level above it. In the root directory, .. still points to the root:

$ cd /
$ pwd
/
$ cd ..
$ pwd
/

This is a definition, not a bug: there is nothing above the top of the tree, so .. refers back to itself. This is what keeps path resolution from ever leaving the root.

The Standard Layout

The names and functions of the directories under the root are set by a layout standard. The standard says what kind of file goes where, so an administrator knows where to look even on an unfamiliar system.

Directory Contents
/bin, /usr/bin Basic commands run by all users
/sbin, /usr/sbin System administration commands
/lib, /usr/lib Shared libraries
/etc System-specific configuration files
/home User home directories
/root The superuser’s home directory
/var Data of changing size: logs, queues, caches
/tmp Temporary files; may be cleared on reboot
/opt Software installed outside the package manager
/dev Device nodes
/proc, /sys Virtual file systems presenting kernel state as files
/boot Kernel image and bootloader files
/mnt, /media Temporary mount points

Two distinctions sit behind the layout. The first is static versus variable: content under /usr changes only when software is installed, while content under /var changes continuously while the system runs. The second is shareable versus local: /usr can be shared across multiple machines, while /etc is specific to that machine. These two distinctions make it possible to place directories on separate partitions and apply different backup policies.

A standard layout merges some branches that were split for historical reasons back into a single point:

$ ls -l /bin /lib
lrwxrwxrwx 1 root root 7 Apr 22  2024 /bin -> usr/bin
lrwxrwxrwx 1 root root 7 Apr 22  2024 /lib -> usr/lib

On this system, /bin is not a separate directory but a link pointing to /usr/bin. /bin/ls and /usr/bin/ls are the same file. Old paths continuing to work is for backward compatibility.

Absolute and Relative Paths

An absolute path starts from the root and is a string beginning with /: /home/student/project/data. It points to the same file no matter where you are.

A relative path starts from the current directory and does not begin with /: data/raw. Which file it points to depends on the current working directory.

The working directory is a property of the process, not the shell; every process has its own working directory, and child processes inherit it. The pwd command prints this value.

Three special names are valid in every directory:

  • . — the current directory itself
  • .. — the parent directory
  • ~ — the home directory; this is not a directory name but an expansion the shell performs

. and .. are real directory entries; the kernel keeps them in every directory. ~ belongs to the shell and is converted to the home directory’s absolute path before it reaches the command.

How a Path Is Resolved

A path is resolved component by component by the kernel. For the path /home/student/project, the order is: the root directory is opened, home is looked up inside it, the directory found is opened, student is looked up inside it, that directory is opened, project is looked up inside it.

This step-by-step walk has three consequences, and all three come up again in later lessons of the course:

Every step requires a lookup. The deeper the path, the more directories are read.

Every intermediate directory is checked for traverse permission. Reaching the file at the end of the path requires execute permission on every directory along the way. The Permission Bits lesson will explain why the x bit means “the right to pass through” for directories.

Intermediate components can be links. If a symbolic link is encountered along the path, resolution continues from its target. This is why the path /bin/ls ends up at the file /usr/bin/ls.

A relative path is resolved the same way; the only difference is that it starts from the working directory, not the root:

$ cd /home/student/project/data
$ pwd
/home/student/project/data
$ cd -
/
$ pwd
/

cd - returns to the previous directory and prints the directory it returns to. This is a shortcut used when switching frequently between two directories.

Hidden Files

Files whose name starts with a dot are not shown by default in a listing:

$ ls
project
$ ls -a
.  ..  .bash_history  .bash_logout  .bashrc  .profile  project

This is not a security feature but a display convention. As far as the kernel is concerned, a name starting with a dot has no special status; only listing tools skip them by default. The convention keeps user configuration files from cluttering the home directory.

The first two entries in ls -a’s output are . and ..; here they appear as the real entries every directory has.

Building the Project Tree

The tree used throughout the course is built under the home directory. Measurement data, processed output, scripts, documentation, and archives are placed in separate branches — a small-scale repetition of the layout under the root:

$ mkdir -p project/data/raw project/data/processed
$ mkdir -p project/scripts project/docs project/archive
$ find project -type d | sort
project
project/archive
project/data
project/data/processed
project/data/raw
project/docs
project/scripts

mkdir normally creates only a single level of directory and errors if the parent does not exist. The -p option also creates missing intermediate directories and does not error if the directory already exists. The second property lets the same command be rerun safely.

The find command walks a tree recursively; -type d selects only directories. The output is not sorted — the file system can return entries in no particular order — so it is piped through sort.

The tree’s meaning is this: data/raw holds measurement files that come from outside and will not be modified; data/processed holds output derived from them. scripts holds the programs that do the processing, docs holds the descriptions, archive holds packaged versions. Keeping raw data separate and untouched will sit at the center of the irreversibility discussion in the deletion lesson.

Summary

  • The file system has a single root; extra disks do not open a new tree, they mount onto an existing directory.
  • The layout under the root is standardized and rests on two distinctions: static/variable and shareable/local.
  • An absolute path starts from the root, a relative path from the working directory; the working directory is a property of the process.
  • A path is resolved component by component; every intermediate directory is checked for traverse permission, and intermediate components can be links.
  • The hiddenness of names starting with a dot is a listing convention, not a kernel-level privilege.
  • mkdir -p creates missing intermediate directories and does not error on an existing directory.

Next Step

The tree is built, but a command’s options and behavior still have to be learned by rote for now. The system itself makes that unnecessary: every command’s own documentation is installed with it, and the documentation follows a consistent structure divided into numbered sections. The next lesson shows how to read this documentation, and why the same name can describe different things in different sections.

To keep your progress and take notes, Log in

My notes

Log in to take notes.

Start typing to search.

↑↓ Esc navigate · open · close