Skip to content
academia.sh

Lesson 16 / 22

Block Devices and Partitions

A partition list is not a map: while the device shows 3,702,784 free sectors, the largest contiguous gap is 2,881,536 sectors, and six of the twelve placement decisions drawn from the list turn out wrong.

Contents

The Logs topic separated which indicator a bottleneck should be read from, and reduced I/O wait to a single number. Beneath that number sits a device, and this topic descends there: to the place where bytes are actually written. Storage is the layer in system administration where failure looks quietest; a process’s stopping is visible, a service’s not running is visible, but a disk filling up is usually visible only once a write fails.

The first question is the lowest one. A file system is built not on an entire device but on a part of it; a small table says where that part begins and ends. This lesson reads that table and counts what the table does not say.

Device, Sector, and Partition

A block device is a storage unit read from and written to in fixed-size units. The smallest addressable unit is called a sector; a device is an array of sectors numbered starting from zero, and it has no structure beyond this array. All of the order inside it is read from a data structure sitting in the device’s first sectors.

A sector has two sizes, and the two need not be equal. The logical sector the device presents to the outside and the physical sector actually read and written in one operation internally are separate numbers; in a common setup, the logical sector is 512 bytes, the physical sector 4096 bytes. So the device accepts 512-byte addresses, but when 512 bytes are to be written, it reads a 4096-byte block, modifies it, and writes it back. This distinction is the source of the difference counted in the alignment section, and it sits nowhere in the partition table; it is the device’s own declaration.

That data structure is the partition table. For every partition, the table holds a start sector, an end sector, and a type label. This is all the table holds, and nothing more: the table does not know what is inside the partition, whether it is full or empty, or even whether a file system exists inside it at all. A partition is nothing but a range drawn on the device’s body.

Two table formats are common. The older format is limited to four primary partitions and carries more only by nesting entries inside one partition; the later format allows hundreds of entries, gives every partition an immutable identity, and keeps a copy of the table at the end of the device. The difference between them is not a ranking of superiority but a difference in capacity, and which one is in use is learned by looking at the device itself.

Tools that list partitions only read, and are harmless: lsblk gives the device tree, blkid writes partitions’ identities and type labels, parted dumps a table in units of sectors. The transcript below is a sample transcript showing the shape of parted’s output; it was not run, and no numeric claim in this lesson comes from it.

Model: fictional block device
Disk /dev/sdb: 12582912s
Sector size (logical/physical): 512B/4096B
Partition Table: gpt

Number  Start     End       Size      File system  Name
 1      2048s     1048575s  1046528s  ext4         root
 2      1048576s  3344383s  2295808s  ext4         data
 3      3753984s  4636671s  882688s   ext4         log
 4      4636672s  6383615s  1746944s  linux-swap   swap

What stands out in the transcript is that the gap between where the fourth row ends and where the next partition begins is not written down. The tool lists partitions; it does not list what remains in between.

Why It Is Partitioned

Building a single file system on a device is also an option, and partitioning is the option set against it. The reason for the distinction is not capacity but limiting the spread of failure. On the fictional server, the units that write logs and the units that produce measurement files share the same device; if both write to the same file system, the moment log rotation stalls once, the filled space floods not just the logs but the root file system too, and the system becomes unable to log in. A separate partition keeps the filling within its own boundary.

The second reason is different mount options: one partition can be mounted read-only, another closed to running programs. The third is that a partition’s file system type can be chosen independently of another’s. There is also a cost, and it is measurable: every partition’s free space serves only itself. In the layout above, while the log partition fills up, 1,320,960 sectors can sit empty in the reserved partition, and these two numbers cannot be transferred to each other. Partitioning gives up flexibility to buy isolation; one of the later lessons measures a layer that loosens this trade-off.

A Partition List Is Not a Map

The only way to compute free space from a partition list is subtraction: the sum of the partitions is taken away from the device’s sector count. This operation gives a correct total and produces a wrong diagnosis, because free space is not a single piece.

  • ST1 — The fictional device is 12,582,912 sectors; every measure is in sectors, and no real device’s size is written down.
  • ST2 — Partitions are aligned to a 2048-sector multiple. The first 2048 sectors are reserved for the partition table and cannot be used as a partition.
  • ST3 — The layout is produced from the shared definition’s generator with seed 20260218; six partitions carry the fictional server’s root, data, log, swap, backup, and reserved areas.
  • ST4 — The tool’s output is the partition list; the gap map is not in the output. The diagnosis is built from the list alone.
  • ST5 — A placement decision consists of two questions: does the requested size fit, and how many of that size fit.
"""M03/K03 storage: the fictional block device's partition layout and gap map."""
SEED = 20260218
ALIGN = 2048                 # partitions start at this multiple of sectors
DEVICE = 12582912            # the fictional device's sector count
NAMES = ("root", "data", "log", "swap", "backup", "reserved")


def generator(seed):
    d = seed

    def next_value(n):
        nonlocal d
        d = (d * 1103515245 + 12345) % 2147483648
        return d % n
    return next_value


def build_layout(seed=SEED):
    """Oracle: the partitions' real start and size, in sectors."""
    r = generator(seed)
    parts, cursor = [], ALIGN
    for name in NAMES:
        size = (r(900) + 300) * ALIGN
        parts.append({"name": name, "start": cursor, "size": size})
        cursor += size + r(4) * 100 * ALIGN
    return parts


def gaps(parts, device=DEVICE):
    """Oracle: the CONTIGUOUS free ranges between and after the partitions."""
    g, cursor = [], 0
    for x in sorted(parts, key=lambda z: z["start"]):
        if x["start"] > cursor:
            g.append(x["start"] - cursor)
        cursor = x["start"] + x["size"]
    if cursor < device:
        g.append(device - cursor)
    return g


def diagnose(parts, requests, device=DEVICE):
    """The tool's output is the partition LIST. The diagnosis drawn from the
    list: free space is the device minus the sum of the partitions."""
    g = gaps(parts, device)
    total_free = device - sum(x["size"] for x in parts)
    wrong, rows = 0, []
    for s in requests:
        tool_fits, actual_fits = s <= total_free, s <= max(g)
        tool_count, actual_count = total_free // s, sum(x // s for x in g)
        wrong += (tool_fits != actual_fits) + (tool_count != actual_count)
        rows.append((s, tool_fits, actual_fits, tool_count, actual_count))
    return {"total_free": total_free, "largest": max(g), "gap_count": len(g),
            "rows": rows, "wrong": wrong}


layout = build_layout()
print("partition  start         size        end")
for x in layout:
    print(f"  {x['name']:9s} {x['start']:9d} {x['size']:9d} {x['start'] + x['size']:11d}")
print("used sectors:", sum(x["size"] for x in layout), "| free sectors:",
      DEVICE - sum(x["size"] for x in layout))
print("contiguous gap count:", len(gaps(layout)), "| gaps:", gaps(layout))
print()
REQUESTS = (204800, 614400, 1228800, 2457600, 3072000, 3686400)
print("requested  tool: fits  actual: fits  tool: count  actual: count")
for s, a, ge, aa, ga in diagnose(layout, REQUESTS)["rows"]:
    print(f"{s:7d}  {str(a):11s}  {str(ge):13s}  {aa:10d}  {ga:12d}")
print()
for seed in (SEED, 20260219):
    t2 = diagnose(build_layout(seed), REQUESTS)
    print(f"seed {seed}: free {t2['total_free']:8d}  largest contiguous {t2['largest']:8d}"
          f"  gaps {t2['gap_count']:2d}  wrong diagnosis {t2['wrong']:2d} / 12")
partition  start         size        end
  root           2048   1046528     1048576
  data        1048576   2295808     3344384
  log         3753984    882688     4636672
  swap        4636672   1746944     6383616
  backup      6793216   1587200     8380416
  reserved    8380416   1320960     9701376
used sectors: 8880128 | free sectors: 3702784
contiguous gap count: 4 | gaps: [2048, 409600, 409600, 2881536]

requested  tool: fits  actual: fits  tool: count  actual: count
 204800  True         True                   18            18
 614400  True         True                    6             4
1228800  True         True                    3             2
2457600  True         True                    1             1
3072000  True         False                   1             0
3686400  True         False                   1             0

seed 20260218: free  3702784  largest contiguous  2881536  gaps  4  wrong diagnosis  6 / 12
seed 20260219: free  3633152  largest contiguous  1787904  gaps  7  wrong diagnosis  6 / 12

Three numbers side by side. The oracle: free space is four pieces — 2048, 409,600, 409,600, and 2,881,536 sectors; the largest contiguous range is 2,881,536 sectors. The tool’s output: a partition list, six rows totaling 8,880,128 sectors; the free space drawn from this is 3,702,784 sectors. Wrong diagnosis: six of twelve decisions.

The distribution of the errors is instructive too. Two decisions say “it fits” when it does not: partitions of 3,072,000 and 3,686,400 sectors stay under the total free space but fit into no contiguous range. Four decisions err in count: the tool says six 614,400-sector partitions fit; in reality four fit. The total is correct, the diagnosis is wrong.

In the layout produced with the second seed, free space drops to 3,633,152 sectors, the largest contiguous range to 1,787,904 sectors, and the gap count rises to seven; the wrong diagnosis is again 6. The number staying the same is a coincidence; its staying in the same order of magnitude is not: as the gap count grows, the diagnosis drawn from the total keeps breaking down.

Alignment

Where a partition begins determines the performance of the file system built inside it. A file system works in 4096-byte blocks, and one block is eight sectors; the device’s own physical block is also eight sectors, counted from zero. If a partition begins at a sector that is not a multiple of eight, every block of the file system spans two physical blocks.

  • ST6 — In sequential writing, the number of physical blocks touched increases by only one under misalignment; the blocks are already consecutive.
  • ST7 — In a random single-block write, a misaligned block touches two physical blocks, and because the unchanged portion of both must be preserved, they are read before being written: four operations.
"""Alignment: does the file system block sit on the same boundary as the device block."""
SECTORS_PER_BLOCK = 8        # 4096-byte block, eight 512-byte sectors
BLOCK_COUNT = 512            # file system blocks written
PARTITION_SIZE = 1048576     # same in all six layouts, sectors


def cost(start, blocks=BLOCK_COUNT, s=SECTORS_PER_BLOCK):
    offset = start % s
    return {"aligned": offset == 0,
            "sequential": blocks + (1 if offset else 0),
            "random": blocks * (4 if offset else 1)}


print("start      partition size  aligned  sequential ops  random ops")
wrong = 0
for start in (63, 2048, 2049, 2052, 2056, 4096):
    m = cost(start)
    wrong += not m["aligned"]
    print(f"{start:9d}  {PARTITION_SIZE:12d}  {str(m['aligned']):6s}"
          f"  {m['sequential']:12d}  {m['random']:14d}")
print("the partition size in the tool's output is the same across all six layouts:", PARTITION_SIZE)
print("how many times the 'same layout' diagnosis drawn from that size is wrong:", wrong, "/ 6")
start      partition size  aligned  sequential ops  random ops
       63       1048576  False            513            2048
     2048       1048576  True             512             512
     2049       1048576  False            513            2048
     2052       1048576  False            513            2048
     2056       1048576  True             512             512
     4096       1048576  True             512             512
the partition size in the tool's output is the same across all six layouts: 1048576
how many times the 'same layout' diagnosis drawn from that size is wrong: 3 / 6

The six layouts’ partition size is the same, and this size appears in the tool’s output; the starting sector also appears but is not read, because whether a number divides by eight is not asked while looking at the list. Sequential writing’s cost rises from 512 to 513 under misalignment, an immeasurable difference; random writing’s rises from 512 to 2048. The “no difference” diagnosis given by a sequential test is wrong on three of the six layouts. This is storage’s counterpart to the course’s second claim: what distinguishes is not multiplying the test, but testing the right load.

The row for 2056 in the table shows that alignment is not specific to 2048: every start that is a multiple of eight is aligned. The reason partitioning tools default to 2048 sectors is not eight but staying aligned on devices that work with much larger internal units too; 2048 is a number divisible by all plausible internal unit sizes. Tools do not make this decision silently; they produce a warning too, and lsblk writes the alignment deviation in a separate column. The warning sits in the output; reading it is a matter of habit, and the difference that appears when it is not read shows up nowhere until the moment it is measured.

Modifying the Partition Table

Tools that read the table are harmless; ones that write it are irreversible. Deleting a partition or changing its boundary overwrites the start and end values in the table; the data inside the partition is not deleted, but the information about how to reach it is lost, and the content does not become reachable again until the table is rewritten. This is why commands that write the table are not given in a runnable, complete form in this lesson.

If a change is unavoidable, three precautions provide measurable safety. The first is taking a text dump of the table; a reading command like sfdisk --dump writes the table to a file, and restoring that same file re-establishes the old layout. The second is a dry run: most tools that write the table carry an option that shows the result without applying the change, and the output from this option must be read before it is run for real. The third is doing the trial on a separate device; mounting a file as a loop device makes it possible to practice partitioning away from real data.

There is also an ordering rule: when a command that changes the table returns, the kernel may still be holding the old layout in memory. In that case, lsblk shows the new table while mounted file systems keep working with the old boundaries. A command returning does not mean the change has taken effect; this is storage’s counterpart to the sentence established for signals in the Process Management topic.

Summary

  • A block device is an array of numbered sectors; all of the order inside it is read from the partition table sitting in the first sectors, and the table carries only start, end, and type information.
  • The free space drawn from a partition list is a total, not a map: the largest contiguous piece of 3,702,784 free sectors is 2,881,536 sectors, and the free space is split into four pieces.
  • Six of twelve placement decisions turn out wrong when only the list is consulted; two think something that does not fit does fit, four overcount how many fit. With the second seed, the gap count rises to seven and the wrong diagnosis is again 6.
  • A partition start being a multiple of eight sectors produces a difference of 513 operations instead of 512 in sequential writing, and 2048 instead of 512 in random writing; the “no difference” diagnosis drawn from output showing the same partition size is wrong on three of six layouts.
  • Commands that write the partition table are irreversible; three measurable precautions are taking a dump of the table, reading the dry-run option, and doing the trial on a separate device.

Next Step

This lesson never looked inside a partition: to the table, a partition was nothing but a range with a start and an end. Yet the moment a file system is built on that range, decisions are made that cannot be reversed without cost — how many files can be held, how much space metadata will take up, and whether corrupted data will be noticed when read. The next lesson measures how file system families make these decisions differently, and counts how many times a design that carries no checksum silently returns corrupted data.

To keep your progress and take notes, Log in

My notes

Log in to take notes.

Start typing to search.

↑↓ Esc navigate · open · close