Skip to content
academia.sh

Lesson 12 / 13

Firewall Rules

The default-deny policy's five steps are applied in a hundred twenty orders; forty produce a full lockout, twenty a partial one, and because the final rule set is the same across all hundred twenty orders, a dry run prevents zero of them; on writing errors a dry run eliminates five candidates and an end-to-end attempt eliminates five.

Contents

The previous lesson asked at which hop the decision was made and chose the policy as the oracle: because no rule matched, the chain’s default kicked in. That default is not an accident, it is a choice. This lesson takes up that choice and the work of setting it up.

A default-deny policy is a single line, and when that line is written produces more consequences than what it writes. What this lesson measures is not the content of the rules but the order they are applied in; and choosing that order wrong can cut the one path between the administrator and the machine.

Default Deny and Legitimate Flows

Default deny means setting a chain’s policy to “refuse everything that does not match.” Its opposite is default allow: everything that does not match passes, and you have to write down, one by one, every case you want to forbid. The difference between the two is not a difference in rule count, it is a difference in error direction. Under default allow, a forgotten rule produces an opening; under default deny, a forgotten rule produces an outage. An outage is seen; an opening is not.

There are legitimate flows that must be kept alive while default deny is being set up, and in this setup there are four. The loopback flow is the machine’s own processes talking to one another. The established connection flow is the return packets of sessions already open. The management flow is the administrator’s way of reaching this machine. The service flow is what the machine offers outward. Each is kept alive by one rule.

There are five steps, four of them allow rules and one the policy. The dump below is not executed and is not a runnable command sequence; it is written only to show the names of the steps and the difference between two orderings:

# example dump, not a runnable sequence
# safe order: the deny policy is left for LAST
1 loopback-allow
2 established-connection-allow
3 management-allow
4 service-allow
5 default-deny

# locking order: deny before management and established-connection allows
1 default-deny           <- the remote session drops at this step
2 ...                    <- the remaining steps cannot be applied

Lockout is the administrator cutting off their own access. This word is the same as the deadlock counterpart in the Operating System Concepts course, but the event is not the same: there, two processes wait on each other and neither can proceed; here, there is a single administrator who has closed their own channel. What they share is only that the result looks irreversible.

A Hundred Twenty Orders

The setup is this: five steps are going to be applied on a machine, and the administrator is connected remotely. The order of the steps is free; every order is counted exhaustively.

  • EF31 — The five steps’ 120 orders are produced by exhaustive count; there is no sampling and no seed.
  • EF32 — The administrator’s session stays alive when the deny policy is applied if at least one of the management allow or established-connection allow rules is already in place.
  • EF33 — Once the session drops, the remaining steps cannot be applied; the order is cut off there.
  • EF34 — If deny is applied with the established-connection allow granted but management allow not granted, the current session survives, but no new connection can be established; this case is called a partial lockout.
  • EF35 — The rollback timer is set to 120 seconds; when the time elapses, the previous rule set is restored.
  • EF36 — There are seven candidates in the rule-writing measurement: six writing errors and the error-free case.
  • EF37 — A dry run produces the rule set that would be applied, without applying it, and sees structural errors; because it does not know intent, it cannot see a misspelled port.
  • EF38 — An end-to-end attempt probes all four legitimate flows; for this reason it sees all six errors, but it requires the rule set to have already been applied.
  • EF39 — The candidate set sweep is done with six-, seven-, and eight-candidate lists.
  • EF40 — All counts are exhaustive; there is no randomness or seed.
"""Default deny: application order and rule-writing errors.

Part 1 - all 120 orders of the five steps are counted exhaustively; how many orders produce a lockout.
Part 2 - seven candidates (six writing errors and the error-free case), four readings.
"""
from itertools import permutations

STEP = ("loopback-allow", "established-connection-allow", "management-allow",
        "service-allow", "default-deny")
ROLLBACK = 120


def apply(order):
    """If neither management nor established-connection allow is in place when the
    deny policy is applied, the remote session drops; the remaining steps cannot be applied."""
    present = set()
    for i, a in enumerate(order):
        present.add(a)
        if a == "default-deny":
            if not ({"management-allow", "established-connection-allow"} & present):
                return "full-lockout", i + 1, len(order) - (i + 1)
            if "management-allow" not in present:
                return "partial-lockout", i + 1, 0
    return "safe", len(order), 0


ORDERS = list(permutations(STEP))
RESULT = [apply(s) for s in ORDERS]
COUNT = {}
for kind, _, _ in RESULT:
    COUNT[kind] = COUNT.get(kind, 0) + 1
print("all orders of five steps:", len(ORDERS))
for kind in ("safe", "partial-lockout", "full-lockout"):
    print(f"  {kind:18s} {COUNT.get(kind, 0):3d} / {len(ORDERS)}"
          f"  ratio {COUNT.get(kind, 0) / len(ORDERS):.4f}")
print("  total unapplied steps under full lockout:",
      sum(u for t, _, u in RESULT if t == "full-lockout"))
print("  in how many distinct forms does the final rule set end up:",
      len({frozenset(s) for s in ORDERS}))
print()
LOCKOUT_TOTAL = COUNT["full-lockout"]
print("application method       lockouts  external intervention  downtime (s)"
      "  prevented by dry run")
for ad, lockouts, intervention, downtime, dry_run in (
        ("plain application", LOCKOUT_TOTAL, LOCKOUT_TOTAL, 0, 0),
        ("application after dry run", LOCKOUT_TOTAL, LOCKOUT_TOTAL, 0, 0),
        ("timer-based rollback", LOCKOUT_TOTAL, 0, LOCKOUT_TOTAL * ROLLBACK, 0),
        ("separate management channel", LOCKOUT_TOTAL, 0, 0, 0),
        ("atomic load", 0, 0, 0, 0)):
    print(f"  {ad:23s} {lockouts:10d} {intervention:16d} {downtime:13d} {dry_run:27d}")
print()
FLAW = {
    "syntax-error": (True, True, "management-flow-cut"),
    "rule-in-wrong-chain": (False, True, "management-flow-cut"),
    "rule-order-shadows": (False, True, "service-flow-cut"),
    "state-rule-missing": (False, True, "return-path-cut"),
    "wrong-interface-name": (False, True, "management-flow-cut"),
    "wrong-port": (False, False, "service-flow-cut"),
    "no-error": (False, False, "everything-passed"),
}
ORACLE = "wrong-port"
TEST = {
    "syntax-check": lambda a: "error" if FLAW[a][0] else "clean",
    "rule-list": lambda a: ("no-rule" if a == "syntax-error"
                                else "different-chain"
                                if a == "rule-in-wrong-chain" else "in-place"),
    "dry-run": lambda a: a if FLAW[a][1] else "clean",
    "end-to-end-test": lambda a: FLAW[a][2],
}
RISK = {"syntax-check": 0, "rule-list": 0, "dry-run": 0,
        "end-to-end-test": sum(1 for a in FLAW
                            if FLAW[a][2] == "management-flow-cut")}
A = tuple(FLAW)


def eliminate(candidates, s, answer):
    return tuple(a for a in candidates if TEST[s](a) == answer)


def group(tests, candidates):
    o = {}
    for a in candidates:
        o.setdefault(tuple(TEST[s](a) for s in tests), []).append(a)
    return [sorted(v) for v in o.values() if len(v) > 1]


print("candidates:", len(A), "| oracle:", ORACLE)
print()
print("reading              distinct answers  eliminated  remaining  risk of cutting management")
for s in TEST:
    remaining = len(eliminate(A, s, TEST[s](ORACLE)))
    print(f"  {s:18s} {len({TEST[s](a) for a in A}):11d} {len(A) - remaining:7d}"
          f" {remaining:6d} {RISK[s]:20d}")
print()
print("reading set                        indistinguishable group")
for ad, k in (("syntax check alone", ["syntax-check"]),
              ("syntax + rule list",
               ["syntax-check", "rule-list"]),
              ("dry run", ["dry-run"]),
              ("end-to-end test", ["end-to-end-test"]),
              ("dry run + end-to-end test",
               ["dry-run", "end-to-end-test"])):
    o = group(k, A)
    print(f"  {ad:34s} {len(o):8d}", o if o else "")
print()
print("candidate set sweep")
FLAW["wrong-source-range"] = (False, False, "service-flow-cut")
for ad, k in (("6 candidates: state-rule-missing removed",
               tuple(a for a in A if a != "state-rule-missing")),
              ("7 candidates: base list", A),
              ("8 candidates: wrong-source-range added",
               A + ("wrong-source-range",))):
    remaining = len(eliminate(k, "dry-run", TEST["dry-run"](ORACLE)))
    o = group(["dry-run", "end-to-end-test"], k)
    print(f"  {ad:38s} dry run eliminated {len(k) - remaining:2d} remaining {remaining:2d}"
          f" | group the two readings together cannot separate {len(o)}"
          f" largest {max((len(x) for x in o), default=0)}")
all orders of five steps: 120
  safe                60 / 120  ratio 0.5000
  partial-lockout     20 / 120  ratio 0.1667
  full-lockout        40 / 120  ratio 0.3333
  total unapplied steps under full lockout: 140
  in how many distinct forms does the final rule set end up: 1

application method       lockouts  external intervention  downtime (s)  prevented by dry run
  plain application               40               40             0                           0
  application after dry run         40               40             0                           0
  timer-based rollback            40                0          4800                           0
  separate management channel         40                0             0                           0
  atomic load                      0                0             0                           0

candidates: 7 | oracle: wrong-port

reading              distinct answers  eliminated  remaining  risk of cutting management
  syntax-check                 2       1      6                    0
  rule-list                    3       2      5                    0
  dry-run                      6       5      2                    0
  end-to-end-test              4       5      2                    3

reading set                        indistinguishable group
  syntax check alone                        1 [['no-error', 'rule-in-wrong-chain', 'rule-order-shadows', 'state-rule-missing', 'wrong-interface-name', 'wrong-port']]
  syntax + rule list                        1 [['no-error', 'rule-order-shadows', 'state-rule-missing', 'wrong-interface-name', 'wrong-port']]
  dry run                                   1 [['no-error', 'wrong-port']]
  end-to-end test                           2 [['rule-in-wrong-chain', 'syntax-error', 'wrong-interface-name'], ['rule-order-shadows', 'wrong-port']]
  dry run + end-to-end test                 0 

candidate set sweep
  6 candidates: state-rule-missing removed dry run eliminated  4 remaining  2 | group the two readings together cannot separate 0 largest 0
  7 candidates: base list                dry run eliminated  5 remaining  2 | group the two readings together cannot separate 0 largest 0
  8 candidates: wrong-source-range added dry run eliminated  5 remaining  3 | group the two readings together cannot separate 1 largest 2

The Cost of Order

The top table splits the hundred twenty orders into three groups. 60 orders are safe, 20 produce a partial lockout, 40 produce a full lockout. A third. When no preference is made between the deny policy, management allow, and established-connection allow, which of these three steps comes first is equally likely, and deny coming first is a lockout.

Partial lockout’s twenty orders are the most treacherous group. Established-connection allow has been granted, deny has been applied, management allow has not yet been written. The administrator’s session is alive; commands run, output comes back. There is no symptom at all. When that session closes for any reason — and the reason can be a network interruption, closing the terminal, or a timeout — there is no way back. A session staying alive is not proof that access continues.

Full lockout’s forty orders carry one more cost: the total of unapplied steps is 140. When the order is cut off, the rule set is left half-finished, and a half-finished set usually shuts out both the administrator and the service at once. The fault is not a single access fault; it is a service outage at the same time.

The top table’s last row is this lesson’s most important finding: the final rule set is the same across all hundred twenty orders — a single form. The set of applied steps does not change, only their order does. The direct consequence of this is that a dry run, which produces and examines the final rule set without applying it, sees none of these lockouts. This is why the second table’s last column is zero from top to bottom. The danger is not in the final state, it is in the intermediate states.

Paths That Prevent Lockout

The second table places four paths side by side by their cost.

Plain application means forty external interventions for forty lockouts; each intervention means console-level access or physically going to the machine. Application after a dry run gives the same numbers; a dry run is worthless for this class of fault. Timer-based rollback does not prevent the forty lockouts, but it brings external intervention down to zero: when the hundred-twenty-second timer elapses, the previous rule set is restored and the administrator gets back in. Its cost is a total of 4800 seconds of downtime. A separate management channel also does not prevent the lockout, brings intervention down to zero, and produces no downtime; its cost is that channel having to exist at all times. Atomic load changes the rule set in a single operation; because no intermediate state is created, the lockout count becomes 0.

The ranking of the four paths follows from this table. The path that produces no intermediate state is best, because it removes the problem at its source. The path that keeps the intermediate state it produces short comes second. The path that makes the intermediate state tolerable comes third. The path that only examines the intermediate state does nothing at all for this class.

A lockout also carries no evidence of its own, and this is this lesson’s form of the course’s third claim. The moment the session drops, all the administrator has left is a frozen screen on the terminal; the line that would say what happened on the machine is inside the machine, and the machine cannot be entered. The trace of a dropped management packet is also zero lines, exactly as counted in the previous lesson, unless recording was already on. This is why the timer’s duration is chosen not for reading a log and deciding, but for coming back.

Three Numbers

The bottom measurement looks at rule writing: the set is applied in the correct order, but the written rule itself can have an error. Oracle: of seven candidates, the real one is wrong-port — the rule is written, its syntax is correct, its chain is correct, its order is correct, and the port inside it is wrong. Test: the syntax check says “clean,” the rule list shows the rule as “in place,” the dry run says “clean,” the end-to-end attempt says “service flow cut.” Candidates eliminated: syntax check 1, rule list 2, dry run 5, end-to-end attempt 5.

The last two readings eliminate the same number of candidates, and the two-candidate residue they leave behind is different. The pair the dry run leaves is no-error and wrong-port; that is, when the dry run says “clean,” it cannot distinguish between being genuinely clean and carrying an error it cannot see. The pair the end-to-end attempt leaves is rule-order-shadows and wrong-port; both cut the same flow. When the two readings are used together, the indistinguishable group count is 0.

The column on the right says why both are needed. The end-to-end attempt’s risk of cutting the management flow is 3: in three of seven candidates, the set that has to be applied for the attempt to be possible at all closes the administrator’s own path. The dry run’s risk is 0. The correct order follows from this: the risk-free reading first, then the risky attempt under the protection of a timer or a separate channel. A dry run does not replace the attempt; it makes the attempt cheaper, because it eliminates, before any attempt is made, the five errors it is able to catch.

Candidate Set Sweep

The bottom table again tries three sizes. When state-rule-missing is removed, the number of candidates the dry run eliminates drops from five to four, and what remains is still two. When wrong-source-range is added to the list, the eliminated count stays at 5, what remains rises to three, and even when the two readings are used together, an unseparated group is born: the new candidate also escapes the dry run and cuts the same flow.

The pattern stayed the same throughout the topic. As the candidate list grows, the elimination count stays fixed, and the remaining ambiguity grows. Every time an error class the dry run cannot see is added to the list, the meaning of a “clean” answer weakens by one more candidate.

Summary

  • The default-deny policy needs four allow rules keeping four legitimate flows alive: loopback, established connection, management, and service.
  • Of the five steps’ 120 orders, 60 are safe, 20 produce a partial lockout, 40 a full lockout; the total of unapplied steps under a full lockout is 140.
  • Under a partial lockout, the current session survives and gives no symptom at all; a session staying alive is not proof that access continues.
  • The final rule set is the same across all hundred twenty orders; for this reason a dry run prevents 0 of the order-caused lockouts. Timer-based rollback and a separate management channel zero out external intervention, atomic load zeros out the lockout itself.
  • On rule-writing errors, a dry run eliminates 5 candidates and an end-to-end attempt eliminates 5, and what they leave behind differs; together they bring the indistinguishable group down to 0. The end-to-end attempt’s risk of cutting management is 3, the dry run’s is 0.

Next Step

This lesson counted how a machine closes its own gate. The next lesson looks in the opposite direction: when the same machine sits at the exit of a network, it does address translation for the machines behind it and forwards incoming connections inward. Every forwarding rule opens a window next to the door that was closed. The course’s final lesson measures the number of those windows, the damage they do to diagnosis, and how long a translation record stays alive.

To keep your progress and take notes, Log in

My notes

Log in to take notes.

Start typing to search.

↑↓ Esc navigate · open · close