Skip to content
academia.sh

Lesson 14 / 19

Circuit Breaker

Measuring where the breaker sits in the code: whether its state is kept per dependency or per call site, the calls reaching the dependency before each placement opens, the same threshold carrying two separate meanings under the two placements, and a rarely called path whose breaker never opens.

Contents

The previous lesson measured how a retry worsens a persistent failure: all 27 of the 27 calls failed, and a dependency that was already down took twenty-seven times the load. What’s missing is not a number but a decision — recognizing that the dependency is down and not making the call at all.

What a circuit breaker is, and its closed, open, and half-open states, were built in the M14/K06 Application Architecture: Routing, State and Data course; the pattern’s effect on the system, the cost of a failure-free day, and the price of a false trip were measured in M19/K05 Resilience and Reliability. Neither repeats here. This lesson’s question is: where is the breaker’s state kept. A breaker is nothing more than a number and a timestamp; where that number is placed decides what the pattern does from the start.

There are two placements, and the difference between them is a single symbol in the code: is state stored under the dependency’s name, or under the call site’s name. The loan service reaches the same notification dependency from three separate call sites; in the first placement, all three share one state, and in the second, each keeps its own.

Mechanism

The notification dependency is permanently faulty and returns 503 on every call. Catalog is healthy and, although it passes through the breaker too, it never opens — it is there to show where the breaker’s scope ends. RS7 — the open duration outlasts the run, so once a breaker opens it does not close, and the half-open transition does not interfere with the measurement.

// services.mjs — 8302 notification (permanently faulty), 8303 catalog (healthy), 8301 loan.
// The breaker is the same code under both placements; the only change is which key holds the state.
import { createServer, request } from "node:http";
const s = { catalog: 0, send: 0, remind: 0, cancel: 0 };
createServer((req, res) => {                        // 8302 notification: 503 on every call
  if (req.url === "/counter") return res.end(JSON.stringify(s));
  if (req.url === "/reset") { for (const k in s) s[k] = 0; return res.end("{}"); }
  s[req.url.slice(1)] += 1; res.statusCode = 503; res.end("{}");
}).listen(8302);
createServer((req, res) => { s.catalog += 1; res.end("{}"); }).listen(8303);

const call = (port, path) => new Promise((resolve) => {
  const r = request({ port, path, agent: false }, (y) => { y.resume(); y.on("end", () => resolve(y.statusCode)); });
  r.on("error", () => resolve(0)); r.end();
});
let THRESHOLD = 5, state = new Map();               // RS7: open duration outlasts the run; half-open transition built in M14/K06
const OPEN_DURATION = 3_600_000;
const breakerFor = (a) => { if (!state.has(a)) state.set(a, { consecutive: 0, openUntil: 0 }); return state.get(a); };
const guardedCall = async (port, path, key) => {    // breaker lives here; the key decides placement
  const d = breakerFor(key);
  if (Date.now() < d.openUntil) return 0;           // open: the call is never made
  const k = await call(port, path);
  if (k >= 500) { d.consecutive += 1; if (d.consecutive >= THRESHOLD) d.openUntil = Date.now() + OPEN_DURATION; } else d.consecutive = 0;
  return k;
};
createServer(async (req, res) => {                  // 8301 loan: three call sites reach the same dependency
  const q = new URL(req.url, "http://x").searchParams;
  if (req.url.startsWith("/configure")) { THRESHOLD = Number(q.get("threshold")); state = new Map(); return res.end("{}"); }
  if (req.url.startsWith("/state"))
    return res.end(JSON.stringify([...state].map(([a, d]) => [a, Date.now() < d.openUntil])));
  const n = Number(q.get("n")), placement = q.get("by");
  const keyFor = (dependency, callSite) => (placement === "dependency" ? dependency : callSite);
  await guardedCall(8303, "/book", keyFor("catalog", "book"));
  await guardedCall(8302, "/send", keyFor("notification", "send"));
  await guardedCall(8302, "/remind", keyFor("notification", "remind"));
  if (n % 20 === 0) await guardedCall(8302, "/cancel", keyFor("notification", "cancel"));
  res.end("{}");
}).listen(8301);

The keyFor function is this entire lesson. The breaker’s code does not change, the threshold does not change, the call sites do not change; the only thing that changes is whether the key passed into guardedCall names the dependency or the call site. Two of the three call sites run on every request; the cancel path is called only every twentieth request — in a real codebase, frequent and rare paths always coexist.

// measure.mjs — where the breaker's state is kept: per dependency or per call site.
import { request } from "node:http";
const N = 60;                                       // 60 user requests; the cancel path is called every 20th
const call = (port, path) => new Promise((resolve) => {
  const r = request({ port, path, agent: false }, (y) => {
    let g = ""; y.on("data", (p) => (g += p)); y.on("end", () => resolve(g));
  });
  r.on("error", () => resolve("")); r.end();
});
const print = (a, ...r) => console.log(String(a).padEnd(24) + r.map((x) => String(x).padStart(20)).join(""));

const run = async (threshold, by) => {
  await call(8302, "/reset");
  await call(8301, `/configure?threshold=${threshold}`);
  for (let n = 1; n <= N; n += 1) await call(8301, `/request?n=${n}&by=${by}`);
  const s = JSON.parse(await call(8302, "/counter"));
  const d = JSON.parse(await call(8301, "/state"));
  return { ...s, objects: d.length, open: d.filter(([, a]) => a).length };
};

console.log(`${N} user requests; send and remind on every request, cancel every 20th request`);
console.log("notification dependency is permanently faulty: 503 on every call");
print("threshold / placement", "state objects", "open objects", "notification calls", "send", "remind", "cancel", "catalog");
const t = {};
for (const threshold of [3, 5, 10])
  for (const [by, name] of [["dependency", "dependency"], ["callSite", "call site"]]) {
    const r = await run(threshold, by);
    t[`${threshold}${by}`] = r;
    print(`${threshold} / ${name}`, r.objects, r.open, r.send + r.remind + r.cancel, r.send, r.remind, r.cancel, r.catalog);
  }
const total = (r) => r.send + r.remind + r.cancel;
console.log(`\nconfiguration surface: 2 numbers (threshold, open duration) and 1 decision (the state's key).`);
console.log(`same threshold, two separate meanings under two placements: for threshold 5, calls reaching the dependency go ` +
  `${total(t["5dependency"])} -> ${total(t["5callSite"])}; for threshold 10, ${total(t["10dependency"])} -> ${total(t["10callSite"])}.`);
console.log(`under per-call-site placement the effective threshold at the dependency level = threshold x openable call sites; ` +
  `at threshold 10 the cancel path never reached the threshold with ${t["10callSite"].cancel} calls and its breaker stayed closed.`);
process.exit(0);
node services.mjs > /dev/null 2>&1 & SP=$!
sleep 1; node measure.mjs; kill $SP
60 user requests; send and remind on every request, cancel every 20th request
notification dependency is permanently faulty: 503 on every call
threshold / placement          state objects        open objects  notification calls                send              remind              cancel             catalog
3 / dependency                             2                   1                   3                   2                   1                   0                  60
3 / call site                              4                   3                   9                   3                   3                   3                  60
5 / dependency                             2                   1                   5                   3                   2                   0                  60
5 / call site                              4                   2                  13                   5                   5                   3                  60
10 / dependency                            2                   1                  10                   5                   5                   0                  60
10 / call site                             4                   2                  23                  10                  10                   3                  60

configuration surface: 2 numbers (threshold, open duration) and 1 decision (the state's key).
same threshold, two separate meanings under two placements: for threshold 5, calls reaching the dependency go 5 -> 13; for threshold 10, 10 -> 23.
under per-call-site placement the effective threshold at the dependency level = threshold x openable call sites; at threshold 10 the cancel path never reached the threshold with 3 calls and its breaker stayed closed.

The State-Object Column

The leftmost numeric column is the placement’s direct definition: 2 state objects under per-dependency placement, 4 under per-call-site placement. Both reach the same two dependencies; the difference is whether state multiplies by dependency count or by call-site count. This number grows with the codebase: adding a call site changes nothing under per-dependency placement, but under per-call-site placement one more state object is born, and it starts learning from zero.

Whether the learning is shared shows up in the notification-calls column. At threshold 5, per-dependency placement makes 5 calls to the broken dependency: send fails three times, remind fails twice, the shared breaker opens on the fifth failure, and the cancel path never reaches the dependency at all — 0 in the column. What one call site learns, the other already knows.

Under per-call-site placement, the same threshold sends 13 calls: send 5, remind 5, cancel 3. Every call site learns the same fact from scratch. While the dependency is already down, the third call site behaves as if it knows nothing.

The Same Threshold, Two Meanings

Reading the threshold rows together reveals the real problem with the configuration surface. For thresholds 3, 5, and 10, calls reaching the dependency are 3, 5, and 10 under per-dependency placement, and 9, 13, and 23 under per-call-site placement. Whoever writes threshold 5 has said “at most five failed calls to the dependency” under one placement, and “five per call site, up to fifteen in total” under the other. The number is the same, its meaning is not, and the only thing in the code that says which meaning applies is the name of a key.

Under per-call-site placement, the effective threshold at the dependency level is the threshold multiplied by the number of call sites that can open. This multiplier appears nowhere in the configuration; it is the call-site count in the source, and it grows silently with every new call site.

The Breaker That Never Opens

The last row is this lesson’s harshest measure. At threshold 10, under per-call-site placement only 2 of the 4 state objects are open — send and remind opened, cancel did not. The reason is arithmetic: the cancel path is called only 3 times across 60 requests, and 3 failures never reach a threshold of 10.

The result is this: even though the dependency is known to be broken, that call site still goes out every time it is called. The breaker is there, its code is correct, its setting matches the others, and it will never do anything useful. A rarely called call site can never fill a threshold on its own — what stays unprotected is exactly the least exercised path. Under per-dependency placement, that same path makes zero calls, because it inherits what its frequently called neighbors already learned.

The catalog column is 60 on every row. The breaker never touches the healthy dependency and behaves the same under both placements; sharing state is not a question between dependencies, it is a question between the call sites of the same dependency.

Testability

// test/breaker.test.mjs — the assertion pins wherever the breaker's state is kept.
import { test } from "node:test";
import assert from "node:assert/strict";
import { request } from "node:http";
const call = (port, path) => new Promise((resolve) => {
  const r = request({ port, path, agent: false }, (y) => {
    let g = ""; y.on("data", (p) => (g += p)); y.on("end", () => resolve(g));
  });
  r.on("error", () => resolve("")); r.end();
});
const run = async (threshold, by, n = 60) => {
  await call(8302, "/reset");
  await call(8301, `/configure?threshold=${threshold}`);
  for (let i = 1; i <= n; i += 1) await call(8301, `/request?n=${i}&by=${by}`);
  return { s: JSON.parse(await call(8302, "/counter")), d: JSON.parse(await call(8301, "/state")) };
};

test("per dependency: what one call site learns, the other already knows", async () => {
  const { s, d } = await run(5, "dependency");
  assert.equal(s.cancel, 0);                        // the cancel path never reached the dependency
  assert.equal(d.length, 2);                        // one assertion: two dependencies, two state objects
});
test("per call site: the rare call site never reaches the threshold", async () => {
  const { s, d } = await run(10, "callSite");
  assert.equal(s.cancel, 3);                        // called three more times while the dependency was broken
  assert.deepEqual(d.find(([a]) => a === "cancel"), ["cancel", false]);
});
test("the breaker's scope is bound to the dependency", async () => {
  for (const by of ["dependency", "callSite"]) assert.equal((await run(5, by)).s.catalog, 60);
});
node services.mjs > /dev/null 2>&1 & SP=$!
sleep 1
node --test --test-force-exit --test-reporter=tap test/breaker.test.mjs |
  grep -E "^(ok|not ok|# (tests|pass|fail) )"
kill $SP
ok 1 - per dependency: what one call site learns, the other already knows
ok 2 - per call site: the rare call site never reaches the threshold
ok 3 - the breaker's scope is bound to the dependency
# tests 3
# pass 3
# fail 0

Triggering the breaker in a test is easy — keep the dependency broken and make as many calls as the threshold. What is hard is deciding which state is under test. Under per-dependency placement there is a single object, and one assertion covers every call site; that is what the first test’s d.length check rests on. Under per-call-site placement, each call site’s state must be tested separately, and as the second test shows, an untested call site stays silently unprotected. Placement decides not only runtime behavior but how many assertions a test must write: one per dependency, or one per call site.

Summary

  • The breaker’s states were built in M14/K06 Application Architecture: Routing, State and Data, its effect on the system in M19/K05 Resilience and Reliability; what is measured here is where its state is kept in the code. The difference between the two placements is the name of a single key.
  • State-object count is the direct result of placement: 2 per dependency, 4 per call site. Adding a call site under the latter births one more state object that starts learning from zero.
  • At the same threshold, calls reaching the broken dependency are 5 against 13; under per-dependency placement, one call site’s learning carries over to the others, and the cancel path never reaches the dependency.
  • The same threshold carries two separate meanings under the two placements: thresholds of 3, 5, and 10 correspond to 3, 5, and 10 calls at the dependency, but to 9, 13, and 23 calls under per-call-site placement. The effective threshold at the dependency level is the threshold times the call-site count, and that multiplier appears in no configuration.
  • A rarely called path cannot fill its own threshold: at threshold 10, the cancel path’s breaker stayed closed after 3 calls and kept going out on every call even though the dependency was known to be broken.
  • The breaker never touches the healthy dependency (catalog is 60 on every row); the sharing question sits between the call sites of one dependency, not between dependencies. In tests, per-dependency placement closes with one assertion, per-call-site placement demands one per call site.

Next Step

The breaker drew a boundary, but the boundary’s material is information: protection only begins after enough failures accumulate, and on rare paths it never begins at all. Beyond that, the breaker understands that a dependency is broken, not that the work waiting on it is consuming the caller’s own resources — all three call sites run in the same process, from the same pool, and when one slows down, what is left for the others shrinks. So one more question remains: can the contagion be cut off from the start, without ever depending on detecting it. The next lesson tries this by splitting the resource pool, and measures where the pool is defined in the code, how many call sites moving a dependency into its own pool touches, which paths bypass the pool’s limit, and what behavior a wrong pool size produces.

To keep your progress and take notes, Log in

My notes

Log in to take notes.

Start typing to search.

↑↓ Esc navigate · open · close