Skip to content
academia.sh

Lesson 18 / 25

Long-Running Jobs

Making a background job that runs for minutes manageable: breaking the job into chunks, writing progress to a durable status record, building cooperative cancellation with an abort signal, and resuming from where cancellation left off.

Contents

The previous lesson built recurring tasks and prevented two runs from overlapping with a lease lock. The task there was short: a report that finishes in three hundred milliseconds. The lease duration could be chosen accordingly.

A share of background jobs, though, run for minutes. A monthly report scanning thirty thousand loan records, regenerating an uploaded cover image at four different sizes, recomputing the branch inventory summary from scratch. As soon as a job like that starts, two questions follow. How does the party that started it learn how far it has gotten. And when the job is abandoned — the user closed the tab, the service is shutting down, the lease expired — where does it stop and what does it leave behind there.

Breaking the Job into Chunks

The first decision that makes a long job manageable is turning it from a single block into a sequence of chunks. A job that has not been chunked cannot do three things at once: it cannot report progress, it cannot be cancelled, and if it is left half-done it leaves no option but starting over.

The rule for chunking is simple: the job turns into a loop that processes a fixed number of units per round, and two things happen at the end of every round. State is written as a checkpoint, and control returns to the event loop. The second is mandatory in a single-threaded runtime: without returning, the process can do nothing else until that job finishes — it cannot even answer a health check.

Progress Reporting and Cancellation

Progress is written to two places at once. A counter kept in memory is not enough, because that counter goes away when the process crashes. This is why progress is written to a durable status record keyed by the job’s identity — here, the job_status table.

Cancellation, meanwhile, is cooperative: it is not possible to forcibly break an outside process’s loop. The outside party only raises a signal; the party doing the work reads that signal at points it chooses and stops there. In JavaScript this signal is produced with AbortController; the signal.throwIfAborted() call throws if the signal has been raised.

// long-job.mjs — monthly loan report: advances chunk by chunk, reports progress, honors cancellation
import { DatabaseSync } from "node:sqlite";

const db = new DatabaseSync(":memory:");
db.exec(`CREATE TABLE loan (loan_id INTEGER PRIMARY KEY, member_id INTEGER NOT NULL,
                             days INTEGER NOT NULL);
         CREATE TABLE job_status (job_id TEXT PRIMARY KEY, status TEXT NOT NULL,
                                 processed INTEGER NOT NULL, total INTEGER NOT NULL)`);
const insert = db.prepare("INSERT INTO loan (loan_id, member_id, days) VALUES (?, ?, ?)");
for (let i = 1; i <= 30_000; i++) insert.run(i, 100 + (i % 500), i % 30);

const RECORD_COUNT = db.prepare("SELECT count(*) AS n FROM loan").get().n;
const CHUNK = 3000;                                  // records processed per round

const readStatus = (jobId) =>
  db.prepare("SELECT processed, total FROM job_status WHERE job_id = ?").get(jobId)
  ?? { processed: 0, total: 0 };
const writeStatus = (jobId, status, processed, total) =>
  db.prepare(`INSERT INTO job_status (job_id, status, processed, total) VALUES (?, ?, ?, ?)
                ON CONFLICT(job_id) DO UPDATE
                SET status = excluded.status, processed = excluded.processed,
                    total = excluded.total`).run(jobId, status, processed, total);

async function generateReport(jobId, signal, onProgress) {
  let { processed, total } = readStatus(jobId);      // from zero or from where it was left
  try {
    while (processed < RECORD_COUNT) {
      signal.throwIfAborted();                       // cancellation is only seen at a chunk boundary
      for (const s of db.prepare(
        "SELECT days FROM loan ORDER BY loan_id LIMIT ? OFFSET ?").all(CHUNK, processed))
        total += s.days;
      processed += CHUNK;
      writeStatus(jobId, "running", processed, total); // write first, then report
      onProgress(Math.round((processed / RECORD_COUNT) * 100), processed);
      await new Promise((c) => setImmediate(c));     // hand-off point back to the event loop
    }
    writeStatus(jobId, "done", processed, total);
    return total;
  } catch (h) {
    writeStatus(jobId, "cancelled", processed, total); // the point it stopped at is recorded
    throw h;
  }
}

const JOB = "report-2026-01";
const printProgress = (percent, processed) => console.log(`  progress %${percent}  processed=${processed}`);

console.log("--- first run: cancelled at %40 ---");
const controller = new AbortController();
try {
  await generateReport(JOB, controller.signal, (percent, processed) => {
    printProgress(percent, processed);
    if (percent >= 40) controller.abort(new Error("user cancelled"));
  });
} catch (h) {
  const d = readStatus(JOB);
  console.log(`cancelled: ${h.message}  processed=${d.processed}/${RECORD_COUNT}  total=${d.total}`);
}

console.log("--- second run: continues from where it was left ---");
const total = await generateReport(JOB, new AbortController().signal, printProgress);
console.log(`done: total=${total}  average days=${(total / RECORD_COUNT).toFixed(2)}`);
--- first run: cancelled at %40 ---
  progress %10  processed=3000
  progress %20  processed=6000
  progress %30  processed=9000
  progress %40  processed=12000
cancelled: user cancelled  processed=12000/30000  total=174000
--- second run: continues from where it was left ---
  progress %50  processed=15000
  progress %60  processed=18000
  progress %70  processed=21000
  progress %80  processed=24000
  progress %90  processed=27000
  progress %100  processed=30000
done: total=435000  average days=14.50

Cancellation fires once progress reaches forty percent; so the measurement can be repeated, the cancellation point is tied to a progress threshold rather than a duration. The output shows three things.

The job stopped at twelve thousand records, not thirty thousand. The stopping point was written to the job_status table: the status column is cancelled, the processed column is 12000. And the second run resumed not from zero but from the fifty-percent point already reached; the overall total came out 435000, meaning no record was counted twice.

The third point matters more than the first two. The cancelled job was left at a resumable point. This is the only meaningful definition of cancellation for a long-running job: cancellation does not undo the work, it freezes it at a consistent point.

Write Order

The writeStatus call happening before the onProgress call in the code is a deliberate choice. If the order were reversed, a process crash after the progress report but before the write would leave progress reported but not saved; resuming would then reprocess records that had actually already been handled.

By the same reasoning, writeStatus and processing the chunk must sit inside a single transaction boundary. This condition is satisfied here because the report total is kept in memory and written together with the status. If the chunk wrote a row to a table instead, that write and the progress update would need to be handled together inside the unit of work introduced in the Data Access Layer and Business Logic curriculum.

Chunk Size

Chunk size looks like a single number, but it sets two magnitudes at once. A small chunk speeds up cancellation, because the signal is only read at a chunk boundary; a large chunk cuts down on progress writes, because every round means one write.

Chunk size Progress writes Cancellation delay (worst case)
100 300 100 records
500 60 500 records
3000 10 3000 records
10000 3 10000 records

For a thirty-thousand-record job the numbers fall out of division: the write count is 30000 / chunk, and the cancellation delay in the worst case is one chunk. A chunk of one hundred shortens how long the user waits after pressing cancel, but it makes three hundred writes to the status table; a chunk of ten thousand cuts writes to three, but cancellation takes as long as processing ten thousand records.

The choice is made according to the nature of the job. For a job the user started and is waiting on, cancellation delay dominates. For a job running overnight that nobody is watching, write count dominates; the progress record there is kept not for cancellation but for resuming from where it was left.

Sources of Cancellation

The cancellation signal does not come from a single place. It has four sources, and all four can be tied to the same signal.

User cancellation is the most visible one: a stop request from the report screen turns into a call carrying the job’s identity, and that call raises the signal. Time limit is the second source; AbortSignal.timeout(ms) produces a signal that raises itself once the given duration elapses, and it keeps the job from running for an unacceptable length of time. Shutdown is the third: a process that receives SIGTERM during the graceful shutdown sequence built in the previous lesson raises the signal for its running jobs and waits for them to finish. Lease expiry is the fourth; if the job’s lock is about to be taken over, the old run must cancel itself, or two runs end up continuing the same job.

When all four are wired to the same structure, a single checkpoint remains in the job’s code. AbortSignal.any([...]) folds several signals into one; the job reads only that one signal and does not need to know who cancelled it.

Where the Job’s Result Goes

Progress reporting has one gap: here it was written with console.log. In a real service, the party that needs to read progress is a different process — the client requesting the report. This is why the job_status table serves not only for resuming but also as a queryable status record: a client can read this record using the job’s identity.

How that read happens — whether the client asks over and over, or the server sends changes on its own — is the question this curriculum’s last topic answers.

Summary

  • A long job is turned into a loop that processes fixed-size chunks; state is written at the end of every round and control returns to the event loop.
  • Progress is written not to memory but to a durable status record keyed by the job’s identity; the record serves both as queryable progress and as the resume point.
  • Cancellation is cooperative: the outside party only raises a signal, and the party doing the work reads that signal at chunk boundaries and stops at a consistent point.
  • The status write happens before the report; the reverse order causes the same records to be processed twice because of progress that was reported but not saved.
  • Chunk size is the trade-off between cancellation delay and progress-write count; delay dominates for a job the user is waiting on, write count dominates for a job running overnight.

Next Step

In this lesson there was a single job, and chunking it was enough. In a real background, jobs flow continuously: every loan transaction spawns a notification job, every cover upload spawns an image-processing job. The arrival rate of jobs and the completion rate of workers are independent of each other, so they are not obligated to be equal. When the arrival rate exceeds the completion rate the queue grows; if the growth continues, memory, disk, or the latency budget runs out. The next lesson names this imbalance and compares the options for a bounded queue with numbers.

To keep your progress and take notes, Log in

My notes

Log in to take notes.

Start typing to search.

↑↓ Esc navigate · open · close