---
title: 'Patch-Based Contribution'
source: 'https://academia.sh/en/courses/branching-and-collaboration/patch-based-contribution'
course: 'Branching and Collaboration'
language: en
updated: '2026-08-17T18:10:45+00:00'
license: 'CC BY-SA 4.0'
---

# Patch-Based Contribution

A patch is a text file and encodes the change, not the graph: push with history leaves 5 questions in 15 commits, the mail-formatted series 4 in 12, the plain diff series 3 in 12, the single plain diff 1 in 1 commit, and under the plain diff, author-preserved commits are 0.

The previous lesson's flow assumed a repository platform: a shared address, a discussion area,
and someone with authority to make the integration decision. There is one more form a
contribution can take with none of these, and it is version control's oldest contribution form:
sending the change **as a file**.

A patch fits into a mailbox, not a repository. This lesson's question is not how a patch is
produced — production amounts to nothing more than two commands. The question is this: what
does the patch carry **to the other side**, what does it not carry, and where does the field it
does not carry fall among the six questions?

## Two Patch Formats

A **patch** is a change encoded as text, and it comes in two separate forms. The first is the
**plain diff**: it carries only the change text, no metadata.

```text
# taught commands — example dump, not executed

git diff upstream/main..metrics > metrics.patch
git apply --check metrics.patch
git apply metrics.patch
```

The `--check` option tests the patch without applying it and does not touch the working
directory. When application succeeds, the changes are written to the working directory and
**no commit is created**: whoever applies the commit does so under their own name and their own
time.

The second is the **mail-formatted patch series**: one file is produced per commit, and a mail
header sits at the top of the file.

```text
# taught commands and file header — example dump, not executed

git format-patch upstream/main..metrics -o patches/
git am --3way patches/0001-read-metrics-threshold-from-configuration.patch

From <commit ID> Mon Sep 17 00:00:00 2001
From: author <author@example.test>
Date: <author date>
Subject: [PATCH 1/4] Read metrics threshold from configuration

The threshold value was embedded in the script; whenever the measurement
point changed, the script had to be edited.
---
 metrics.py | 22 ++++++++++++++++++++++
 1 file changed, 22 insertions(+)
```

Three fields from the header pass straight into the commit: author, author time, and message.
`git am` writes one commit per file and reads these fields not from the local side but **from
the file**. The `--3way` option, when the patch does not apply directly, tries a three-way
merge against the merge base; if a conflict comes up, the decision is again made by hand.

## What a Patch Cannot Carry

There is a class of field neither format can carry, and the reason is structural. A patch is a
**text file** and encodes a change; history, on the other hand, is a **graph**, and a text file
carries no edges. This is why the commit a patch produces on the other side always has a **new
ID**, its parent link is the tip of wherever it was applied, and the branch record never crosses
over — there is no such thing as a branch in a patch, only a file.

For this reason, the measurement **derives** history from the set of carried metadata; it does
not describe the formats by hand.

## The Measurement's Assumptions

- **RR25** — The same twelve-commit development is carried to the other side under four
  separate contribution formats; the code produced is the same in all four, and the only thing
  measured is the **metadata** that remains on the other side.
- **RR26** — Each format is defined by a **set of carried metadata**; the history on the other
  side is computed from this set, not written by hand.
- **RR27** — The mail-formatted series carries the author, author time, and message; it does
  not carry the branch record or the parent link.
- **RR28** — The plain diff series carries only the change text and the message; the author and
  time are written locally at the moment of application and count as "not carried" in this
  measurement.
- **RR29** — The single plain diff collects twelve commits' difference into a single file and
  becomes a single commit on the other side; it is the most extreme form of squashing.
- **RR30** — The "author kept" column counts commits on the other side whose author field shows
  the original contributor; merge commits are written under the integrator's name and do not
  enter the count.

## Measurement

```python
"""Patch-based contribution: carried metadata and the cost of what is not carried.

Part 1 - the metadata fields carried by four contribution formats.
Part 2 - the history the same formats leave and the questions they answer.
"""
SEED = 20260813
BRANCHES = ("metrics", "report", "identity")
FILES = {"metrics": "metrics.py", "report": "report.py", "identity": "identity.py"}
BUGGY = ("report", 2)
SHARED_FILE = "config.py"
METADATA = ("author", "author time", "subject and body", "branch record", "parent link")
CARRIED = {"push with history": set(METADATA),
           "mail-formatted series": {"author", "author time", "subject and body"},
           "plain diff series": {"subject and body"},
           "single plain diff": set()}


def rng(seed):
    state = seed % 2147483646 + 1

    def draw(n):
        nonlocal state
        state = (state * 48271) % 2147483647
        return state % n
    return draw


def development():
    """Real development: three branches, four commits each, one carries the bug."""
    draw, record, time = rng(SEED), [], 0
    for step in range(1, 5):
        for branch in BRANCHES:
            time += 1 + draw(3)
            file = SHARED_FILE if step == 3 and branch in ("metrics", "report") else FILES[branch]
            record.append({"branch": branch, "step": step, "file": file,
                          "buggy": (branch, step) == BUGGY, "time": time})
    return record


def apply(record, kind):
    """The history the contribution leaves on the other side: derived from carried metadata."""
    carried = CARRIED[kind]
    if kind == "push with history":
        t = []
        for branch in BRANCHES:
            for k in [x for x in record if x["branch"] == branch]:
                t.append({**k, "branch_record": branch, "time_kept": True,
                          "author_kept": True})
            t.append({"branch": branch, "step": 0, "file": None, "buggy": False,
                      "time": max(x["time"] for x in record if x["branch"] == branch),
                      "branch_record": branch, "merge": True, "time_kept": True,
                      "author_kept": False})
        return t
    if kind == "single plain diff":
        return [{"branch": None, "step": 0,
                 "file": sorted({x["file"] for x in record}),
                 "buggy": any(x["buggy"] for x in record),
                 "time": max(x["time"] for x in record), "branch_record": None,
                 "time_kept": False, "author_kept": False,
                 "squashed": len(record)}]
    return [{**k, "branch_record": k["branch"] if "branch record" in carried else None,
             "time_kept": "author time" in carried,
             "author_kept": "author" in carried}
            for branch in BRANCHES for k in record if k["branch"] == branch]


def question1_grouping(t):
    return all(x.get("branch_record") for x in t if not x.get("merge"))


def question2_bug(t):
    k = [x for x in t if x["buggy"]]
    return len(k) == 1 and not k[0].get("squashed")


def question3_order(t):
    return all(x.get("time_kept") for x in t)


def question4_file(t):
    return all(isinstance(x["file"], str) or x.get("merge") for x in t)


def question5_integrity(t):
    spots = {}
    for i, x in enumerate(t):
        if x.get("merge"):
            continue
        spots.setdefault(x["branch"], []).append(i)
    return all(y[-1] - y[0] == len(y) - 1 for y in spots.values())


def question6_conflict(t):
    return any("chosen" in x for x in t)


QUESTIONS = (("feature grouping", question1_grouping), ("bug isolation", question2_bug),
             ("true order", question3_order), ("file trail", question4_file),
             ("branch integrity", question5_integrity), ("conflict decision", question6_conflict))

record = development()
print(f"real development {len(record)} commits, {len(BRANCHES)} branches, "
      f"buggy commit {sum(k['buggy'] for k in record)}")
print()
print("metadata field  " + "".join(f"{name:>23s}" for name in CARRIED))
for field in METADATA:
    print(f"  {field:17s}" + "".join(f"{('carries' if field in CARRIED[k] else 'does not carry'):>23s}"
                                    for k in CARRIED))
print()
print(f"{'contribution format':<24s}{'commits':>9s}{'author kept':>14s}"
      f"{'readable':>10s}{'answered':>10s}{'lost':>7s}")
for kind in CARRIED:
    t = apply(record, kind)
    readable = sum(1 for x in t if not x.get("merge") and not x.get("squashed"))
    print(f"  {kind:<22s}{len(t):9d}{sum(x['author_kept'] for x in t):14d}"
          f"{readable:10d}{sum(f(t) for _, f in QUESTIONS):10d}{len(record) - readable:7d}")
print()
print("unanswered questions:")
for kind in CARRIED:
    t = apply(record, kind)
    print(f"  {kind:<22s} " + ", ".join(name for name, f in QUESTIONS if not f(t)))
```

```
real development 12 commits, 3 branches, buggy commit 1

metadata field        push with history  mail-formatted series      plain diff series      single plain diff
  author                           carries                carries         does not carry         does not carry
  author time                      carries                carries         does not carry         does not carry
  subject and body                 carries                carries                carries         does not carry
  branch record                    carries         does not carry         does not carry         does not carry
  parent link                      carries         does not carry         does not carry         does not carry

contribution format       commits   author kept  readable  answered   lost
  push with history            15            12        12         5      0
  mail-formatted series        12            12        12         4      0
  plain diff series            12             0        12         3      0
  single plain diff             1             0         0         1     12

unanswered questions:
  push with history      conflict decision
  mail-formatted series  feature grouping, conflict decision
  plain diff series      feature grouping, true order, conflict decision
  single plain diff      feature grouping, bug isolation, true order, file trail, conflict decision
```

## Reading the Steps

The bottom table gives four steps, and each step has a metadata counterpart.

**Push with history** leaves fifteen commits and **5** questions; this is the reference row,
every field has passed through.

**The mail-formatted series** leaves twelve commits and answers **4** questions. The only
question that drops is **feature grouping**; the reason is written in the top table's fourth
row — the branch record is not carried. Because the author and time are carried, "true order"
stays standing.

**The plain diff series** also leaves twelve commits, but answers **3** questions. The
difference comes from a single field: because the author time is not carried, every commit
carries the time of the moment it was applied, and **true order** cannot be read. The same
number of commits answers one fewer question, because of a single metadata field.

**The single plain diff** reduces twelve commits to one, answers **1** question, and the lost
commit count is **12**. This is the lowest value measured in the course; squashing reduced
twelve commits to three and answered **2**. The sixth question, meanwhile, is unanswered in all
four rows — which transport the contribution took does not change this.

## The Field the Measure Does Not See

The third row carries a warning aimed at the course's own measure. Between the plain diff series
and the mail-formatted series, the **author kept** column falls from **12** to **0**: none of
the twelve commits show who wrote them anymore, all are written under the applying person's
name. Yet the answered question count falls by only **1**, and the question that drops has
nothing to do with authorship.

The reason is plain: **none of the six questions ask about authorship.** The questions ask about
grouping, isolation, order, file trail, branch integrity, and conflict decision; who the
contribution belongs to is not in that set, so the measure cannot see it. This is not a flaw in
the measure, it is a declaration of the measure's **scope**. Contribution ownership and
attribution sit outside that window, and the plain diff series silently erases them.

## Summary

- A patch is a text file and encodes the change; because it does not encode the graph, no patch
  format carries the commit ID, the parent link, or the branch record. The mail-formatted series
  carries the author, the time, and the message; the plain diff carries only the change text.
- The four contribution formats leave **15**, **12**, **12**, and **1** commits on the other
  side and answer **5**, **4**, **3**, and **1** questions respectively.
- The mail-formatted series and the plain diff series leave the same number of commits; the
  one-question difference between them comes only from the **author time** field not being
  carried.
- When the number of author-preserved commits falls from **12** to **0**, the answered question
  count falls by only one: none of the six questions ask about authorship, and this is the
  measure's scope boundary.

## Course Wrap-Up

The course ran on a single measure: an integration format's number is not the code it produces,
it is the **number of questions** the history it leaves can answer. Fifteen lessons applied this
measure to separate places.

| lesson | measured | answered question or loss |
|---|---|---|
| The Concept of a Branch | the cost of branching | 0 new commits and a 123-byte reference; the copy model does the same job by copying 12,000 files |
| Branch Creation and Switching | what switching does to changes left in hand | 3 of 12 states are carried over, 5 are blocked, 4 are left behind; silent loss 0 |
| Branch Naming | the naming convention's contribution to the first question | the group reads in 3 of 8 combinations; contribution is 0 in a format that keeps no record |
| Detached HEAD State | the cost an unreachable commit charges history | 15 entries fall to 13, the answer falls from 5 to 4; all 3 of the three counting queries are answered incompletely, and history does not report it |
| Fast-Forward and Merge Commit | the history the two merge formats leave | 15 commits/5 questions and 12 commits/3 questions; the difference is created by grouping and branch integrity |
| Conflict Resolution | whether the resolution passes into history | all four of the four formats fail to answer the sixth question, 0/4; if the rationale is written to the message, 1 of the 15 commits carries it |
| Rebase | the history linearization leaves | 12 commits / 3 questions; the total matches fast-forward, the loss is separate, the difference is 1/6 |
| Cherry-Pick | the same change standing in two places | file trail stays yes, touched commits rise from 3 to 4; in the buggy copy the answer falls from 5 to 4 |
| Merge Strategies | four formats side by side | 15/12/3/12 commits and 5/3/2/3 questions; under squash, 12 are lost |
| Branch Strategies | what branch lifetime writes into history | 12, 15, 18, and 24 commits; all three branch layouts answer 5 questions — the six questions do not distinguish between branch layouts |
| Defining a Remote Repository | what defining adds to history | three copies carry 5/10/15 commits and answer 4/5/5 questions; after two definitions, local stands at 5 commits / 4 questions |
| Fetch, Pull, and Push | the history the three operations change | fetch leaves the branch at 5, pull carries it to 16 and 5 questions, push changes the other side from 10 to 16 |
| Forking and Cloning | what the upstream repository can answer | 15/5, 12/3, and 3/2; under squash, 3 questions remain only in the fork repository |
| Pull Request Flow | what the flow's chosen format leaves | 5/3/2/3 questions; 7 of the 11 decisions never enter history, and under squash even the written 4 become unreadable |
| Patch-Based Contribution | the metadata the patch carries and does not carry | 15/5, 12/4, 12/3, and 1/1; under plain diff, author-preserved is 0 and the six questions never ask about it |

Three readings follow from the table. **First:** readability and answerability move in opposite
directions — the cleanest histories are the ones that answer the fewest questions, and cleaning
up history erases information. **Second:** the sixth question is answered in no row; how a
conflict was resolved is, by the record's own definition, not in the record, and the only
remedy is writing the rationale to a commit message. **Third:** loss comes from two separate
sources. Format loss erases information and cannot be recovered; scope loss is only a copy's
absence, and it closes with a fetch.

Throughout the course, history was always **read**: questions were asked, answers were counted,
and what the formats left behind was taken as given. The next course — Advanced Git — takes up
the tools for **rewriting** that same record: interactive rebase, bulk history transformation,
force push discipline, and the reflog that rescues lost commits. The distinction has to be set
up front: **reading history and rewriting it are separate jobs.** Reading changes nothing;
rewriting can permanently change the answers to the questions counted in this course. The
measure this course leaves behind is the unit that will be used to compute the cost of every
tool there.
