FlowingDev

Diff, explained: The secret sauce of Git and every code review

Learn how diff algorithms find precise additions, deletions, and changes between texts, the core engine behind version control and code collaboration.

Try the tool: Diff Checker

In one sentence

A 'diff' is a computed summary of the precise differences between two files or blocks of text, showing exactly what was added, removed, or changed to get from the "before" to the "after".

The problem it solves

Picture the world of computing in the early 1970s. Storage is mind-bendingly expensive, and connecting to a remote computer happens over a modem that's slower than a sleepy snail. You're a developer at Bell Labs, and you need to update a source code file on a server across campus. The file is a few thousand lines long, but you only changed three of them.

Do you send the entire file again over that molasses-slow connection? Heck no. That's a waste of time and resources. What you really want is to send only the changes.

This was the exact problem that led Douglas McIlroy to create the original diff command for the Unix operating system in 1974. His goal was to create a tool that could programmatically find the minimum set of line-by-line changes needed to turn one file into another. The output of this diff command, a "patch" file, was tiny and could be sent quickly. The recipient could then use a companion program, patch, to apply these changes to their own copy of the original file, bringing it up to date.

This simple, powerful idea—isolating change itself as a piece of data—was revolutionary. It's the fundamental building block of all modern version control systems like Git, Subversion, and Mercurial. It's the engine behind code reviews, document collaboration tools (like Google Docs' "Suggesting" mode), and configuration management systems. It solves the core problem of tracking and communicating evolution in any digital text.

How it works under the hood

At first glance, diffing seems simple: just scan two texts and flag what's different. But to do it efficiently and produce the smallest, most readable set of differences is a classic computer science problem. The secret isn't looking for what's different, but for what's the same.

The Longest Common Subsequence (LCS)

Most diff algorithms, including the famous Hunt–McIlwain algorithm that powered the original diff, are based on solving the "Longest Common Subsequence" (LCS) problem.

A subsequence is a sequence of items that appear in the same order as in the original sequence, but not necessarily next to each other. The LCS is the longest such subsequence that two sequences have in common.

Let's use a simple, non-code example.

  • Original: The quick red fox
  • New: The slow red cat

The algorithm, working line by line (or in this case, word by word), finds the Longest Common Subsequence is: The red.

Once the LCS is found, the logic is simple:

  • Any item in the Original that is not in the LCS must have been deleted. (quick, fox)
  • Any item in the New text that is not in the "LCS" must have been added. (slow, cat)

By finding the longest bedrock of shared content, the algorithm can clearly and concisely identify the islands of change around it. This method produces a minimal set of differences, which is what we want for a clean, understandable diff.

From LCS to a Readable Diff

Finding the changes is only half the battle. The other half is presenting them in a standardized, readable format. You've probably seen this if you've ever looked at a GitHub pull request. The most common format is the "unified diff format."

Let's apply it to a slightly different example:

  • File A (old):
    An apple a day.
    Keeps the doctor away.
    Or so they say.
    
  • File B (new):
    An apple a day,
    Keeps the doctor away.
    For what it's worth.
    

A diff tool would generate something like this:

--- a/file_a.txt
+++ b/file_b.txt
@@ -1,3 +1,3 @@
-An apple a day.
+An apple a day,
 Keeps the doctor away.
-Or so they say.
+For what it's worth.

Let's break that down:

  • --- a/file_a.txt: The "from" file. The - indicates the source of deletions.
  • +++ b/file_b.txt: The "to" file. The + indicates the source of additions.
  • @@ -1,3 +1,3 @@: This is the "hunk header." It's a bit cryptic, but it gives you context. -1,3 means "this hunk starts at line 1 and is 3 lines long in the original file." +1,3 means "this hunk starts at line 1 and is 3 lines long in the new file."
  • Lines starting with a space ( ) are context lines. They are identical in both files and are shown to help you understand where the change occurred.
  • Lines starting with - are deletions. They exist only in the "before" text.
  • Lines starting with + are additions. They exist only in the "after" text.

The tool shows the change from An apple a day. to An apple a day, not as a modification of a single line, but as a deletion of the old line and an addition of the new one. This line-based approach is a core characteristic of most traditional diff tools.

Beyond Plain Text: Semantic Diffs

A standard line-based diff is great for prose or code, but it falls apart with structured data like JSON, XML, or YAML.

Consider this JSON:

// Original
{
  "name": "Alex",
  "role": "Developer"
}

And this one:

// New
{
  "role": "Developer",
  "name": "Alex"
}

A text-based diff would see this as a complete wipe and rewrite:

-  "name": "Alex",
-  "role": "Developer"
+  "role": "Developer",
+  "name": "Alex"

This is technically true but semantically useless. The order of keys in a JSON object generally doesn't matter. A semantic diff tool is smarter. It parses the text into a data structure first, then compares the structures. It would correctly identify that these two JSON objects are identical, resulting in no differences. This is crucial for comparing configuration files, API responses, or any other structured data where you care about meaning, not just text formatting.

Real-world stories

The One-Character Bug That Crashed the Checkout

A junior developer was integrating a new payment provider. They copied the example API request from the documentation, plugged in their keys, and ran it. It failed. They tried again. Failed. They spent hours staring at their code and the documentation, convinced they were identical. Frustrated, they pasted the "working" documentation example into one side of a diff checker and their own code into the other.

At first, it looked identical. But then they noticed a subtle highlight at the end of their API key line. A single, invisible trailing space. The copy-paste from a web page had included it, and their code was dutifully sending it along, invalidating the key. The diff tool, which saw the space as just another character, was the only "pair of eyes" that could spot it.

Lesson: A diff is your ultimate microscope. It has no assumptions and will show you exactly what's there, including the invisible characters that can bring a system to its knees.

The Configuration Drift Debacle

A high-traffic website started experiencing bizarre, intermittent errors. The on-call engineer, Maya, was stumped. The last deployment was a week ago and had been stable. Nothing in the logs pointed to a clear cause. Her spidey-sense told her something on the server had changed manually.

She pulled the official Nginx configuration file from their Git repository, then SSH'd into the production server and copied the actual running configuration. She pasted both into a diff tool. Bingo. Three lines were different. Someone had added a "temporary" redirect rule directly on the server to fix a minor issue last week and completely forgot about it. This "fix" was now clashing with new traffic patterns. Maya removed the rogue lines, and the errors vanished. The team immediately implemented a policy to audit server configs against Git daily.

Lesson: Your version control system is your source of truth. Diffing reality against that source of truth is the best way to detect "configuration drift" and find unauthorized or forgotten changes.

The "Looks Good To Me" Code Review

A senior developer, Ben, received a pull request from a new hire. The title was "Updates." The diff was a sea of red and green across 20 files and over 3,000 lines. It contained a new feature, a fix for an unrelated bug, a massive code reformatting to switch from tabs to spaces, and a library upgrade. It was impossible to review. Was the bug fix correct? Did the new feature introduce a security hole? It was hidden in a blizzard of whitespace changes.

Ben rejected the PR with a kind note: "Welcome! A PR's diff tells a story. This one is trying to tell four different stories at once. Can you please split this into four separate PRs?" The new hire did. The reformatting PR was instantly approved. The bug fix was easy to verify. The library upgrade was straightforward. And the new feature could finally be reviewed on its own merits.

Lesson: A diff's value is inversely proportional to its size and complexity. Small, focused diffs representing a single logical change are easy to review, understand, and debug later.

Common mistakes and traps

  • Ignoring whitespace. A change from tabs to spaces or the addition of a trailing newline can look like a massive, file-wide change to a diff tool. While sometimes this is intentional, it often just creates noise that hides the real, meaningful changes. Configure your tools to ignore or highlight whitespace changes appropriately.
  • The "semantic-blind" diff. As mentioned earlier, using a plain text diff on structured data like JSON or XML can be incredibly misleading. Reordering attributes or keys can look like a huge change when, functionally, nothing is different. Always reach for a semantic-aware diff tool for these formats.
  • Forgetting context. A diff shows you what changed, but it never tells you why. That's the job of the commit message or the pull request description. A diff without context is like an answer without a question; it's hard to judge if it's right or wrong.
  • Creating "Frankenstein" diffs. Lumping unrelated changes into a single commit (a bug fix, a feature, and a typo correction) makes the diff a nightmare to read. It makes it impossible to revert one of those changes later without affecting the others. Each commit should be a single, logical, atomic change.

Why it belongs on your radar

Understanding diffing is not optional for a modern developer; it's as fundamental as knowing how to use a keyboard. You will encounter diffs multiple times a day, every day:

  • When you run git status or git diff to see your own uncommitted work.
  • When you create a pull request for your colleagues to review.
  • When you review someone else's pull request.
  • When you use git blame to figure out who wrote a specific line of code and why.
  • When you're debugging a problem by comparing a working configuration with a broken one.

Even for non-developers, the concept is powerful. It's the "Track Changes" in your Word document. It's the version history of your Wikipedia article. It's the ability to see how a contract has evolved between drafts. Understanding diffing is understanding how we manage and communicate change in the digital world. It's the auditable, verifiable log of progress.

Go deeper

Theory done. Time to get your hands dirty — 100% in your browser.

Try the tool: Diff Checker