TypeBack

Guide · Updated

Comprehension debt: what it is and how to pay it down

Comprehension debt is code in your repository that nobody on the team understands. What the research says, including Anthropic's 2026 study, and how to measure and reduce it.

Technical debt is code that is hard to change. Comprehension debt is code that nobody understands, however clean it is. Coding agents make the second kind cheap to take on: a working change arrives faster than anyone can read it, and it gets merged because it works.

What comprehension debt is

Addy Osmani defines comprehension debt (March 2026) as “the growing gap between how much code exists in your system and how much of it any human being genuinely understands.” Margaret-Anne Storey describes the same problem from the team's side as cognitive debt (February 2026): debt that lives in the developers' heads rather than in the code. She builds on Peter Naur's idea that a program is a theory held by the people who work on it. When the theory is lost, even well-written code is hard to change safely.

The two terms point at one thing. Code you shipped that no person on the team can explain.

How it differs from technical debt

Technical debtComprehension debt
Where it livesIn the codeIn the gap between the code and the people
What it looks likeDuplication, tangled modules, missing testsClean code nobody wants to touch
Found byLinters, reviews, slow changesIncidents, and questions nobody can answer
Paid down byRefactoringSomeone understanding the code

What the evidence says

  • Delegating costs understanding. In Anthropic's randomized study (January 2026), engineers who learned a new library with an AI assistant scored 50% on a follow-up quiz, against 67% for those who coded by hand. The biggest gap was in debugging. Participants who delegated averaged under 40%; those who asked the assistant conceptual questions scored 65% or more.
  • Trust without verification. In Sonar's January 2026 survey of more than 1,100 developers, 96% said they do not fully trust that AI-generated code is correct, yet only 48% always check it before committing. Sonar sells code quality tools, so read it as a vendor survey.
  • Less refactoring, more copying. GitClear found that moved lines, a rough stand-in for refactoring, fell from 24.8% of changed lines in 2021 to 9.5% in 2024, while copy-pasted lines rose. Code that is added and never reworked is code nobody had to understand.
  • Speed is not the whole story. METR found experienced developers were 19% slower with AI tools in early 2025, while believing they were faster. Its February 2026 follow-up suggests a speedup now, with data METR itself calls unreliable. Whichever way speed goes, understanding has to be checked separately.

How to measure it

You cannot measure understanding directly, but you can count the changes nobody has been asked about. For each commit or branch, track:

  1. AI-written hunks: how many hunks an agent wrote. Agent hooks can capture this as each edit happens.
  2. Hunks a person answered a question about: predicted the output, traced a path, found a planted bug, or explained the purpose.
  3. Skips, with a reason, such as low risk.

The gap between the first number and the second is your comprehension debt for that change. Put it in the commit, where it stays with the code:

Reviewed-by-human: 14/17 hunks (questions), 2 skipped (low-risk)

Over a branch, the same counts tell a reviewer how much of the pull request its author has actually worked through, and which high-risk files were never looked at.

How to reduce it

  1. Review when the code is written, not when it is merged. A pull request with 40 AI-written hunks is too late. Review each hunk as the agent produces it, while you still remember what you asked for.
  2. Ask questions instead of rereading. Predict an output, trace a path, find the bug, say why. The four-question method takes a minute a hunk.
  3. Spend the effort where the risk is. Authentication, payments, concurrency, data deletion and CI config first. A renamed variable can be skipped, as long as the skip is recorded.
  4. Ask the agent conceptual questions. Both Osmani and Anthropic's data point the same way: people who interrogate the tool keep their understanding, and people who delegate lose it.
  5. Make it a team norm, not a gate. Storey suggests that at least one person fully understands each AI change before it ships. A policy can ask for a review record on every pull request without blocking anyone's editor.

What not to do

  • Do not block merges on a quiz. People route around gates, and a gate invites the obvious joke: ask the AI to answer the AI's quiz. A record someone chose to keep is worth more than a pass someone was forced into.
  • Do not score individuals. Report comprehension per change or per team. Per-person scores turn a learning tool into surveillance.
  • Do not treat the record as proof. It is the author's attestation, like a DCO sign-off. It does not replace code review.