Guide · Updated
How to understand code your AI agent wrote
A hunk-by-hunk method for reading code Claude Code, Copilot or Codex wrote: predict it, trace it, find the bug, say why. With the evidence for why it works.
The agent says it is done. The diff is two hundred lines across six files. You scroll through it, nothing looks wrong, the tests pass, and you commit. Two weeks later a bug report lands in that code, and you realise you cannot explain what it does. You read it. You never understood it.
Reading and understanding feel the same while you do them. They are not. This guide is a method for the second one: small enough to do on every change an agent makes, and concrete enough that you know when you are done.
Why skimming AI-written code fails
Code an agent writes looks right by construction. It is fluent, formatted and plausible, which is exactly what makes a skim pass it. Developers know this. In the 2025 Stack Overflow Developer Survey, the most common frustration with AI tools, named by 66% of developers, was “AI solutions that are almost right, but not quite”. 45% said debugging AI-generated code takes longer.
The cost shows up later, in what you can still do on your own. In a randomized study Anthropic published in January 2026, 52 mostly junior engineers learned a new Python library, with or without an AI assistant. On a quiz afterwards the AI group averaged 50% and the hand-coding group 67%. The widest gap was on debugging questions. How people used the assistant mattered more than whether they used it: those who delegated the work averaged under 40%, while those who asked it conceptual questions scored 65% or higher.
That last finding is the useful one. Using an agent does not rot your understanding. Accepting its output without asking anything of it does.
Work in hunks, not diffs
A diff is too big to hold in your head. A hunk is one contiguous block of changed lines: a function body, a branch, a config block. It is small enough to understand completely, and it is the unit you would point at in a code review.
- Review hunk by hunk, in the order the risk suggests, not the order of the files.
- Start with high-risk hunks: authentication, money, concurrency, anything that deletes data, and anything that runs in CI or on deploy.
- Low-risk hunks, such as a rename, a log line or a test fixture, get a glance.
- Do it while the change is fresh. The best time is while the agent is still working on the next part.
Ask each hunk four questions
Understanding is being able to answer questions you did not see coming. So ask them. Each of the four below tests something different, and each has an answer you can check. Here is a hunk an agent wrote, to work through:
1. Predict: what does it do for this input?
Pick one concrete input and say what comes out, before you run anything. Choose an input near an edge.
The cart already holds 3 of this product and its stock is 4. What does addToCart(user, product, 2) return?
nextQty is 3 + 2 = 5, which is more than the 4 in stock, so the function returns { ok: false, reason: 'out_of_stock' } before the upsert runs. If you guessed { ok: true, qty: 5 }, you read the happy path and skipped the guard. That is the most common way a skim goes wrong.
2. Trace: which path does the code take?
Follow one input through the branches, line by line, to where the call ends. Tracing catches the cases prediction skips: early returns, thrown errors, and branches that never run.
The product id does not exist. Which line ends the call? The findUnique on line 14 returns null, so line 15 throws Unknown product. The cart is never created. Is that what the caller expects, or should an unknown product be a result like out-of-stock is? You now have a real question for the code, which is better than a vague feeling.
3. Spot the bug: which line would hurt most if it were wrong?
Assume one line is wrong and find it. This is mutation thinking: if a single character changed, would you notice? Here is the same hunk with one line changed.
Line 20 now says >=. With it, a cart can never take the last unit in stock: 4 of 4 is refused. Tests that only add one item to a large stock pass either way. Off-by-one comparisons, inverted conditions and a missing await are the bugs this question is for.
4. Say why: what is it for, and why this way?
Say the purpose in one sentence, then name the obvious alternative and why the agent did not use it. Why does running out of stock return a result instead of throwing? Because it is an expected outcome the caller should show to the user, not a failure. Why upsert the cart line instead of inserting it? Because the product may already be in the cart. If you cannot answer, the hunk is not yours yet.
When you cannot answer
Rereading harder rarely helps. Change what you are doing instead.
- Ask the agent a conceptual question, not “is this right?”. Ask what happens when the product is missing, or why it chose upsert. In Anthropic's study this style of use kept scores high.
- Run it with the input you could not predict, in a test or a REPL, and compare with your guess.
- Type it back. Retyping a hunk forces you to read every character, and the questions come easier afterwards. Retrieval beats rereading for retention; this is the testing effect, and it is part of why answering questions works better than reading again.
Keep an honest record
A week later you will not remember which hunks you understood and which you waved through. Write it down where it stays with the code: in the commit. A trailer line works well because git already knows how to read one:
Count skips as skips. The record is your own attestation, like a DCO sign-off. It is not proof that you understood anything, and it does not replace code review by someone else. What it gives you is a number you cannot fool yourself about, and a team a way to ask for one.
The checklist
| Step | Done when |
|---|---|
| Split the change into hunks | You have a list, high-risk first |
| Predict | You said the output for an edge input before running it |
| Trace | You can name the line where each path ends |
| Spot the bug | You checked every comparison, condition and await |
| Say why | One sentence of purpose, and why not the obvious alternative |
| Record it | The commit says how many hunks you reviewed and skipped |