Six questions to ask before you approve a pull request your agent wrote

Six of them, and none is about whether the code is correct. Correctness is what the diff shows you and what the tests argue about. The questions below are the ones that decide whether the team still understands the system after this merges, which is the part nobody is checking.

Backthread is built around that gap: agents write the change, a person merges it, and the reasoning behind it has to be captured somewhere or the codebase quietly becomes a thing nobody can explain. You do not need the product to ask the questions. You need ten minutes and a habit.

Why the diff cannot answer them

A diff is a complete record of what changed and a very poor record of why. Line-level attention catches the null check that is missing. It does not catch the approach that was rejected, the assumption the change makes about a service two teams away, or the case that was knowingly punted. Those never appear as lines, so no amount of reading harder surfaces them.

The questions are ordered by payoff. If you only ever ask the first two, you will still be ahead of where you are now.

The six questions

1. What did this decide that it could have decided otherwise?

The one that matters most, and the one with the most uncomfortable answer. Often there is no alternative to name, because none was weighed. Measured on our own repository, 62 percent of the decisions captured from agent sessions record no alternative and no trade-off at all, and we wrote up why that reasoning usually never formed rather than getting lost.

"Nothing was weighed" is a legitimate answer, and recording it is worth as much as recording a rejected option. What is not legitimate is a rationale invented after the fact to fill the box.

2. What does this assume about the rest of the system?

Every change takes something for granted: that the queue is ordered, that this table is small, that the caller already validated the input. The assumption is what breaks somebody else's code six weeks later, and it is cheapest to write down while the session that made it is still open.

3. Which area does this land in, and who else has read that area lately?

Not who touched it — who has actually read it. If the honest answer is "only the person who opened this pull request", the merge has just concentrated the team's model a little further into one head. That is a routing decision, not a blocking one: the fix is usually a second pair of eyes on the next change in that area, chosen deliberately.

4. What did it deliberately not handle?

Agents are good at scoping and bad at announcing the scope. The retry that covers timeouts but not 5xx, the migration that handles new rows but not existing ones, the endpoint that works for one tenant. All cheap to write in a sentence now, all expensive to rediscover from the outside.

5. Does the context file the agent loaded still tell the truth?

The CLAUDE.md or AGENTS.md in your repository is read into every session automatically and almost never re-checked, because it reads as authoritative. Ours described one of our own diagrams as roughly a thousand nodes for weeks. The real figure, measured against production: the current version is 75 nodes and 222 edges, and the largest across all 348 versions we have ever rendered is 77. That false line mis-scoped a piece of work before anyone checked it.

Once a month, pick one factual claim in that file and verify it. A stale context file does not fail loudly; it just quietly points work in the wrong direction.

6. If this breaks in six weeks, who can diagnose it without the session?

Answer with a name, not a role. If the only name is the person who opened it, and the session is gone after the squash merge, the answer is really "nobody" and you have just accepted that.

How to ask these without inventing paperwork

  • Put them in the pull request, not in a meeting. An answer written into the PR body is searchable by the next person. An answer given in standup is gone.
  • Make the agent write them at open time. The session that made the change still holds the reasoning. Our open-source hook does this by blocking gh pr create until the same agent writes the block itself — add-reasoning-to-prs, local, nothing uploaded.
  • Three of six is a good day. This is not a gate. A pull request that answers two questions well is better than one that answers six in template language.
  • Never score people on it. The moment answers become a performance measure, they become uniformly excellent and worthless. The value of question one is that "nothing was weighed" can be said out loud.
  • Re-read the answers when something breaks. That is the only test of whether the record is any good, and it takes one incident to find out.

Answered across a few hundred merges, these stop being a per-PR ritual and start being a map: which areas accumulate decisions with recorded reasoning, and which accumulate merges with blanks. Backthread builds that picture continuously from the sessions and the PR discussion, and shows it per area of the system, with the blanks left as blanks.

Connect one repo and you can see which of your recent merges carry any recorded reasoning at all, before you decide whether the habit is worth it; the trial runs fourteen days.

In short

The diff answers correctness, not comprehension
Reading lines catches the missing null check. The rejected approach, the assumption about a neighbouring service, and the case knowingly left unhandled never appear as lines, so they survive any amount of closer reading.
Ask what was decided that could have gone the other way
It is the highest-value question and frequently has no answer, because nothing was weighed. On our own repository, 62 percent of decisions captured from agent sessions record no alternative and no trade-off. Recording that honestly beats inventing a rationale afterwards.
Check the context file your agents load
An auto-loaded CLAUDE.md is injected into every session and read as authoritative. Ours claimed a diagram had around a thousand nodes when the largest ever rendered was 77. Verify one claim in it a month.
Keep it a habit, never a score
Answers written into the pull request are searchable later; answers given in a meeting are gone. Three questions answered honestly beat six answered in template language, and scoring people on them guarantees the second outcome.

Sources

  1. The Substrate Collapse: AI Code Generation Invalidates Authorship-Based Knowledge Metrics — Brett Wheeler, arXiv, June 2026
  2. Comprehension Debt: The Hidden Cost of AI-Generated Code — Addy Osmani, O'Reilly Radar, April 2026
  3. add-reasoning-to-prs — the open-source hook that writes the reasoning into the pull request

Backthread shows how much of what your agents built your team really understands. See how it works