Reviewing the plan is not reviewing the decisions

Approving the plan tells you what the agent was asked to do. It does not tell you what it decided while doing it, and that is where the trade-offs that shape a system get made. A plan review can feel like oversight and still miss every choice worth knowing about.

That gap is what Backthread works on: agents write the change, the reasoning that surfaced during the session and the pull-request discussion is captured, and a team can see which areas of the codebase now rest on choices nobody read.

The thread

On 24 September a Show HN for Whiteboard — an open-source canvas where a coding agent draws what it built, with the diagrams linked back to the code — reached 417 points and 141 comments. The launch text names the problem plainly: the team found it difficult to reason about which decisions their agents had made autonomously.

Most of the comments were about the tool. One was about where human attention should go, and that is the one worth arguing with:

So I think the main human interaction surfaces to target in the future will be in the planning process.

That is 2001zhaozhao, in the thread on 24 September 2026, having argued that coding agents are now reliable enough during implementation that there is no need to look at the resulting code beyond a glance. It is a widely held position with a tidy logic: if the plan is right and the implementation is faithful to it, then inspecting the implementation is redundant.

The reply came from one of Whiteboard's own founders, who had every commercial reason to agree that planning is the surface that matters:

i think reviewing a plan without an implementation doesn’t feel that useful anymore, at least to me, because key tradeoffs often only surface during implementation that effect the top-level spec.

That is sidharthkmenon, in the same thread. It is the sentence to keep from the week.

Why the plan is not where the decisions are

A plan is written before the cost of any option is known. It names an intent — put a cache in front of the lease broker — and the intent is usually the uncontroversial part. What determines whether the system stays intelligible is everything discovered afterwards: which key, what happens on a miss, what the change now assumes about ordering two services away, which failure mode was accepted because handling it properly meant a migration nobody had budgeted.

None of that is in the plan, because none of it was known when the plan was written. A limitation is found, not scheduled.

What plan review gives youWhat it cannot give you
The intent, stated before work beganThe alternative discovered and dropped mid-implementation
The scope everyone agreed toThe assumption the change now makes about a neighbouring service
A shared vocabulary for the changeThe trade-off accepted because the proper fix was out of scope
A record you can point at afterwardsAny indication of which of the three above actually occurred

The last row is the expensive one. Plan review produces an artefact that looks like evidence of understanding and is not, which is the same trap as treating a merge that way. A merged pull request is not evidence that anyone understood it, and an approved plan is weaker evidence still, because it was written before anyone knew anything.

The practitioner's version

Further down the thread, someone building a game described running exactly the upstream-only regime:

I have (like all of us I assume) tried moving forward over weeks of not code reviewing and only plan reviewing, and it’s amazing how badly things fell apart, I ended up having to reset weeks of work

That is jenniferhooley, on 25 September. One person's month is not a study. It is, however, precisely the failure the argument predicts: approve intents for long enough and the accumulated implementation-time choices are a system nobody has described.

What our own record says about it

If deliberation mostly happened at planning time, you would expect thin capture afterwards to be harmless. Our numbers do not support that reading. On our own repository, 1,613 of 4,302 decisions captured from agent sessions carry any recorded deliberation at all, so 62 percent record no alternative and no trade-off — frequently because nothing was ever weighed. The deliberated minority that does exist turns up in sessions and in pull-request discussion, both of which happen after the plan was signed off, not before.

So the honest version of the upstream argument is narrower than it sounds. Move planning effort upstream by all means. Do not expect the plan to tell you what was decided.

What to do with that

  • Keep plan review, and stop counting it as comprehension. It buys agreement on intent, which is worth having on its own terms. It is not a record of the system.
  • Ask for the delta, not the diff. The question after a merge is what turned out differently from the plan, and why. It is one question, and it is answerable while the session is still open and expensive to answer a month later.
  • Keep the answer per area, not per pull request. Nobody searches pull requests in March for something merged in October. They ask what this part of the system assumes, which is a question about the record you hold for each area.

Connect one repo and you can see, per area, which merges arrived with a recorded reason and which arrived with none; the trial runs fourteen days.

In short

A plan states intent, and intent is rarely the contested part
The choices that decide whether a system stays intelligible are made once the cost of each option is visible, which is during implementation. Which key, what happens on a miss, what the change now assumes two services away.
An approved plan is weaker evidence of understanding than a merge
Both get treated as proof that somebody held the change in their head. The plan is worse, because it was written before the work revealed anything, so it cannot contain the trade-off that was actually accepted.
Capture the delta between plan and implementation while the session is open
What turned out differently, and why, is one question with a cheap answer at merge time and an expensive reconstruction later. On our own repository, 62 percent of decisions captured from agent sessions record no alternative and no trade-off at all.

Sources

  1. Show HN: Whiteboard – An open-source IDE for thoughtful software design — Hacker News, 24 September 2026 (417 points, 141 comments)
  2. The exchange on plan review — sidharthkmenon replying to 2001zhaozhao, Hacker News, 24 September 2026
  3. devdotfast/whiteboard — the project's repository on GitHub

Backthread shows how much of what your agents built your team really understands. See how it works