How to onboard an engineer into a codebase that agents wrote
Start with the decisions, in the order they were made, and read them against a map. The usual plan of README, code tour, then ask whoever wrote it has a missing piece in a codebase agents wrote: there is no author to ask, and whoever merged it may never have held it.
Backthread was built around that absence: agents produce the change, the reasoning behind it is captured from the session and the pull-request discussion, and the result is a record a new engineer can read in merge order instead of interviewing people who were not there.
Why the standard onboarding plan stalls
The traditional plan has three legs, and in a codebase written largely by agents two of them are broken.
The README and the architecture document describe a system as it was understood at the moment someone sat down to write about it. Under agents, that moment is further in the past than it used to be, and the document rots at the rate the code moves rather than at the rate people revise prose. Ours carried a false claim for weeks: it described one of our own diagrams as roughly a thousand nodes when the largest version we have ever rendered, across all 348 of them, is 77. Nobody was careless. The file simply reads as authoritative and is therefore never re-checked.
The code tour survives, but it is expensive and it teaches the shape of the code rather than the shape of the decisions. A new engineer can follow a request through six files and still not know that the retry lives there because a queue two teams away is not ordered.
"Ask the person who wrote it" is the leg that has actually snapped. The change came out of a session that was discarded at merge. The reviewer approved a diff. Neither party is holding what the newcomer needs, which is why a merged pull request is not evidence that anyone understood it.
The substitute for an author
What replaces an author is not more documentation. It is the sequence of decisions that produced the current system, read in the order they were made, against a picture of where in the system each one landed.
Three properties make that sequence work where a document does not.
- It is ordered. A system is intelligible as a series of choices far more readily than as a snapshot. Reading fifty decisions in merge order gives you the shape of how the team thinks, which transfers to the fifty-first.
- It is attached to a place. A decision read against a map of the system tells the reader both what was chosen and which area now depends on that choice. Detached from a location, it is trivia.
- It is honest about absence. Most changes have no deliberation behind them, and a record that says so is more useful than one that pads. On our own repository, 62 percent of the decisions captured from agent sessions record no alternative and no trade-off at all. A newcomer reading a blank is learning something true: that change was cheap and reversible and nobody weighed anything.
The decisions mined from pull-request discussion rather than sessions do somewhat better on that measure — 213 of 508, or 42 percent, carry a weighed alternative — which is a reason to make the discussion itself part of what a newcomer reads.
The first two weeks
The plan below assumes a team shipping with agents daily, and one new engineer who will own an area within a quarter. Adjust the durations; do not reorder the stages, because each one supplies the vocabulary for the next.
| When | What they do | What they should be able to do at the end |
|---|---|---|
| Days 1–2 | Read the map of the system at its current version, top level only. No code. | Name the seven or eight areas and say which two talk to everything else |
| Days 3–5 | Read the last 40 to 60 recorded decisions in merge order, across all areas | Say what kind of change this team makes without deliberating, and what it argues about |
| Days 6–8 | Pick the single area they will own. Read every recorded decision in it, oldest first | Explain why the area is shaped the way it is, including two choices that could have gone the other way |
| Days 9–10 | Read the code of that area against the decisions they now hold | Point at the lines that exist because of a decision they read, and at lines nobody can account for |
| Days 11–14 | Ship one small change in the area, with an agent, and record the reasoning themselves | Answer a reviewer's question about the change without re-deriving it |
The order matters more than the timings. Map before decisions, because a decision without a location does not stick. Decisions before code, because the code answers "what" and the reader needs "why" first. Code before shipping, because the first change is where the model gets tested.
What to hand them on day one
- The map at the current version, not a drawing from a wiki. If your diagram is regenerated per commit and reshuffles every time, it teaches nobody anything; that is a fixable layout problem, not a fact about your system.
- A list of the areas with nobody assigned. The honest version of the org chart. A newcomer should know on day one which parts of the system are currently unowned, because that is where they are most likely to be sent.
- The context file your agents load. Have them read
CLAUDE.mdorAGENTS.mdcritically and flag anything they cannot verify. A newcomer is the last person in the company who will read that file with fresh eyes, and it is being injected into every agent session. - One incident write-up from the last quarter. Nothing conveys the real failure modes faster, and it gives the decisions they are about to read something to hang on.
- The name of one person per area, even if the answer is uncomfortable. If an area has no name against it, say so rather than leaving the newcomer to discover it in week six.
How to tell whether it worked
Set the test before they start, and make it a question rather than a checklist. At the end of the two weeks, ask them to explain one area to someone outside it: what it does, why it is shaped that way, what it assumes about its neighbours, and what it deliberately does not handle. Those are the same four things worth asking before approving any agent-written change, which is not a coincidence. Onboarding and review are the same comprehension problem viewed from different ends.
A second test, cheaper and slightly cruel: give them a bug in their area three weeks in and see whether they reach for the code or for the record. Reaching for the record is the behaviour you were trying to install.
What this does not fix
Reading decisions builds an accurate model of how the system got here. It does not build the muscle memory of having debugged the thing at three in the morning, and nothing except time does. Nor does it help where no record exists: an area that predates any capture of reasoning will still have to be read the hard way, line by line, and that is worth planning for explicitly rather than discovering in week two.
It also does not survive a team that treats the record as paperwork. If the reasoning is written to satisfy a process, a newcomer reading fifty entries learns only the house style of the template. The value of the blanks is that they are allowed to stay blank.
An engineer who can read a system's decisions in order, against a map that holds still between renders, is doing in two weeks what used to take a quarter of asking people who happened to be around. That picture — per area, with the gaps left visible — is what Backthread keeps for a team shipping with agents, and the onboarding case is only the most obvious thing to do with it.
Connect one repo and a new joiner can read the recorded reasoning for your busiest area on their first morning; the trial runs fourteen days.
In short
- There is no author to ask, so the decisions have to stand in for one
- The session that produced an agent-written change is discarded at merge and the reviewer only ever saw a diff. What a newcomer needs is the sequence of decisions that produced the current system, in the order they were made.
- Map first, decisions second, code third
- A decision without a location in the system does not stick, and code read before the reasoning answers "what" while leaving "why" unanswered. Reordering those stages is the most common way this plan fails.
- Blanks in the record are information, not a defect
- On our own repository 62 percent of decisions captured from agent sessions record no alternative and no trade-off. A newcomer reading a blank learns that the change was cheap and reversible, which is true and useful.
- Test comprehension by asking them to explain an area to an outsider
- What it does, why it is shaped that way, what it assumes about its neighbours, what it deliberately does not handle. A checklist of files read measures attendance; those four questions measure whether a model formed.
Sources
- The Substrate Collapse: AI Code Generation Invalidates Authorship-Based Knowledge Metrics — Brett Wheeler, arXiv, June 2026
- Comprehension Debt: The Hidden Cost of AI-Generated Code — Addy Osmani, O'Reilly Radar, April 2026
- AI Engineering Report 2026: The Acceleration Whiplash — Faros AI, 21 May 2026
Backthread shows how much of what your agents built your team really understands. See how it works