Comparison

Backthread vs HumanLayer

One puts the human in before the code exists; the other keeps what the human decided after it ships.

The two act at opposite ends of one change. HumanLayer is a shared workspace for coding agents that walks a team through questions, research, design and plan before any code exists, the team commenting at each phase. Backthread starts where that ends: it captures the reasoning behind each merged change from the agent session, shows the leader how much of the system the team understands per area, and teaches it back in the coding agent.

They do not compete for the same minute of an engineer's day. This page is for the CTO trying to work out whether they need one, the other, or both, and in which order.

What is HumanLayer, and what does it believe?

HumanLayer is the company behind the 12-factor agents essay (26,000-plus stars on GitHub as of 2026-09-18, Apache 2.0 for the code) and the CodeLayer IDE that grew out of it. Its current product, on humanlayer.com as fetched on 2026-09-18, is "the multiplayer control plane for your software factory": a workspace where a local daemon runs several Claude Code, Codex, Copilot or Fireworks sessions in parallel, and where the team's plan artefacts, agent sessions and diffs sit side by side for comment. The original open-source repository (11,600 stars) now carries a note that most of its code is deprecated in favour of the rebuilt product.

The belief is the one Dex Horthy laid out in his July 2026 AI Engineer talk, Harness Engineering is not Enough: Why Software Factories Fail: coding models are rewarded for passing tests, not for keeping a system maintainable, so an unattended factory decays; his own company's lights-off experiment in July 2025 is the evidence he cites. The fix is to put the human back where one hour of thinking changes the most, which is before the code, and to have a person read what comes out. We agree with the diagnosis and wrote up where we think the remedy stops short.

The product encodes the fix as a workflow HumanLayer calls QRSPI: Questions, Research, Design, Structure, Plan, Implement. Each phase produces an artefact (a design doc, a phased structure, a plan with file paths and test cases) that teammates comment on inline before the agent moves to the next. The site's own summary is the line to remember: "Do not outsource the thinking. Every phase is a place to push back."

Where the two products part ways

Both start from the same fact, that agent-written code arrives faster than anyone's understanding of it, and each puts the human at a different point in the lifecycle.

HumanLayer puts the human earlier. The thinking happens in the design and plan documents, with the whole team in the room, before the agent touches a file. By the time a pull request exists it is already aligned with a plan somebody argued about, so a reviewer can read it quickly and with context. The reasoning is real and it is written down; it lives in the plan artefact.

Backthread keeps the human's reasoning after the merge, and asks a different question about it: not "was this planned well" but "who on the team still holds this, three months on". The decision, the alternatives that were weighed, the trade-off accepted and the assumption made are captured while the session that produced them still exists, held until the work merges so the record is of what shipped, and then attached to the area of the system they concern. The leader gets a map of coverage per area; the engineer gets the reasoning taught back inside the coding agent when they next touch that code.

The gap between the two is time. A plan document is a snapshot of what the room believed on the day. Three months later the code has moved, the doc has not, and the person who wrote it may be on another team. A well-run HumanLayer workflow produces one good review per change. It does not by itself tell you whether the twelve people who did not write the plan could explain the subsystem it describes. That second thing is what a per-area map of who understands what is for, and it is why a merged pull request is not evidence of understanding even when the pull request came out of a good plan.

There is a pleasant consequence for teams that run both: HumanLayer's design and plan artefacts are exactly the kind of reasoning Backthread captures. A plan that was argued about, then implemented and merged, is the best possible input to a decision record; without something downstream it produces one review and then goes quiet.

Backthread vs HumanLayer: the comparison table

HumanLayerBackthread
What it is forGetting a team's thinking in front of the agent before it codes, and running many agent sessions in one shared workspaceRecording the reasoning behind merged changes and showing who on the team understands which part of the system
When in the lifecycle it actsBefore and during the change: questions, research, design, structure, plan, implementAfter the merge: a decision is held until its work ships, then attached to the area it concerns
What the engineer doesWorks inside the workspace; writes and comments on plan artefacts; reviews the diff that followsNothing extra; reasoning is captured from the agent session and PR discussion, and taught back in the flow of work
What the leader seesThe team's sessions, artefacts and diffs in one place; audit logs on EnterpriseA live map of the system with knowledge coverage per area, estimate first, earned above it
What is measuredNothing about people's understanding; the artefacts are the outputCoverage per area: inferred from git and PR history and capped at 50%, then earned by explaining decisions
PricingStarter free for up to 3 members and 200 sessions a month; Pro $100 per user per month; Enterprise custom (fetched 2026-09-18)14-day trial, then $25 per seat per month, one repository included, +$10 per additional repository
Security postureSOC 2 Type II; BYOK for model subscriptions and keys; ZDR and DPA on Enterprise; on-prem and private VPC options listedSource stripped on the engineer's machine before anything is sent; analysis in a sandbox destroyed after each job; source never stored; redaction library is open source

Two rows need a note. On what is measured, Backthread's first picture is deliberately modest: having touched an area is not understanding it, so the inferred estimate is capped at 50 percent and labelled as such, and an area with nothing on record says "nothing on record here" rather than showing a zero. On security, HumanLayer's claims are theirs from their pricing page; ours are that the redaction fence is public so anyone can read what leaves the machine, with the rest at backthread.dev/security. How capture, the map and in-agent teaching fit together is on the product page.

How do the prices compare?

They price different things, so a single multiplier would mislead.

HumanLayer, as published on 2026-09-18: a free Starter tier for teams of up to three members and 200 sessions a month; Pro at $100 per user per month with unlimited sessions, multi-repo workspaces, remote daemons and real-time collaboration; Enterprise on request with SSO, audit logs and private deployment options. There is no per-token billing; you bring your own Claude, Codex or other subscription. Third-party write-ups from June 2026 describe the product as waitlisted and macOS/Linux only at the time; check the current state on their site, as that may have changed since.

Backthread is $25 per seat per month after a fourteen-day trial, one repository included, $10 per additional repository. No free tier, no usage meter, no enterprise tier.

A 20-engineer team pays $2,000 a month for HumanLayer Pro plus its existing agent subscriptions, and $500 a month for Backthread plus repositories. But the HumanLayer number buys the place the team does its work all day; the Backthread number buys a record and a map alongside whichever agent the team already uses. If your engineers are going to live in HumanLayer's workspace, its price is a share of the tooling budget; Backthread is an addition to it in either case.

When HumanLayer is the better choice

Choose HumanLayer, and be glad of it, when:

  • You want a structured planning workflow inside the agent, and you want the team in it. QRSPI is opinionated in a good way. If your engineers are running agents from bare terminals with no shared design step, HumanLayer gives them one, with inline comments from teammates before a line is written. Nothing in Backthread does this and nothing in it is meant to.
  • Everyone still reads every line. If your team is small enough, or disciplined enough, that a person reads every diff that merges, then the understanding is being rebuilt in each reviewer's head as they go. A coverage map will mostly tell you what you already know. Horthy's remedy works for that team, and HumanLayer is built for it.
  • Your problem is running many agent sessions at once, across repositories and worktrees, without losing track. That is a workspace problem and HumanLayer is a workspace.
  • You are three people. Starter is free at that size, and at three engineers the whole system fits in one conversation.

Choose Backthread when the reading has stopped keeping up. The signal is the one described in what to do when review cannot keep up with agent pull requests: approvals that nobody fully understands, a subsystem whose model lives in one head, a leader who cannot say which parts of the system the team could rebuild from memory. HumanLayer improves the quality of each change going in. Backthread tells you what the team holds of what came out.

Can you run both?

Yes, and the two fit better than most pairings on this site. HumanLayer's phases produce written reasoning before the code; Backthread's capture step is what keeps reasoning after the merge and puts it on a map. A plan document that was argued about in HumanLayer, implemented, and merged is precisely the material that turns into a decision with recorded alternatives and trade-offs. On our own repository most agent-written decisions carried no recorded alternative at all, because nobody weighed one; a team that plans the way HumanLayer prescribes would have a far richer record to capture. What neither tool does alone is close the loop: the plan is written once and read once, and the record is only useful if it is captured. Together, the plan gives you a good change and the record tells you, per area, who could still explain it.

Connect one repo and the first coverage map is drawn from your git and PR history within the hour, capped at an estimate until the team's recorded decisions raise it. The trial runs fourteen days with everything on, and if your team really is reading every line, the map will say so and you can stop there.

In short

HumanLayer acts before the code; Backthread acts after the merge.
HumanLayer is a shared workspace that walks a team through questions, research, design, structure and plan before the agent implements, with people commenting at each phase. Backthread captures the reasoning behind each change once it merges, maps knowledge coverage per area, and teaches that reasoning back inside the coding agent.
A good plan gives you one good review, not a team that holds the system.
The plan artefact records what the room believed on the day. It is not a record of what the rest of the team has retained about that subsystem months later. That is the gap Backthread measures, and Anthropic's trial (50 percent versus 67 percent on comprehension, measured shortly after the task) suggests it widens under AI assistance.
The prices buy different things.
HumanLayer is free for up to three members and 200 sessions a month, then $100 per user per month, with your own model subscriptions brought in, as published on 2026-09-18. Backthread is $25 per seat per month, one repository included, $10 per extra repository. One is the workspace the team works in; the other is a record and a map alongside whatever agent it already uses.
HumanLayer is the better choice for a team that wants structured planning inside the agent and still reads every line.
If every diff that merges is read by a person, understanding is being rebuilt as you go and a coverage map adds little. Backthread is for the 10 to 100-engineer team where that stopped being true.
Running both is coherent, and the order is HumanLayer first.
The plan documents HumanLayer produces are exactly the reasoning Backthread captures at merge. Plan the change there; keep and measure what the team holds of it here.

Sources

  1. HumanLayer — Pricing (fetched 2026-09-18)
  2. HumanLayer — home page, "the multiplayer control plane for your software factory" and the QRSPI workflow (fetched 2026-09-18)
  3. Dex Horthy — Harness Engineering is not Enough: Why Software Factories Fail, AI Engineer, July 2026
  4. humanlayer/12-factor-agents on GitHub (26.2k stars as of 2026-09-18)
  5. humanlayer/humanlayer on GitHub (11.6k stars; README marks most of the code deprecated, as of 2026-09-18)
  6. Pane — Pane vs HumanLayer (June 2026): waitlist status and platform support at the time
  7. Anthropic — How AI assistance impacts the formation of coding skills (2026-01-29)

Backthread shows how much of what your agents built your team really understands. See how it works