简体中文
I used to think multi-agent parallelism was a solution shopping for a problem. The coordination tax is fixed, but the serial wait grows with every capability step — the default has quietly flipped.文章 · standmeet2026.10.07 · 文章
2026.10.07·5 分钟阅读#essay#agents

What Can't Be Done Serially?

I used to think multi-agent parallelism was a solution shopping for a problem. The coordination tax is fixed, but the serial wait grows with every capability step — the default has quietly flipped.

My old skepticism

For a long time I thought multi-agent parallelism was a solution shopping for a problem. What work is so urgent it can't wait in line? A serial agent has a virtue nobody puts on slides: one context that accumulates everything, where every step sees all that was learned before it. There is no information exchange to design, because there is nothing to exchange — the memory is total and free. Parallelism, by contrast, charges a coordination tax up front: split the task, isolate the contexts, define what each branch returns, reconcile the contradictions at the merge. And the merge itself is serial — someone has to read every branch's output before synthesizing, so the more branches you open, the heavier the one step you cannot parallelize. For small tasks the arithmetic is not close: the tax is fixed, the time saved is nearly zero.

I was right about all of that. I was wrong about it staying true.

What serial actually costs

The bill for running everything in series has four lines, and only the first is "slow." The first is wall-clock: a long run keeps a human waiting, or worse, teaches the human to stop asking. The second is context pressure: a serial run must carry everything it has ever seen in one window, and long windows degrade — the middle gets lost, compaction quietly eats detail. A parallel branch starts with a small, focused context. Parallelism is a context-management strategy that also happens to be faster. The third is restart cost: a serial run that dies at step forty of fifty starts over; branches fail independently and can be re-run one at a time, from the fork, not from the beginning. The fourth is path dependence: a serial run is greedy. An early wrong turn poisons everything downstream, and you never see the road not taken. Parallel branches can walk genuinely different routes and let a judge pick — that is parallelism buying quality, not just time.

The crossover

The reason the conclusion flips is that the two sides of the ledger scale differently. The coordination tax is roughly fixed: you design the branch interface once and reuse it. The serial wall-clock grows linearly with the size of the task you dare to delegate. And the size of the task grows with the model's capability — every capability step moves the unit of work from "answer this" to "research this," "audit this codebase," "run this evaluation." At the same time, rising per-step reliability lowers the classic risk of parallelism, branches wandering off in five directions at once. A fixed cost and a falling risk on one side; a linearly growing wait on the other. Crossing is not a matter of taste. It is a matter of time.

I have crossed it twice in my own systems without noticing at first. My evaluation harness runs hundreds of prompts in parallel because serially it would take days — so its driver is deliberately stateless, built so that N runs share nothing. And my visitor agent, whose serial crawl through a corpus made people wait a minute for an answer, is getting the same treatment one level down: the independent retrievals inside a single layer now fan out, and only the synthesis waits for all of them.

The exchange is the real design problem

What parallelism actually demands is not compute, it is an answer to: what do the branches say to each other? My rule, learned from building the other way first, is that branches should exchange artifacts, not conversation. Each branch writes durable output into a shared place — a blackboard, a store, a corpus — and the merge reads the artifacts, not the transcripts. Written records survive compaction, can be inspected, and can be checked by something other than the author's confidence. The conversational alternative — agents chatting their state to each other — scales the token bill and loses exactly the information the merge needs most.

And there is a wall behind this wall. Parallel generation meets serial verification: N branches produce N outputs that someone, or some eval, must check, and verification bandwidth is the scarcest resource in the whole system. Teams have already hit this with code review, drowning in generated changes until they built a second machine to review the first. How far you dare to parallelize is set by your gates — by whether your benchmarks can actually go red — not by how many branches you can spawn.

Where I still run in series

Depth cannot be widened. A reasoning chain where each step depends on the full result of the previous one has no layer to fan out; parallelism there is theater. Small tasks stay serial because the tax exceeds the prize. And anything whose branches need fine-grained shared mutable state will spend its savings on merge conflicts. But those are now the exceptions I have to argue for, one by one. The default flipped while I wasn't looking: the question used to be "why would this need more than one agent?" It is now "what about this task is so urgent it can't be parallel — and what exactly was I saving by keeping it in line?"

就这篇文章向 AI 提问·上下文:“what can't be done serially?”
›