한국어
0.95^20 was right arithmetic aimed at the wrong target: reliability is a property of architecture, not of the model.essay · standmeet2026.10.06 · essay
2026.10.06·2 min read#ai-predictions

Errors Compound, So Agents Can't Work

0.95^20 was right arithmetic aimed at the wrong target: reliability is a property of architecture, not of the model.

What was said

In 2023–2024, as the agent concept heated up, the hardest counter-argument was arithmetic: if a model is 95% reliable per step, a 20-step task succeeds 36% of the time; at 50 steps, 8%. Errors compound, so a model autonomously completing a coherent piece of work is mathematically untenable. The argument was quoted everywhere and sounded unanswerable — multiplication does not lie.

What actually happened

By 2025–2026, agents had become the industry's central form: coding agents work for hours inside real repositories and deliver running results. Not because per-step reliability reached 100% — it still hasn't — but because the whole system changed how it plays: verification layers (tool checks, tests and type systems intercept errors on the spot), retry and rollback (take a wrong step back and try another path), cutting long tasks short (subtasks accepted one by one), human checkpoints (stop at the critical step and wait for a person), and making the model check before it acts. Errors still compound, but behind every step stands a gate, and the compounding chain is cut into segments.

Why it didn't come true

The math was right; the static assumption was wrong. The 0.95^n formula presumes every step is independent, identically distributed, and uncorrected — it computes a bare model multiplied by itself. Engineering systems are never bare models: reliability is a property of the architecture, not of the model. The same arithmetic could "prove" that human organisations cannot complete any complex project: everyone makes mistakes, so a long enough chain must collapse. Organisations survive on review, acceptance and rework; agents survive on the same things, running at machine speed.

The unexpected part

This prophecy has a concrete counterexample: my own platform (StandMeet) is agent infrastructure — every agent action wrapped in capability boundaries, signed credentials, evidence gates and an evaluation line. The eight controls are not decoration; they are the engineering answer to the compounding argument: trust no single step, and cut the compounding off with gates. The people doing that arithmetic at the time did not foresee that the answer was not making steps perfect, but — as in all mature engineering — designing around imperfection.

ask the AI about this essay·context: “errors compound, so agents can't work”
›