Agents coordinate through explicit, mechanically checkable interface specifications rather than through shared context or prose handoff. The contract is the coordination primitive, and it is the only thing two agents are required to agree on.
Humans working this way get a nice-to-have. Agents get a necessity, for one reason: there is no shared working memory. Each agent's context is private, finite, and lossy. Two humans at a whiteboard hold a common picture; two agents hold two pictures and no way to diff them.
the only shared surface │ agent A ▼ agent B ┌─────────┐ ┌──────────────┐ ┌─────────┐ │ private │ │ CONTRACT │ │ private │ │ context │◄──────►│ schema │◄──────►│ context │ │ │ │ invariants │ │ │ └─────────┘ │ examples │ └─────────┘ └──────────────┘ never merged checkable by both never merged without talking
| Property | Why it follows from the contract |
|---|---|
| Parallelism | Producer and consumer are built at the same time. Neither waits on the other's implementation, only on the agreement. |
| Verifiability | Conformance is a test, not a judgement. No reviewer opinion required. |
| Replaceability | Any agent can be swapped, restarted, or run on a different model. The contract survives the worker. |
| Bounded context | The subagent is handed the contract, not the repository. Context spend drops from the whole tree to one interface. |
Everything downstream in this reference follows from that one test. Section 03 grades contract kinds by whether they pass it.
Worth knowing before pitching it, because someone in the room will have shipped this in 2009 under a different name. The idea is forty years old. What changed is the constraint, not the technique.
| Origin | Year | What it contributed |
|---|---|---|
| Conway's Law | 1967 | Melvin Conway: system structure mirrors the communication structure of the org that built it. Applies unchanged to agent topology -- fan out by service and you get service-shaped code, whether or not that is the right architecture. |
| Design by Contract | 1986 | Bertrand Meyer, Eiffel. Preconditions, postconditions, class invariants as first-class language constructs. The vocabulary everything since uses. |
| N-version programming | 1977 | Chen and Avizienis. Build it independently N times, compare. See section 06 for the 1986 result that broke it, which is the single most important caveat in multi-agent verification. |
| Consumer-driven contracts | 2006 | Ian Robinson. The consumer states what it needs; the producer proves it still satisfies every consumer. Inverts who owns the spec. Pact (2013) is the tooling. |
| Property-based testing | 2000 | QuickCheck (Claessen and Hughes). Contracts as generators plus invariants rather than as examples. The closest thing to a machine-checkable behavioural contract in wide use. |
| Formal specification | 1999 | TLA+ (Lamport), Alloy (Jackson). Contracts strong enough to model-check. Expensive, and occasionally the only thing that works. |
Contracts form a ladder. Each rung catches more and costs more to write. The useful skill is picking the lowest rung that catches the failure you actually fear.
weaker, cheaper stronger, dearer ├──────────────────────────────────────────────────────────────┤ 1 type 2 schema 3 behavioural 4 property 5 formal signature JSON Schema pre/post invariants TLA+ fn(A) -> B required, "nonempty in, "sorted for model-checked enum, range sorted out" ALL inputs" state space catches: catches: catches: catches: catches: wrong shape bad values wrong logic edge cases concurrency you did not interleavings think of nobody would
| Rung | Machine-checkable | Good for | Cost |
|---|---|---|---|
| Type signature | yes, at compile time | Structural agreement. The floor, not a strategy. | free |
| Schema | yes, at runtime and in CI | The workhorse. JSON Schema, Protobuf, Pydantic. Enums and ranges kill a surprising share of integration bugs. | minutes |
| Behavioural | yes, as assertions | Preconditions and postconditions. Encodes intent the type cannot. | hours |
| Property | yes, by generation | Invariants over generated inputs. Finds cases no author enumerates. | hours to days |
| Formal | yes, exhaustively | Concurrency, consensus, anything where the bug is an interleaving. Rarely worth it. Occasionally the only option. | weeks |
Contracts let you parallelize, but only across the seams they isolate. Split anywhere else and you have bought coordination cost without buying independence.
CONTRACT-FIRST FAN-OUT the serial part is the agreement, and it is not compressible ┌── negotiate the contract ──┐ serial │ schema + invariants + │ one agent, or a human │ one valid, one invalid │ └─────────────┬──────────────┘ │ ┌─────────┼─────────┐ parallel ▼ ▼ ▼ n agents, no talking producer consumer conformance impl impl tests │ │ │ └─────────┼─────────┘ ▼ integrate: the tests serial again already exist and neither side wrote them for itself
Writing the conformance tests in a third agent is the part people drop, and it is the part that makes the pattern work. A producer that writes its own tests tests what it built. Section 06.
| Split | Parallelizes when | Fails when |
|---|---|---|
| By interface | The seam is a real API, queue, or file format. Contract is obvious and narrow. | Rarely. This is the good case. |
| By layer | Layers are genuinely independent and the contract between them is stable. | The contract changes as you learn, so every agent redoes work. Common. |
| By feature | Features touch disjoint files. | They never do. Merge conflicts become the dominant cost. |
| By review dimension | Always. Each reviewer reads the same diff for a different failure class. | Never fails, and it is the cheapest real win. Start here. |
Note also that Conway's Law runs in this direction: the agent topology you pick becomes the architecture you get. Fan out by microservice and you will be handed microservices whether or not that was the design.
The heart of it. Each style is a lens tuned to a different failure class, and the reason to name them is that a reviewer told "review this" defaults to style checking. A reviewer told "assume it is broken and find out how" does something else entirely.
WHICH LENS FOR WHICH FAILURE the design is the plan will the group is wrong fail agreeing too fast │ │ │ ▼ ▼ ▼ adversarial pre-mortem devil's advocate red team fault tree dialectic inversion FMEA wideband delphi we do not know it broke and we we are about to why this exists need the cause delete something │ │ │ ▼ ▼ ▼ socratic five whys chesterton's fence rubber duck blameless PM
The reviewer's job is to break it, not improve it. Brief is "assume the author is wrong; produce the input that proves it". Structurally different from ordinary review because success is a counterexample, not a comment.
Catches: wrong designs, missing validation, happy-path thinking. Trap: unbounded adversaries produce unfalsifiable objections. Bound it: "within the documented threat model" or "reachable from a real caller".
Gary Klein, HBR 2007. Before starting, state that the project has already failed completely, then work backwards to explain why. The mechanism is prospective hindsight: Mitchell, Russo and Pennington (1989) found that imagining an outcome as certain rather than possible raises the number of causes people can generate by roughly thirty percent.
Catches: plan-level risk, optimism bias, the thing everyone privately doubted and nobody said. Trap: it is not risk assessment. Do not let it become a probability-weighted register. The framing "it failed, why" is the active ingredient, and softening it removes the effect.
Literally the advocatus diaboli, the office of Promoter of the Faith established 1587 to argue against canonization; the role was cut back in 1983. One party is assigned to argue the opposing case regardless of belief.
Catches: premature consensus. Trap, and it is a serious one: Charlan Nemeth's work found that authentic minority dissent produces better divergent thinking than a known-assigned advocate. Everyone discounts the person role-playing. For agents this partly dissolves, since there is no social cost to disagreeing and no visible name badge -- but an agent told "argue against this" also has no genuine belief, so treat its output as a checklist of objections rather than as evidence anyone is actually unconvinced.
Strengthen the opposing position to its best form before attacking it. The inverse of a strawman.
Catches: arguing against a weak version of the alternative, which is how teams talk themselves into the thing they already wanted. Use it when rejecting an approach, not when accepting one.
Carl Jacobi's "man muss immer umkehren", invert, always invert; popularized by Munger. Instead of "how do we make this work", ask "what would guarantee this fails", then check whether you are doing any of it.
Catches: the difference between a plan with no known flaws and a plan that has been attacked. Cheap, fast, and unreasonably effective as a first pass.
Sakichi Toyoda, Toyota Production System. Ask why five times, each answer feeding the next question.
Catches: proximate-cause fixation. Trap: it produces a single causal chain, and real incidents are trees. Use fault-tree or FMEA when there is more than one contributing branch, which there usually is.
Systematic enumeration rather than inspiration. FMEA walks every component and asks how it fails, how that is detected, and what it costs. Fault tree starts at the undesired event and decomposes backwards through AND and OR gates.
Catches: failure modes nobody would think to name. Cost: tedious, which is precisely why it is a good thing to hand an agent.
Chesterton, 1929. Do not remove what you cannot explain the purpose of. State the reason the thing exists before proposing its deletion.
Catches: agents confidently deleting load-bearing code that looks redundant. This is a live failure mode, not a theoretical one, and it is worth a standing rule rather than a review pass.
Edward de Bono, 1985. Parallel thinking: the whole group wears one hat at a time -- facts, feelings, caution, benefits, creativity, process -- rather than each person defending a position.
Catches: debate-as-identity, where changing your mind is a loss. For agents: maps cleanly onto one pass per hat over the same artifact, and the caution pass is the one that pays.
Barry Boehm, from RAND's Delphi method. Reviewers assess independently and without seeing each other, then the spread is examined and only then discussed.
Catches: anchoring and cascade. Why it matters here more than anywhere: the moment agent B sees agent A's review, B's independence is gone. Sequential review is contaminated review. Fan out blind, converge after.
Thesis, antithesis, synthesis. Two agents argue opposing positions; a third, which has not argued, synthesizes.
Catches: false binaries. Requirement: the synthesizer must not have participated. An arguer asked to synthesize will find its own position persuasive.
Interrogate the assumptions rather than the conclusion. "What must be true for this to work?" then test each item.
Catches: unstated premises, which is where most bad designs actually live. Pairs well with pre-mortem: Socratic finds the assumptions, pre-mortem kills them.
John Allspaw and Etsy, around 2012. Separate learning from accountability, on the grounds that people asked to explain a failure they may be punished for will explain something else.
Catches: the real sequence of events. Note for agents: there is nobody to blame, so the discipline is free -- but the other half of the practice still applies: write down what was believed at the time, not what is known now.
Two lineages specific to models. Debate (Irving, Christiano and Amodei, 2018) has two agents argue to a judge, on the theory that lying is harder to defend than telling the truth. Critique-revise (Constitutional AI, Bai et al. 2022) has a model critique an output against written principles, then revise against the critique.
Trap for both: they inherit the judge's biases and the critic's blind spots. Section 06 is about exactly that.
You can only verify what you can check. The hard part of verification is almost never running the test; it is having something to compare against. That thing is the oracle, and most of the time you do not have one.
| When you have | Use | How it works |
|---|---|---|
| A known-correct answer | Assertion | The easy case. Rare outside toy problems. |
| A second implementation | Differential testing | McKeeman, 1998. Same input to both, diff the output. Any disagreement is a bug in one of them. |
| A known relationship | Metamorphic testing | Chen, 1998. No oracle needed. If sort(x) is right then sort(shuffle(x)) equals it; if a search is right, narrowing the query cannot add results. |
| An invariant | Property testing | Generate inputs, assert the invariant. The generator finds the edge case you would not. |
| Only judgement | Adversarial review | The weakest case. Bound it and get independent passes. Section 05. |
Metamorphic testing is the underused one. It applies exactly where agent output is hardest to check -- no ground truth, but plenty of known relationships between outputs. Most tasks have one if you look for it.
Not a matter of taste. Huang et al. (2023), Large Language Models Cannot Self-Correct Reasoning Yet, found that intrinsic self-correction -- revising with no external signal -- often degrades performance rather than improving it. The model has no new information at revision time, and its confidence is not calibrated to its correctness.
Model-as-judge carries its own documented biases. Zheng et al. (2023), the MT-Bench and Chatbot Arena work, identified position bias (order of presentation changes the verdict), verbosity bias (longer answers score higher independent of quality), and self-enhancement bias (judges prefer text from their own model).
N-version programming assumed independently built implementations fail independently, so a majority vote is safe. Knight and Leveson (1986) tested it and it is false. Twenty-seven independently developed programs, from separate teams, to the same specification, showed correlated failures -- because the hard parts of the problem are hard for everyone, and shared misreadings of the spec cluster.
WHAT INDEPENDENCE WAS ASSUMED TO LOOK LIKE agent A ████░░░░░░░░░░░░░░░░ fails here agent B ░░░░░░░░████░░░░░░░░ and here agent C ░░░░░░░░░░░░░░░░████ and here majority vote is safe everywhere WHAT KNIGHT-LEVESON FOUND, AND WHAT ONE MODEL FAMILY DOES agent A ░░░░░░████░░░░░░░░░░ agent B ░░░░░░███░░░░░░░░░░░ the hard part is agent C ░░░░░█████░░░░░░░░░ hard for all of them majority vote CONFIRMS the shared error
This transfers directly and it is worse for agents than for humans. Five agents on one base model are not five independent reviewers; they are one reviewer sampled five times. Agreement between them measures consistency, not correctness, and rising consensus feels exactly like rising confidence.
Named so they can be spotted in a standup rather than discovered in a post-mortem.
| Failure | What it looks like | Counter |
|---|---|---|
| Context telephone | Fidelity drops at every prose handoff. By hop four the brief has drifted and nobody can point at where. | Hand over the contract artifact, not a summary. Handoffs carry files, not paraphrase. |
| Sycophantic convergence | Agents agree with each other and with you. Confidence climbs, accuracy does not. | Blind independent passes before any sharing. Wideband Delphi, section 05. |
| Correlated failure | Five agents, one blind spot. The vote confirms the error. Section 06. | Vary model family, vary the lens, prefer a mechanical oracle. |
| Contract drift | The implementation moved, the contract did not. Everyone still passes tests written against last month's agreement. | Contract lives in the repo, in CI, and fails the build. Never in a doc. |
| The verification gap | Nobody checks the checker. A broken conformance test passes everything, silently. | Mutation testing, or one deliberately invalid fixture per contract that MUST fail. |
| Over-decomposition | Nine agents, six merge conflicts, and it would have been faster alone. | State the contract before splitting. If you cannot, do not split. |
| Confident wrong | Output is fluent, well-structured, and false. Fluency reads as competence. | Never accept a claim that is checkable without checking it. Fluency is not evidence. |
| Cost blowup | Fan-out multiplies spend linearly and quality sublinearly. | Fan out on review dimensions, which is cheap and pays. Be stingy fanning out on implementation. |
Almost all of it predates agents. That is the point -- the contract layer is mature and boring, which is exactly what you want load-bearing.
| Layer | Tools | Note |
|---|---|---|
| Contract testing | Pact, Schemathesis, Dredd | Pact is the consumer-driven original: the consumer records expectations, the provider is verified against every consumer's recording. |
| Interface schema | OpenAPI, AsyncAPI, Protobuf, JSON Schema, Avro | Pick one and generate both sides from it. Hand-written clients drift; generated ones cannot. |
| Runtime validation | Pydantic, zod, dataclasses, serde | The contract enforced at the boundary at runtime, not just in CI. |
| Property testing | Hypothesis, QuickCheck, fast-check, proptest | The highest value per hour of anything in this table for agent-written code. Generators find what enumeration misses. |
| Mutation testing | mutmut, Stryker, PIT | Answers "who checks the checker" mechanically. Deliberately breaks the code and asks whether any test noticed. |
| Formal | TLA+, Alloy | For interleavings and consensus. Expensive; occasionally the only thing that works. |
| Structured output | JSON-schema-constrained decoding, tool/function schemas | The contract enforced during generation rather than checked after. Turns "please return JSON" into a guarantee. |
| Tool contracts | MCP | A tool interface is a contract between an agent and a capability. Same discipline, same failure modes. |
| Isolation | git worktrees, containers | Parallel agents need separate filesystems or they fight. Cheap, and skipping it produces baffling bugs. |
| Code maps | call-graph indexers | Hand an agent a ranked map before it reads. Cuts orientation cost by an order of magnitude on code-heavy trees. |
The positions people are actually arguing from. Worth being able to name the one you hold, and the one the person disagreeing with you holds.
Small, composable, one job each, text at the boundaries. An agent that does one thing against a stated contract can be tested, replaced and reasoned about. An agent that "owns the feature" cannot. The oldest position here and still the strongest.
The direct analogue of least privilege: give an agent the minimum context the task requires. Not primarily an efficiency argument -- extra context is extra surface for the agent to be confused by, and it makes the failure harder to attribute afterwards.
The workflow is code: ordinary, versioned, reviewable, replayable. The steps inside it are agents. The argument is that you want exactly one source of nondeterminism, at a known place, rather than an emergent system whose control flow is itself a model output.
Postel's law says be liberal in what you accept. For contracts between agents, invert it: be strict, and fail loudly. Liberal acceptance hides drift until it is expensive, and the protocol-ossification critique of Postel's law applies double when the party on the other end will happily invent a plausible field name.
Since you cannot fully verify agent output, optimise for cheap undo instead: small commits, feature branches, idempotent operations, staged rollout. Correctness you cannot guarantee, recoverability you can.
Your agent topology becomes your architecture. Choosing how to fan out is an architectural decision, made early, usually implicitly, and usually by whoever wrote the orchestration script. Worth making on purpose.