Contract-Driven Multi-Agent Development

Contracts · Decomposition · Adversarial Review · Pre-mortem · Oracles · Correlated Failure
coordination without shared memory // the idea is from 1986, the constraint is new // what to actually run, and what it costs

The Core Idea

Agents coordinate through explicit, mechanically checkable interface specifications rather than through shared context or prose handoff. The contract is the coordination primitive, and it is the only thing two agents are required to agree on.

Humans working this way get a nice-to-have. Agents get a necessity, for one reason: there is no shared working memory. Each agent's context is private, finite, and lossy. Two humans at a whiteboard hold a common picture; two agents hold two pictures and no way to diff them.

                 the only shared surface
                          │
   agent A                  ▼                  agent B
  ┌─────────┐        ┌──────────────┐        ┌─────────┐
  │ private │        │  CONTRACT    │        │ private │
  │ context │◄──────►│  schema      │◄──────►│ context │
  │         │        │  invariants  │        │         │
  └─────────┘        │  examples    │        └─────────┘
                      └──────────────┘
   never merged        checkable by both        never merged
                      without talking

What the contract buys

PropertyWhy it follows from the contract
ParallelismProducer and consumer are built at the same time. Neither waits on the other's implementation, only on the agreement.
VerifiabilityConformance is a test, not a judgement. No reviewer opinion required.
ReplaceabilityAny agent can be swapped, restarted, or run on a different model. The contract survives the worker.
Bounded contextThe subagent is handed the contract, not the repository. Context spend drops from the whole tree to one interface.
The load-bearing claim A contract that an agent cannot check mechanically is not a contract. It is a suggestion, and it will be interpreted differently by every agent that reads it. The test of a contract is whether conformance can be decided without a conversation.

Everything downstream in this reference follows from that one test. Section 03 grades contract kinds by whether they pass it.

Not New

Worth knowing before pitching it, because someone in the room will have shipped this in 2009 under a different name. The idea is forty years old. What changed is the constraint, not the technique.

OriginYearWhat it contributed
Conway's Law1967Melvin Conway: system structure mirrors the communication structure of the org that built it. Applies unchanged to agent topology -- fan out by service and you get service-shaped code, whether or not that is the right architecture.
Design by Contract1986Bertrand Meyer, Eiffel. Preconditions, postconditions, class invariants as first-class language constructs. The vocabulary everything since uses.
N-version programming1977Chen and Avizienis. Build it independently N times, compare. See section 06 for the 1986 result that broke it, which is the single most important caveat in multi-agent verification.
Consumer-driven contracts2006Ian Robinson. The consumer states what it needs; the producer proves it still satisfies every consumer. Inverts who owns the spec. Pact (2013) is the tooling.
Property-based testing2000QuickCheck (Claessen and Hughes). Contracts as generators plus invariants rather than as examples. The closest thing to a machine-checkable behavioural contract in wide use.
Formal specification1999TLA+ (Lamport), Alloy (Jackson). Contracts strong enough to model-check. Expensive, and occasionally the only thing that works.

So what is actually new

Framing that survives contact with a skeptic Do not pitch it as a new methodology. Pitch it as consumer-driven contract testing where the consumers happen to be agents, and the reason it matters more than it used to is that the coordination medium got worse: private lossy context instead of a shared whiteboard.

What A Contract Is

Contracts form a ladder. Each rung catches more and costs more to write. The useful skill is picking the lowest rung that catches the failure you actually fear.

  weaker, cheaper                                    stronger, dearer
  ├──────────────────────────────────────────────────────────────┤

  1 type        2 schema      3 behavioural   4 property     5 formal
  signature     JSON Schema   pre/post       invariants    TLA+
  fn(A) -> B    required,     "nonempty in,   "sorted for    model-checked
                enum, range   sorted out"     ALL inputs"    state space

  catches:      catches:      catches:       catches:      catches:
  wrong shape   bad values    wrong logic    edge cases    concurrency
                                             you did not   interleavings
                                             think of      nobody would
RungMachine-checkableGood forCost
Type signatureyes, at compile timeStructural agreement. The floor, not a strategy.free
Schemayes, at runtime and in CIThe workhorse. JSON Schema, Protobuf, Pydantic. Enums and ranges kill a surprising share of integration bugs.minutes
Behaviouralyes, as assertionsPreconditions and postconditions. Encodes intent the type cannot.hours
Propertyyes, by generationInvariants over generated inputs. Finds cases no author enumerates.hours to days
Formalyes, exhaustivelyConcurrency, consensus, anything where the bug is an interleaving. Rarely worth it. Occasionally the only option.weeks

The three parts worth writing every time

Prose in a contract is a defect "Handles errors gracefully" is not a contract term. Neither is "reasonable performance". If an agent cannot turn the clause into an assertion, it will turn it into an assumption, and two agents will assume differently. Either make it checkable or move it out of the contract and into the notes.

Decomposition

Contracts let you parallelize, but only across the seams they isolate. Split anywhere else and you have bought coordination cost without buying independence.

  CONTRACT-FIRST FAN-OUT
  the serial part is the agreement, and it is not compressible

  ┌── negotiate the contract ──┐            serial
  │  schema + invariants +     │            one agent, or a human
  │  one valid, one invalid    │
  └─────────────┬──────────────┘
                │
      ┌─────────┼─────────┐                      parallel
      ▼         ▼         ▼                      n agents, no talking
  producer  consumer  conformance
  impl      impl      tests
      │         │         │
      └─────────┼─────────┘
                ▼
        integrate: the tests                      serial again
        already exist and
        neither side wrote
        them for itself

Writing the conformance tests in a third agent is the part people drop, and it is the part that makes the pattern work. A producer that writes its own tests tests what it built. Section 06.

Slicing

SplitParallelizes whenFails when
By interfaceThe seam is a real API, queue, or file format. Contract is obvious and narrow.Rarely. This is the good case.
By layerLayers are genuinely independent and the contract between them is stable.The contract changes as you learn, so every agent redoes work. Common.
By featureFeatures touch disjoint files.They never do. Merge conflicts become the dominant cost.
By review dimensionAlways. Each reviewer reads the same diff for a different failure class.Never fails, and it is the cheapest real win. Start here.
The over-decomposition trap Coordination cost grows with the number of agents; the work does not shrink as fast. Three agents with a clean contract beat nine with a muddy one, every time. Before fanning out, ask what the contract between the pieces actually is. If you cannot state it in a schema, you have found the reason not to split there.

Note also that Conway's Law runs in this direction: the agent topology you pick becomes the architecture you get. Fan out by microservice and you will be handed microservices whether or not that was the design.

Styles Of Review

The heart of it. Each style is a lens tuned to a different failure class, and the reason to name them is that a reviewer told "review this" defaults to style checking. A reviewer told "assume it is broken and find out how" does something else entirely.

  WHICH LENS FOR WHICH FAILURE

  the design is          the plan will        the group is
  wrong                  fail                 agreeing too fast
       │                      │                    │
       ▼                      ▼                    ▼
  adversarial            pre-mortem           devil's advocate
  red team               fault tree           dialectic
  inversion              FMEA                 wideband delphi

  we do not know         it broke and we      we are about to
  why this exists        need the cause       delete something
       │                      │                    │
       ▼                      ▼                    ▼
  socratic               five whys            chesterton's fence
  rubber duck            blameless PM

Adversarial review / red team

The reviewer's job is to break it, not improve it. Brief is "assume the author is wrong; produce the input that proves it". Structurally different from ordinary review because success is a counterexample, not a comment.

Catches: wrong designs, missing validation, happy-path thinking. Trap: unbounded adversaries produce unfalsifiable objections. Bound it: "within the documented threat model" or "reachable from a real caller".

Pre-mortem

Gary Klein, HBR 2007. Before starting, state that the project has already failed completely, then work backwards to explain why. The mechanism is prospective hindsight: Mitchell, Russo and Pennington (1989) found that imagining an outcome as certain rather than possible raises the number of causes people can generate by roughly thirty percent.

Catches: plan-level risk, optimism bias, the thing everyone privately doubted and nobody said. Trap: it is not risk assessment. Do not let it become a probability-weighted register. The framing "it failed, why" is the active ingredient, and softening it removes the effect.

Devil's advocate

Literally the advocatus diaboli, the office of Promoter of the Faith established 1587 to argue against canonization; the role was cut back in 1983. One party is assigned to argue the opposing case regardless of belief.

Catches: premature consensus. Trap, and it is a serious one: Charlan Nemeth's work found that authentic minority dissent produces better divergent thinking than a known-assigned advocate. Everyone discounts the person role-playing. For agents this partly dissolves, since there is no social cost to disagreeing and no visible name badge -- but an agent told "argue against this" also has no genuine belief, so treat its output as a checklist of objections rather than as evidence anyone is actually unconvinced.

Steelman

Strengthen the opposing position to its best form before attacking it. The inverse of a strawman.

Catches: arguing against a weak version of the alternative, which is how teams talk themselves into the thing they already wanted. Use it when rejecting an approach, not when accepting one.

Inversion

Carl Jacobi's "man muss immer umkehren", invert, always invert; popularized by Munger. Instead of "how do we make this work", ask "what would guarantee this fails", then check whether you are doing any of it.

Catches: the difference between a plan with no known flaws and a plan that has been attacked. Cheap, fast, and unreasonably effective as a first pass.

Five Whys

Sakichi Toyoda, Toyota Production System. Ask why five times, each answer feeding the next question.

Catches: proximate-cause fixation. Trap: it produces a single causal chain, and real incidents are trees. Use fault-tree or FMEA when there is more than one contributing branch, which there usually is.

FMEA and fault tree analysis

Systematic enumeration rather than inspiration. FMEA walks every component and asks how it fails, how that is detected, and what it costs. Fault tree starts at the undesired event and decomposes backwards through AND and OR gates.

Catches: failure modes nobody would think to name. Cost: tedious, which is precisely why it is a good thing to hand an agent.

Chesterton's Fence

Chesterton, 1929. Do not remove what you cannot explain the purpose of. State the reason the thing exists before proposing its deletion.

Catches: agents confidently deleting load-bearing code that looks redundant. This is a live failure mode, not a theoretical one, and it is worth a standing rule rather than a review pass.

Six Thinking Hats

Edward de Bono, 1985. Parallel thinking: the whole group wears one hat at a time -- facts, feelings, caution, benefits, creativity, process -- rather than each person defending a position.

Catches: debate-as-identity, where changing your mind is a loss. For agents: maps cleanly onto one pass per hat over the same artifact, and the caution pass is the one that pays.

Wideband Delphi

Barry Boehm, from RAND's Delphi method. Reviewers assess independently and without seeing each other, then the spread is examined and only then discussed.

Catches: anchoring and cascade. Why it matters here more than anywhere: the moment agent B sees agent A's review, B's independence is gone. Sequential review is contaminated review. Fan out blind, converge after.

Dialectic

Thesis, antithesis, synthesis. Two agents argue opposing positions; a third, which has not argued, synthesizes.

Catches: false binaries. Requirement: the synthesizer must not have participated. An arguer asked to synthesize will find its own position persuasive.

Socratic questioning

Interrogate the assumptions rather than the conclusion. "What must be true for this to work?" then test each item.

Catches: unstated premises, which is where most bad designs actually live. Pairs well with pre-mortem: Socratic finds the assumptions, pre-mortem kills them.

Blameless post-mortem

John Allspaw and Etsy, around 2012. Separate learning from accountability, on the grounds that people asked to explain a failure they may be punished for will explain something else.

Catches: the real sequence of events. Note for agents: there is nobody to blame, so the discipline is free -- but the other half of the practice still applies: write down what was believed at the time, not what is known now.

Debate and critique-revise

Two lineages specific to models. Debate (Irving, Christiano and Amodei, 2018) has two agents argue to a judge, on the theory that lying is harder to defend than telling the truth. Critique-revise (Constitutional AI, Bai et al. 2022) has a model critique an output against written principles, then revise against the critique.

Trap for both: they inherit the judge's biases and the critic's blind spots. Section 06 is about exactly that.

The one rule that matters more than the style Whoever built it does not grade it. Self-review shares the blind spot that caused the defect -- that is what a blind spot is. Every style above is weakened, and some are neutralized entirely, when the reviewer and the author are the same agent.

The Oracle Problem

You can only verify what you can check. The hard part of verification is almost never running the test; it is having something to compare against. That thing is the oracle, and most of the time you do not have one.

When you haveUseHow it works
A known-correct answerAssertionThe easy case. Rare outside toy problems.
A second implementationDifferential testingMcKeeman, 1998. Same input to both, diff the output. Any disagreement is a bug in one of them.
A known relationshipMetamorphic testingChen, 1998. No oracle needed. If sort(x) is right then sort(shuffle(x)) equals it; if a search is right, narrowing the query cannot add results.
An invariantProperty testingGenerate inputs, assert the invariant. The generator finds the edge case you would not.
Only judgementAdversarial reviewThe weakest case. Bound it and get independent passes. Section 05.

Metamorphic testing is the underused one. It applies exactly where agent output is hardest to check -- no ground truth, but plenty of known relationships between outputs. Most tasks have one if you look for it.

Why self-verification is weak

Not a matter of taste. Huang et al. (2023), Large Language Models Cannot Self-Correct Reasoning Yet, found that intrinsic self-correction -- revising with no external signal -- often degrades performance rather than improving it. The model has no new information at revision time, and its confidence is not calibrated to its correctness.

Model-as-judge carries its own documented biases. Zheng et al. (2023), the MT-Bench and Chatbot Arena work, identified position bias (order of presentation changes the verdict), verbosity bias (longer answers score higher independent of quality), and self-enhancement bias (judges prefer text from their own model).

Three cheap mitigations, all of them mechanical Swap the presentation order and require the verdict to survive it. Strip formatting and length cues before judging. Have a different model family judge than the one that produced the work. None of these need cleverness, and skipping them is how a review pipeline becomes theatre.

The 1986 result that everyone rediscovers

N-version programming assumed independently built implementations fail independently, so a majority vote is safe. Knight and Leveson (1986) tested it and it is false. Twenty-seven independently developed programs, from separate teams, to the same specification, showed correlated failures -- because the hard parts of the problem are hard for everyone, and shared misreadings of the spec cluster.

  WHAT INDEPENDENCE WAS ASSUMED TO LOOK LIKE

  agent A  ████░░░░░░░░░░░░░░░░  fails here
  agent B  ░░░░░░░░████░░░░░░░░  and here
  agent C  ░░░░░░░░░░░░░░░░████  and here
           majority vote is safe everywhere

  WHAT KNIGHT-LEVESON FOUND, AND WHAT ONE MODEL FAMILY DOES

  agent A  ░░░░░░████░░░░░░░░░░
  agent B  ░░░░░░███░░░░░░░░░░░   the hard part is
  agent C  ░░░░░█████░░░░░░░░░   hard for all of them
           majority vote CONFIRMS the shared error

This transfers directly and it is worse for agents than for humans. Five agents on one base model are not five independent reviewers; they are one reviewer sampled five times. Agreement between them measures consistency, not correctness, and rising consensus feels exactly like rising confidence.

Failure Modes

Named so they can be spotted in a standup rather than discovered in a post-mortem.

FailureWhat it looks likeCounter
Context telephoneFidelity drops at every prose handoff. By hop four the brief has drifted and nobody can point at where.Hand over the contract artifact, not a summary. Handoffs carry files, not paraphrase.
Sycophantic convergenceAgents agree with each other and with you. Confidence climbs, accuracy does not.Blind independent passes before any sharing. Wideband Delphi, section 05.
Correlated failureFive agents, one blind spot. The vote confirms the error. Section 06.Vary model family, vary the lens, prefer a mechanical oracle.
Contract driftThe implementation moved, the contract did not. Everyone still passes tests written against last month's agreement.Contract lives in the repo, in CI, and fails the build. Never in a doc.
The verification gapNobody checks the checker. A broken conformance test passes everything, silently.Mutation testing, or one deliberately invalid fixture per contract that MUST fail.
Over-decompositionNine agents, six merge conflicts, and it would have been faster alone.State the contract before splitting. If you cannot, do not split.
Confident wrongOutput is fluent, well-structured, and false. Fluency reads as competence.Never accept a claim that is checkable without checking it. Fluency is not evidence.
Cost blowupFan-out multiplies spend linearly and quality sublinearly.Fan out on review dimensions, which is cheap and pays. Be stingy fanning out on implementation.
The failure that eats the others A review pipeline that never returns "this is wrong" is not working, it is decorative. If every adversarial pass comes back with polish suggestions, the brief is too soft or the reviewer is grading its own work. Track how often review changes the outcome. If that number is near zero, the pipeline is a cost with no product.

Tooling

Almost all of it predates agents. That is the point -- the contract layer is mature and boring, which is exactly what you want load-bearing.

LayerToolsNote
Contract testingPact, Schemathesis, DreddPact is the consumer-driven original: the consumer records expectations, the provider is verified against every consumer's recording.
Interface schemaOpenAPI, AsyncAPI, Protobuf, JSON Schema, AvroPick one and generate both sides from it. Hand-written clients drift; generated ones cannot.
Runtime validationPydantic, zod, dataclasses, serdeThe contract enforced at the boundary at runtime, not just in CI.
Property testingHypothesis, QuickCheck, fast-check, proptestThe highest value per hour of anything in this table for agent-written code. Generators find what enumeration misses.
Mutation testingmutmut, Stryker, PITAnswers "who checks the checker" mechanically. Deliberately breaks the code and asks whether any test noticed.
FormalTLA+, AlloyFor interleavings and consensus. Expensive; occasionally the only thing that works.
Structured outputJSON-schema-constrained decoding, tool/function schemasThe contract enforced during generation rather than checked after. Turns "please return JSON" into a guarantee.
Tool contractsMCPA tool interface is a contract between an agent and a capability. Same discipline, same failure modes.
Isolationgit worktrees, containersParallel agents need separate filesystems or they fight. Cheap, and skipping it produces baffling bugs.
Code mapscall-graph indexersHand an agent a ranked map before it reads. Cuts orientation cost by an order of magnitude on code-heavy trees.
If you adopt exactly two things Property-based testing and mutation testing. The first generates the cases nobody enumerated; the second proves the tests can actually fail. Together they close the verification gap without any agent-specific machinery at all, and they work whether or not the code was written by a model.

Philosophies

The positions people are actually arguing from. Worth being able to name the one you hold, and the one the person disagreeing with you holds.

Unix philosophy, applied to agents

Small, composable, one job each, text at the boundaries. An agent that does one thing against a stated contract can be tested, replaced and reasoned about. An agent that "owns the feature" cannot. The oldest position here and still the strongest.

Least context

The direct analogue of least privilege: give an agent the minimum context the task requires. Not primarily an efficiency argument -- extra context is extra surface for the agent to be confused by, and it makes the failure harder to attribute afterwards.

Deterministic orchestration, stochastic execution

The workflow is code: ordinary, versioned, reviewable, replayable. The steps inside it are agents. The argument is that you want exactly one source of nondeterminism, at a known place, rather than an emergent system whose control flow is itself a model output.

Strict contracts over robustness

Postel's law says be liberal in what you accept. For contracts between agents, invert it: be strict, and fail loudly. Liberal acceptance hides drift until it is expensive, and the protocol-ossification critique of Postel's law applies double when the party on the other end will happily invent a plausible field name.

Reversibility as the safety property

Since you cannot fully verify agent output, optimise for cheap undo instead: small commits, feature branches, idempotent operations, staged rollout. Correctness you cannot guarantee, recoverability you can.

Conway's Law awareness

Your agent topology becomes your architecture. Choosing how to fan out is an architectural decision, made early, usually implicitly, and usually by whoever wrote the orchestration script. Worth making on purpose.

The honest state of the field The contract half is mature: forty years of design-by-contract and twenty of consumer-driven contract testing, and it transfers essentially unchanged. The multi-agent half is not. Correlated failure across one model family is unsolved, model-as-judge is measurably biased, and the returns on fan-out are poorly characterised. Adopt the contract discipline with confidence. Treat the orchestration patterns as promising and instrument them, because the field does not yet know which ones pay.
// contracts · decomposition · adversarial · pre-mortem · oracles · correlated failure · philosophies //