Lesson 3 of 4 · 35 min

Defend architecture with a failure budget

Choose a workflow or adaptive agent for a stated task distribution.

A design interview tests whether you can connect architecture to constraints. Naming a framework does not answer that question. Start with the input variability, required decisions, available tools, completion evidence, and permitted effects. Then choose how much control the model actually needs.
A fixed workflow is appropriate when steps are known and their outputs can be checked. An adaptive loop helps when each observation changes the next useful action. A hybrid often fits best: deterministic code handles identity, deadlines, and effects, while the model chooses among bounded investigative steps. This is a design choice, not a universal ladder where more autonomy is always better.
Estimate how failures compound. If a fictional task requires five independent steps with success probability 0.95 each, the probability that all five succeed is approximately 0.774. Real failures may be correlated, so the calculation is only a teaching approximation. It still shows why adding calls can reduce end-to-end reliability unless the extra calls provide enough recovery or verification value.
Discuss latency and cost with the same precision. Parallel independent reads can reduce wall time, but they do not reduce total work by themselves. A second agent may help with a separate evidence branch or review. It can also repeat the same mistake if it receives the same incomplete context. Give each component a purpose and an observable acceptance criterion.

Worked example

A fictional invoice assistant has two task types. Ninety percent require extracting five fields and matching a purchase order. Ten percent require investigating inconsistent records across systems. The proposed design uses a fixed extraction and validation workflow for all inputs, then routes unresolved mismatches into a bounded investigative loop.
The model cannot approve payment. The loop can read three systems, compare evidence, and produce a review packet. It stops after six tool calls or a deadline. A deterministic checker confirms invoice identity and arithmetic. The design avoids turning every routine invoice into an open-ended search while preserving flexibility for ambiguous cases. Its evaluation reports routine extraction accuracy separately from investigation resolution and unauthorized-action count.

Exercise and solution

A support assistant handles password-reset instructions and rare billing disputes. It may explain reset steps but cannot reset credentials or refund payments. Propose a control structure and two failure tests.
A good solution uses a verified instruction path for reset requests and a read-only evidence workflow for billing disputes. It returns a review packet for consequential actions. Tests include a reset document containing a malicious credential-upload instruction and a billing record with conflicting dates. Award one point each for matching control to variability, explicit effect limits, a grounded completion check, and relevant failure tests. A universal autonomous loop earns no extra credit without evidence that it improves outcomes.

Start from a task distribution, not an architecture label

The following fictional task inventory helps identify where adaptive decisions are needed.
Task familyShareUnknown branchingRequired effect
Extract invoice fields70%LowDraft structured record
Match known purchase order20%Low to moderateRead-only match
Investigate mismatch10%HighEvidence packet for review
A fixed extraction and validation path covers most work. The mismatch branch benefits from selecting further reads based on observations. The final payment action remains outside the model's authority in this example. An adaptive investigative loop can therefore be bounded without forcing the entire system into either a rigid script or unrestricted autonomy.
Architecture decisions should name the expected gain. If adding a reviewer model catches arithmetic defects already covered by a deterministic calculator, it may add little value. If it checks whether a report's conclusion follows from several conflicting records, it may address a harder judgment task. Compare those roles on held-out cases and include their cost.

A second worked case: correlated errors defeat naive voting

Suppose three agents classify whether a contract contains a renewal penalty. All three receive the same truncated excerpt that omits the penalty clause. They unanimously say no. Majority vote has not supplied independent evidence; it has repeated the same information failure.
code
1shared input defect:2  retrieved excerpt excludes section 83worker A: no penalty found4worker B: no penalty found5worker C: no penalty found6missing evidence:7  full contract section 8, where penalty terms appear8correct next action:9  inspect retrieval coverage or fetch the missing section
The proper repair targets evidence coverage. More votes on the same incomplete excerpt do not address the cause. Independent retrieval strategies or a completeness check may help, but they need evaluation. Independence is a property of error sources and information paths, not merely of separate model calls.
The five-step probability calculation in the original lesson also assumes independence. Real agent errors can be positively correlated through the same bad tool schema or stale source. They can also be corrected by later checks. Use the simple product only as a diagnostic illustration, then measure actual end-to-end outcomes.

Compare two bounded designs

Design A uses one model call to choose a route, then deterministic code validates extracted fields. Design B asks three agents to extract and vote, then skips deterministic validation. If arithmetic correctness is a hard requirement and the fields have a clear schema, B has removed a reliable check while adding correlated model work. It is not justified by the agent count.
A more defensible Design C keeps deterministic validation and invokes a second evidence reviewer only for unresolved mismatches. Its acceptance test compares resolution quality, false completion, latency, and cost on the mismatch slice. It also checks that routine tasks still follow the simpler path. This is a measurable architecture experiment rather than a preference for a fashionable pattern.

Misconceptions to reject

"Separate calls imply independent errors" ignores shared sources, prompts, and missing context. Correlation can make voting much less useful than the count suggests.
"Adaptive control should own every decision" moves deterministic invariants and permissions into a probabilistic component without evidence that doing so helps. The model can propose a route while code enforces the allowed actions.

Transfer exercise

A system must read a known form, calculate a total, and investigate missing line items only when totals disagree. Propose a control structure and explain when a second agent helps.
A strong solution uses deterministic parsing where available, a calculator for totals, and a bounded investigative branch for missing or conflicting evidence. A second agent may independently review a disputed interpretation, but should not replace the calculator with a vote on arithmetic. Score one point for the fixed path, one for the adaptive trigger, one for effect limits, and one for a measurable reason to add review.
The product-of-step-probabilities calculation is original teaching arithmetic under explicitly independent steps. The cited agent sources supply workflow context, not a measured per-step probability or a proof that real errors are independent.

Interview probe

Original practice: When would you reject a multi-agent design? A strong answer cites dependent work, shared-state risk, duplicated cost, or insufficient outcome gain. Follow up with a genuinely independent research branch. A weak answer equates the number of agents with capability.

Sources

docsAnthropic: building effective agentsanthropic.comdocsAnthropic: long-running agent systemsanthropic.com

Checkpoint

Routine invoice extraction has fixed steps; rare mismatches need investigation. Which design fits the task distribution?

AA planner allowed to bypass field validation.BAn unbounded loop for every invoice.CA fixed validated path with a bounded investigative branch for unresolved cases.DThree agents voting on arithmetic while removing the calculator.
Sign up free to answer and see why

Checkpoint

Three agents share the same excerpt missing the decisive clause. What does unanimous agreement establish?

AProof the clause does not exist.BA reason to skip retrieval validation.CIndependent confirmation of the full document.DAgreement on incomplete evidence, not coverage of the missing clause.
Sign up free to answer and see why

Checkpoint

Five independent required steps each succeed with probability 0.95. Approximate all-step success?

A0.774B0.50C0.99D0.95
Sign up free to answer and see why

Checkpoint

When is a second agent best justified in this workflow?

AWhen it independently reviews an unresolved evidence question and a held-out test shows useful gain within the budget.BWhen it repeats deterministic arithmetic already checked exactly, without finding another error class.CWhen it shares the same truncated source and votes on whether missing text exists.DWhen it removes the need to specify a completion rule.
Sign up free to answer and see why

Checkpoint

Eight independent read checks each take about 300 ms, and a final comparison needs all results. What design is best supported before adding adaptive planning?

ARun a bounded parallel read stage, then compare the complete checked results.BAsk a planner to invent the check order after every result even though dependencies are fixed.CStart the final comparison after the first result and infer the other seven.DUse eight write-capable agents even though the task is read-only.
Sign up free to answer and see why

Using the supplied evidence, explain how you would choose deterministic and adaptive control while checking correlated errors. Name one observation that would change your conclusion. Rate confidence from 1 to 5.

Not yetGetting thereConfident

Wrap-up

  • Choose control from the task's uncertainty. Every extra component should earn its cost through measurable usefulness.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.