Skip to content
  1. Home
  2. Guides
  3. Guide

How to Design Multi-Agent Systems with Handoffs, Shared State, and Verification

Design multi agent systems around explicit contracts: a handoff changes who owns the work, shared state preserves what later steps need, and verification checks the actual end state. The ownership distinction appears in OpenAI’s agent orchestration guidance: agents as tools handle bounded subtasks, while handoffs let a specialist own the rest of the current turn; Anthropic’s workflow and agent patterns recommend starting with the simplest workable solution and adding complexity only when needed. In our setup, that means keeping planning, file writes, deployment, and verification in the main session while routing bounded research or drafting elsewhere.

Set up handoffs

Start with an ownership rule, not an agent list. Decide whether the specialist is helping the coordinator or taking over from it. A bounded contribution belongs behind an agent-as-tools boundary; a change of responsibility belongs behind a handoff. This prevents every specialist from receiving control merely because its capabilities overlap with the main task.

The OpenAI Agents SDK handoff documentation says handoffs are represented to the LLM as tools. For an agent named Refund Agent, the documented tool name is:

transfer_to_refund_agent

That exact name is an SDK-facing identifier, not a universal convention you should copy into another framework. Route the handoff through the product’s documented construction API rather than inventing a matching function or configuration key in your own stack.

By default, the receiving agent sees the entire previous conversation history. The same documentation says an input_filter can change that inherited context. Decide whether the specialist genuinely needs the full history or only the task, current state, and relevant prior results. Filtering without understanding the dependency can remove context the specialist needs; inheriting everything does not make the important context clear.

The handoff itself needs a durable brief. Anthropic says each subagent should receive an objective, an output format, guidance about tools and sources, and clear task boundaries; without that detail, agents can duplicate work, leave gaps, or miss necessary information in its multi-agent research system guide.

In our setup, the written handoff records what is finished, what remains partly done, what has been ruled out, and the exact next action. We also record where durable artifacts live rather than copying large results into every message. A fresh subagent knows only its brief, and the parent sees only its returned summary, so the handoff record is the interface that must carry the missing context.

Set up shared state

Shared state should represent the current application snapshot, not become an accumulating transcript that every agent must reread. LangGraph defines State as a shared data structure representing the current application snapshot; it can use any data type and is typically defined through a shared state schema, according to the LangGraph Graph API overview.

The relevant documented concepts are:

State
shared state schema
reducer function

Each key in LangGraph State has an independent reducer. When no reducer is specified, updates to that key override its previous value. That default is simple, but it can discard work when separate branches update the same field. Before parallel workers touch a shared key, decide whether replacement is correct or whether the field needs explicit merge behavior. The documentation does not prescribe one reducer for every workload, so copying a schema without understanding its update semantics creates hidden data loss.

Separate conversational context from durable artifacts. Anthropic describes subagents storing work in external systems and returning lightweight references to the coordinator, which avoids information loss during multi-stage processing and reduces the need to copy large outputs through conversation history in its multi-agent research system architecture.

Our operating boundary is equally concrete:

  • Shared working-tree changes belong to every active session, not only the session that notices them.
  • Each session writes temporary files in its own scratch directory.
  • Output directories carry a unique task marker, and parallel workers receive disjoint task IDs.
  • Only one writer changes a shared file at a time.
  • Shared-file edits are computed first, written to a temporary file, and then renamed over the original.

These practices are not product features. They are controls we need because shared state does not imply shared ownership.

Boundary Written contract Continuation check
Handoff Current status, durable references, ruled-out paths, next action The receiver can continue without reconstructing hidden context
Shared state Current snapshot and defined update behavior Required content survives every relevant update
Verification Acceptance condition, observed result, verdict The real end state has been read and checked

Set up verification

Define the acceptance condition before delegating work. Otherwise, a worker can produce a plausible artifact while optimizing for the wrong result. Anthropic’s evaluator-optimizer pattern uses one LLM call to generate a response and another to evaluate it and return feedback in a loop, as described in Building effective agents. That pattern is useful only when the evaluator can inspect something meaningful about the result; another pass over the same unsupported narrative is not independent verification.

During execution, agents need “ground truth” from the environment at each step so they can assess progress. Anthropic gives tool-call results and code execution as examples in Building effective agents. In practice, that means choosing checks that observe the task rather than checks that merely repeat the worker’s account.

For agents that mutate state, Anthropic recommends end-state evaluation: decide whether the correct final state was reached instead of requiring one particular process. Its multi-agent research system guide also pairs agent adaptability with deterministic safeguards such as retry logic and regular checkpoints. The guidance does not provide a universal retry limit, so define a task-appropriate stopping condition rather than copying an unexplained retry count.

In our setup, verification stays with the main session and the strongest model assigned to root-cause work. We do not delegate a check that we can run directly, especially when the subagent’s conclusion is the thing that needs testing.

Check it worked

Check the whole chain rather than isolated artifacts. Start at the handoff: confirm that the receiving agent can identify the current state, resolve every durable reference, explain what remains, and name the next action without asking the original worker to reconstruct its reasoning.

Next, read the shared state from where the next component will consume it. A reference is useful only when its target exists and contains the expected material. Then inspect the user-visible result against the acceptance condition. The following signals are useful, but none is sufficient by itself:

Signal What it establishes What remains unchecked
Agent says it is done A summary was produced The real result and side effects
Process exits with 0 The process ended Whether the work was accepted
Green build The build completed Whether the rendered output is correct
HTTP 200 The request returned that status Whether its content and state changes are correct
Handoff note exists A packet was recorded Whether the receiver can continue from it

Our rule is simple: done means checked. An exit code, build result, status code, or agent summary is evidence of execution, not proof of completion. We once accepted a screenshot file that actually displayed a connection-refused page; file existence told us nothing useful. We now read the displayed content itself, including headings and whether the main result loaded.

A new verification gate must be exercised with a known-bad case that must fail. We do that once when introducing the gate, because a copied check that always passes can protect nothing. For a new batch, we also put one representative, high-risk item through the full production chain and audit it before running the rest. Established recurring work can proceed through the full batch without repeating that pilot.

When deployment is part of the task, the final-state check includes the live service, not just the repository or working tree. Our deployment check keeps a line diff, sends one real request, and confirms that an unrelated tenant remains unchanged. For content work, we count accepted outputs rather than treating a successful process as a successful batch.

Where multi agent systems break

The architecture becomes a poor fit when every agent needs the same changing context or when work contains many dependencies between agents. Anthropic says those domains are not a good fit for multi-agent systems today in its multi-agent research system guide. A single agent with direct tool access may preserve context and ownership more cleanly.

Coordination also has a measurable cost. Anthropic reports that agents in its data use about 4× more tokens than chat interactions, while multi-agent systems use about 15× more than chats. Those figures describe Anthropic’s data rather than predict your workload, but they make handoff length, repeated context, and verifier loops worth examining directly.

Several failure modes recur in our own setup:

  • A handoff preserves a conclusion but loses the constraints that made the conclusion valid.
  • Two workers update the same state under replacement semantics, so one result disappears.
  • Generic temporary filenames or output directories collide across concurrent sessions.
  • Overlapping task IDs let parallel workers overwrite each other.
  • A subagent handed a hypothesis tends to return that hypothesis confirmed rather than challenge it independently.

The last problem is not solved by adding another conversational turn. Give the checker a different observable result to inspect, state what has already been ruled out without presenting the desired conclusion as settled, and keep final verification where the main session can inspect the real state.

Retries and checkpoints are safeguards, not a substitute for acceptance criteria. Without a stopping condition, an evaluator can keep requesting changes to an answer that no check has grounded in the environment. Without clear update rules, shared state can make concurrent work less visible rather than more reliable.

We have not run or benchmarked the OpenAI Agents SDK or LangGraph for this guide, so their documented behavior is not presented as our test result. A multi-agent architecture is justified when work has separable objectives, explicit tool and source boundaries, durable state, and a verification path. Without those boundaries, added agents mostly create more handoffs, conflicting updates, and summaries to reconcile.

Sources