Skip to content
  1. Home
  2. Guides
  3. Guide

Why Multi-Agent LLM Systems Fail: Lost State, Broken Handoffs, and False Completion

The practical reason why multi-agent LLM systems fail is not usually one bad model decision: execution state can disappear, handoffs can omit decision-critical context, or a run can end without a valid check against the task. MAST places these mechanisms in system design, inter-agent misalignment, and task verification, so repair the execution contract before adding another agent. Make state recoverable, handoffs explicit, and completion testable.

What broke

The MAST abstract introduces MAST-Data as 1600+ annotated traces collected across 7 multi-agent frameworks. The full text gives the distinct exact total of 1642 annotated execution traces from 7 frameworks, spanning coding, math, and generic tasks. The taxonomy was developed from 150 traces with expert human annotators, with inter-annotator agreement of kappa = 0.88. It identifies 14 failure modes in 3 categories: system design issues, inter-agent misalignment, and task verification.

FM-1.4 and FM-2.1

The title phrase “Lost State” is not a MAST category name. It maps to FM-1.4, loss of conversation history, and FM-2.1, unexpected conversation resets. The paper defines FM-1.4 as context truncation that disregards recent interaction history and reverts the agent to an earlier conversational state (MAST full text).

We have also hit this pattern outside model conversations. After a crash or forced restart, scratch files and background child processes were gone, and a command-line login had expired. Our rule is to inspect APIs, repositories, registries, and half-written shared files before repeating work. A restarted process may have landed only part of an intended change, so its old intention is not proof of current state.

FM-2.4 and FM-2.5

“Broken Handoffs” is also title shorthand. It maps primarily to FM-2.4, information withholding, and FM-2.5, ignoring other agents’ input. FM-2.4 means failing to communicate important data or insights that could change another agent’s decision-making. FM-2.2, proceeding on wrong assumptions instead of seeking clarification, and FM-2.3, task derailment, belong to the same inter-agent weakness (MAST full text).

In our setup, two research workers received overlapping topic identifiers and overwrote each other’s files. We now give parallel workers disjoint identifiers. A written handoff records what is finished, what is half done, what has been ruled out, and the exact next step; otherwise the receiving agent must reconstruct decisions from partial output.

FM-1.5 and FM-3.1–FM-3.3

“False Completion” maps to FM-1.5, not recognizing task completion, and the task-verification modes FM-3.1, premature termination; FM-3.2, no or incomplete verification; and FM-3.3, incorrect verification. FM-3.2 covers omitted or partial checking of task outcomes, allowing errors and inconsistencies to propagate undetected (MAST full text).

The same paper reports that many verifiers perform only superficial checks, even when instructed to verify thoroughly. Examples include checking whether code compiles or whether leftover to-do comments remain. Those signals can coexist with an output that does not satisfy the actual task.

We have seen the same operational mistake. A copied verification gate returned PASS without testing anything relevant. Elsewhere, a screenshot file existed but displayed a connection-refused page. An overnight batch exited with code 0 even though only about one draft in eight had been accepted. We now inspect the result itself and count accepted outputs; file existence and exit status are not completion evidence.

Why

The paper’s modes form a causal chain: the system can fail to preserve the task and state, agents can fail to share or interpret the work, and the final verifier can fail to detect the resulting error. The reported shares and practical first checks are:

Category Paper mode Reported share First thing to inspect
System design FM-1.1 Not following task requirements 11.8% Objective and acceptance criteria
System design FM-1.2 Not following agent roles 1.5% Role and ownership contract
System design FM-1.3 Step repetitions 15.7% Repeated or non-progressing actions
System design FM-1.4 Loss of conversation history 2.80% Context assembly and truncation
System design FM-1.5 Not recognizing task completion 12.4% Terminal condition
Inter-agent misalignment FM-2.1 Conversation resets 2.20% Session continuity and reset events
Inter-agent misalignment FM-2.2 Wrong assumptions instead of clarification 6.80% Uncertainty and clarification policy
Inter-agent misalignment FM-2.3 Task derailment 7.40% Drift from the active objective
Inter-agent misalignment FM-2.4 Information withholding 0.85% Handoff payload completeness
Inter-agent misalignment FM-2.5 Ignoring other agents’ input 1.90% Input receipt and acknowledgement
Inter-agent misalignment FM-2.6 Reasoning-action mismatch 13.2% Difference between stated plan and action
Task verification FM-3.1 Premature termination 6.20% Stop conditions
Task verification FM-3.2 No or incomplete verification 8.20% Required outcome checks
Task verification FM-3.3 Incorrect verification 9.10% Independence and evidence quality

The taxonomy and shares in the table come from the MAST full text; the checks are operational translations, not additional measurements by Agent Talk. These figures describe the paper’s dataset, not a forecast for every architecture.

System-design failures begin when requirements, roles, history, or completion rules are implicit. Inter-agent failures appear when agents reset, assume, derail, withhold, ignore, or act inconsistently with their stated reasoning. Task-verification failures then allow an incorrect result to pass. More agents increase the number of state transitions and handoffs that need observable controls; they do not remove the need for those controls.

Long-running tool use makes compounding errors especially dangerous. Anthropic describes agents as stateful and says errors can compound while an agent maintains state across many tool calls in its multi-agent research system guide. Cognition separately argues that distributed decision-making and insufficiently shared context make multi-agent systems fragile in its “Don’t Build Multi-Agents” post. These are explanatory accounts, not additional benchmark results.

Fix

Recover state before restarting the workflow. Read the real system of record: APIs, repositories, shared registries, running processes, and partially written files. Do not repeat an action until you know whether it already landed. In a shared codebase, keep one writer per file and use disjoint task identifiers so parallel work cannot overwrite the same state.

Then require a handoff record like this plain operating template, which is not a framework configuration schema:

Goal:
Done when:
Finished:
Half done:
Ruled out:
Owner:
Exact next step:
Evidence location:

The receiving agent should be able to continue without guessing what the sender knew. Keep planning, consequential file writes, deployment, and final verification with the main session when subagents are handling noisy searches or independent drafting. In our setup, the parent sees only a subagent’s summary, so consequential details otherwise disappear at the boundary.

Anthropic reports that agents without detailed task descriptions can duplicate work, leave gaps, or fail to find necessary information in its multi-agent research system guide. Include the objective, constraints, interfaces, ownership, and observable completion conditions rather than asking an agent to infer them from a short role label.

Add production tracing around the complete run. Anthropic reports that full production tracing enabled its team to diagnose why agents failed and fix issues systematically in the same guide. Capture context assembly, agent inputs and outputs, handoffs, tool calls, stop decisions, and verification evidence. Keep secret values out of traces; record key names instead.

Finally, verify the task objective, not merely syntax or surface signals. In the paper’s intervention study, adding a high-level task objective verification step to ChatDev yielded a +15.6% improvement in task success on ProgramDev (MAST full text). That is a reported study result, not a general performance guarantee.

No framework-neutral flag or command is specified for this repair. The documentation does not say that one setting fixes all three MAST categories. For framework-specific tracing, stopping, or verification controls, check the product’s current documentation and use only its supported interface.

Guard

Treat tracing as LLM observability, not decorative logging. A useful trace should let you reconstruct:

  • The active objective, role, and acceptance criteria.
  • What context each agent received and whether it was truncated.
  • What crossed each handoff and whether the receiver acknowledged it.
  • Which actions followed the stated reasoning.
  • Why execution stopped and what verifier examined the result.

Turn those failure modes into agent evals. Include cases where required context is missing, another agent supplies decisive information, an assumption remains unresolved, a verifier must reject incorrect output, and a superficial check would pass. Read the actual artifact produced by each case; do not grade only a summary, screenshot, or PASS field.

Test every new gate once with a case that must fail. Keep one writer per file, use disjoint work identifiers, and require atomic replacement for shared files. After an interruption, verify real state before re-running anything. For a new batch, send one representative high-risk item through the full chain, then compare accepted outputs rather than process exit codes.

Our completion rule is direct: done means checked. A summary, exit code 0, green build, and HTTP 200 are signals. The task is complete only after the result has been read back from the surface where a user would see it. That control catches the errors the other checks can miss.

Sources