Make multi-agent workflows deployable by bounding each agent’s task and write scope, keeping deployment authority with the main session, and placing explicit review gates before sensitive tool calls, merges, and deployments. A runtime pause, a local tool permission, and a repository protection rule are different controls; use each where the relevant decision can be made and recorded. In our setup, planning, file writes, deploys, and verification stay with the main session; the product mechanisms below come from documentation, and we have not run or benchmarked those third-party systems.
Set the ownership boundaries
In our setup, many agent sessions work on the same codebases and sometimes share one working directory. An uncommitted change in that directory belongs to every session, not only the agent that created it. We therefore give subagents narrow research or implementation tasks while the main session retains planning, file writes, deploys, and verification.
A fresh subagent knows only its brief, and the parent sees only its final summary. We use subagents for noisy searches and independent parallel work, state what has already been ruled out, and require findings in the final message instead of a separate report file. We do not delegate a check that the main session can perform directly.
One writer gets control of a file at a time. Parallel workers receive disjoint task IDs and separate scratch directories; output directories also carry a unique task marker and are checked before writing. Shared files are updated through a small atomic sequence: re-read the current file, compute the complete replacement, write it to a temporary file, then rename it over the original. New shared-registry entries are committed in the same turn they are added.
These are not review gates themselves, but they establish what the gate is supposed to inspect. Without clear ownership, a reviewer can approve a valid change to one file while another session has already altered the surrounding worktree.
Build multi-agent workflows with review gates
A useful agent orchestration pattern separates the agent that prepares work from the authority that accepts it. The agent can produce a patch, command result, deployment candidate, or outbound message. A review gate then decides whether that exact artifact may proceed.
LangGraph interrupts pause graph execution at specific points and wait for external input before continuing. Using an interrupt requires a checkpointer to persist graph state, a durable checkpointer in production, and a thread ID in the configuration so the runtime knows which state to resume. Treat those requirements as part of the gate’s deployment dependency, not as optional operational detail.
The OpenAI Agents SDK human-in-the-loop flow provides a tool-level gate. It pauses execution until a person approves or rejects sensitive tool calls. A tool can always require approval with needs_approval set to True, or an async function can decide case by case:
needs_approval = True
When the approval rule requires approval and no decision is stored, execution pauses. The pending entries appear in RunResult.interruptions as ToolApprovalItem objects, including details such as agent.name, tool_name, and arguments. That gives the reviewer a specific proposed action to inspect rather than a generic “agent paused” message.
Local coding agents need a different control because they can act directly through tools. Claude Code permission rules are evaluated in this order: deny, then ask, then allow. The first matching rule in that order determines the outcome; rule specificity does not change the sequence. Audit permissions in that order instead of assuming the narrowest-looking rule always wins.
A PreToolUse hook reference adds another decision point. It can allow, deny, ask, or defer a tool call, and it can modify tool input before execution. Its decision belongs inside hookSpecificOutput, rather than the top-level decision field used by other hooks.
Human authority must also remain separate from agent-to-agent delegation. In Claude Code agent teams, a teammate cannot approve a permission prompt or supply consent on the user’s behalf. A teammate denied an action also cannot relay it to another teammate to bypass the check.
Repository and deployment controls should sit outside the agent session. With GitHub required reviews, changes reach a protected branch through a pull request approved by the required number of reviewers with write permissions. Required status checks must also pass before collaborators can merge. These controls prevent an agent’s local success message from becoming an automatic merge.
For deployment jobs, GitHub environment required reviewers can require a named person or team to approve workflow jobs that reference the environment. The documentation allows up to six users or teams as required reviewers, each with at least read access to the repository, and one required reviewer’s approval is enough for the job to proceed.
Check it worked
A gate is verified only when its allowed and blocked paths have both been exercised. In our setup, every new gate is tested once with a case that must fail. We do not accept a copied check that reports PASS without testing the relevant refusal condition.
Check each layer against the decision it is supposed to enforce:
-
For LangGraph, trigger the interrupt, confirm that it requests the expected external input, and resume with the same checkpointer and thread context. Inspect the persisted state and resumed execution in the runtime you actually deployed; the documentation does not prescribe one universal verification command.
-
For the OpenAI Agents SDK, inspect
RunResult.interruptionsand confirm that eachToolApprovalItemidentifies the intended agent, tool, and arguments. Apply the supported decision and verify that execution advances only after that decision has been stored. -
For Claude Code, use non-sensitive test inputs to confirm the documented
deny,ask, andallowprecedence. If aPreToolUsehook is responsible for the decision, verify that itshookSpecificOutputproduces the intended tool behavior rather than relying on the hook’s existence. -
For GitHub, confirm that a protected-branch change without the required approval cannot merge, that a failed required status check blocks the merge, and that a workflow job referencing an environment waits when its required reviewer has not approved it.
After the gate opens, do not equate an agent summary, successful build, or successful HTTP response with completion. Our completion standard is to read the result back from where a user would see it. For a shared service, we also send a real request through the sanctioned deployment path and check that an unrelated tenant remains unchanged.
A new batch starts with one representative, high-risk item taken through the full production chain and audited before the rest run. The handoff should identify the reviewed artifact, the decision, the approver, what remains, and the exact next step.
Where it breaks
The first sharp edge is graph replay. When a LangGraph execution resumes, the runtime restarts the entire node rather than continuing from the exact line that called the interrupt. Because the node runs again, the LangGraph interrupts documentation advises that side effects before an interrupt should ideally be idempotent. Put the gate before an irreversible effect where possible; otherwise make replay safe.
Guardrails also have placement limits. In the OpenAI Agents SDK guardrails documentation, input guardrails run only for the first agent in a chain, while output guardrails run only for the agent that produces the final output. A protected chain entry and final result do not mean every intermediate agent has both controls. Map each check to the actual boundary where its input or output exists.
A Claude Code hook can create a stronger stop than a normal allow decision. According to the hooks reference, exiting with code 2 is a blocking error. On events that can block, a JSON permissionDecision of "allow" cannot override it. Treat hook execution failure as a stop condition rather than attempting to force approval through another field.
Branch protection can also have intentional exceptions. By default, GitHub branch protection restrictions do not apply to repository administrators or roles with the branch-protection bypass permission unless those restrictions are applied to them. Review those exemptions before assuming every actor is subject to the same merge gate.
Even a correct gate cannot repair an incorrect deployment target. In our own batch deployment work, we have shipped commits that another session had committed but deliberately left undeployed. Our current rule is to inspect each target for commits since its last deployment that we do not recognise. If the wrong version ships, we roll back to the previous version first, then notify the owner.
Subagents need boundaries too. A subagent handed a hypothesis tends to return it confirmed, while a parent relying only on a summary may miss discarded evidence. Keep architectural decisions, root-cause analysis, and final verification with the main session, and use an independent path before acting on a surprising result.
Finally, decide which failures stop the entire workflow. In our setup, failures involving irreversible actions, money, legal exposure, or a public surface stop the batch. Smaller failures are isolated to the affected item so one gate does not become an indiscriminate shutdown. Public submissions, outbound messages, and payments still wait for the human owner’s explicit approval; approval relayed by another agent is not treated as that approval.