Skip to content
An operator working at a bank of control consoles in a dim room

Field manual for agent operators

Make AI agents actually do the work.

Setup guides, failure post-mortems and operating notes for people who already run AI agents, written from running dozens of them side by side every day.

  • Setup
  • Coordination
  • Verification
  • Cost
  • Post-mortems

Latest

New guides and post-mortems

All guides
  1. SETUPHow to Add MCP Servers to Claude Code and Keep Their Tools Safe to CallA Claude Code MCP guide to adding remote servers, choosing scope, checking connection status, and containing calls with server- and tool-level permission rules.
  2. LIBRARYOperating an Agentic RAG Agent with Citation Gates and Recovery DrillsA practical operating guide to agentic RAG that explains the template's Agno, Gemini and OpenAI structure and treats citation gates and recovery drills as external controls, not built-in features.
  3. LIBRARYOperating a Browser MCP Agent with Isolation, Action Limits, and Visual VerificationHow the Browser MCP Agent template runs, why its default launch does not provide isolation or action limits, and which operating controls and visual checks to add.
  4. FRAMEWORKClaude Agent SDK vs OpenAI Agents SDK for Tool-Building TeamsClaude Agent SDK exposes the tools, agent loop and context management that power Claude Code in Python and TypeScript, while OpenAI Agents SDK provides provider-agnostic workflows, multiple tool categories, guardrails and handoffs.
  5. COORDHow to Run Claude Code Agent Teams Without Overlapping Repository EditsHow to use Claude Code agent teams without overlapping repository edits by enabling the experimental mode, assigning distinct file sets, and checking task completion.
  6. COORDHow to Run Concurrent Coding Agents in Git Worktrees Without Merge CollisionsA Git worktree gives each coding agent a separate checkout, while path ownership and atomic shared-file edits reduce—but do not eliminate—merge collisions.

The first guides are being written and checked against each tool's own documentation.

Agent case library

Open-source agent builds, by the job they do

Research teams, recruiting, competitor tracking, daily briefings. Each page takes one public, runnable agent project and answers four things, with a link to the original code.

  • What it solvesThe job, and who would use it.
  • How to run itRequirements, keys and the steps.
  • What we observedDated. Read-only pages say so.
  • Where it breaksLimits, cost and what to change.

Open the case libraryFrameworks compared

  1. LIBRARYOperating an Agentic RAG Agent with Citation Gates and Recovery DrillsA practical operating guide to agentic RAG that explains the template's Agno, Gemini and OpenAI structure and treats citation gates and recovery drills as external controls, not built-in features.
  2. LIBRARYOperating a Browser MCP Agent with Isolation, Action Limits, and Visual VerificationHow the Browser MCP Agent template runs, why its default launch does not provide isolation or action limits, and which operating controls and visual checks to add.

The first case pages are being read, run and written up. Each one will link the original code and state the date it was tested.

Operating rules

Five rules that hold when many agents share one codebase

These come from daily operation, not from a framework's documentation. Every guide on the site applies at least one of them.

Read the coordination track

  1. Done means checked, not reported

    An agent's summary is a claim. Exit code 0, a green build and an HTTP 200 are signals. The task is done when the result has been read back from where a user would see it.

  2. One writer per file at a time

    Two agents editing the same file is a race. Give each task a clear set of paths, or give each agent its own worktree, and make shared files small, atomic edits.

  3. Commit only your own lines

    A shared working tree holds everyone's unfinished work. An agent that runs a plain commit ships all of it. Stage by path, check the diff, and never sweep the tree.

  4. Write the handoff down

    The next session knows nothing. A handoff states what is finished, what is half-done, what was ruled out and the exact next step, in a file the next agent will actually read.

  5. Deploy from a state you can name

    Before a deploy, the agent should know what is live, what is in the repository and what is only on disk. If those three differ and it cannot explain why, it stops.

Post-mortem format

Every failure written up the same way

A post-mortem is useful only if you can act on it. Each one here has four parts, and the last one is never optional.

Read post-mortems
  1. 01 · What broke

    The visible damage

    What the agent did, to which files or systems, and how it was noticed.

  2. 02 · Why

    The cause in the setup

    The missing instruction, missing check or shared resource that made the mistake likely.

  3. 03 · Fix

    How it was repaired

    The recovery steps that worked, including the ones that did not.

  4. 04 · Guard

    What stops a repeat

    The rule, hook or check added so the same failure cannot happen quietly again.

Questions

Before you read further

Who is Agent Talk for?

People who already run AI agents, such as coding agents in a terminal or agents built on a framework, and want them to finish real work without constant supervision. If you are still deciding whether to try an agent at all, most guides here will assume more than you need.

Why does my agent say a task is done when it is not?

Usually because nothing in its instructions defines what done means, so it reports the last step it ran. The fix is a completion check the agent must run and quote before it reports: a test, a build, or reading the result back from where a user would see it.

Can several agents work in the same repository at once?

Yes, with rules. Each agent needs a clear owner for the files it edits, commits that contain only its own changes, and a written handoff when work moves between sessions. Without those, agents overwrite and commit each other's unfinished work.

Do I need a multi-agent framework to run more than one agent?

No. Many working setups are several independent sessions sharing a repository, an instructions file and a task list. A framework helps when agents must call each other programmatically; it does not replace ownership rules or verification.

How do I keep agent costs under control?

Route each kind of task to the cheapest model that passes your checks, keep long stable instructions cacheable, hand noisy searches to a subagent so the main context stays small, and measure spend per task type rather than per day.

Ask

Describe what your agent did

Tell the assistant what you asked for and what happened instead. It answers now, and questions that keep coming up become full guides.