Skip to content
  1. Home
  2. Guides
  3. Guide

Operating a Browser MCP Agent with Isolation, Action Limits, and Visual Verification

Operating a Browser MCP Agent with explicit control means treating the template as a starting point rather than a complete control plane: its checked-in Playwright MCP command has no extra arguments, so neither isolated storage nor action limits should be assumed. Use structured accessibility state to choose actions, require human approval for consequential tool calls, and inspect screenshots as visual evidence rather than as a basis for clicking. The controls below preserve the template’s browsing flow while making session boundaries, stop conditions, and verification part of your own orchestration.

What it solves: Browser MCP

The original Browser MCP Agent code provides a compact example of a browser automation agent built around MCP, mcp-agent, Playwright, and Streamlit. Its template README says the application accepts natural-language commands for browsing and interacting with websites through MCP and mcp-agent with Playwright integration.

That makes the project useful as an implementation example, especially if you need to see how the Streamlit interface, language model, MCP server, and browser tool fit together. It does not make the template a production policy layer for credentials, session state, approvals, or accepted outcomes.

The browser interaction model also matters. Playwright MCP’s documentation says it works through structured accessibility snapshots rather than screenshots or vision models. A screenshot can help a person inspect the rendered result, but it is not the representation from which the next browser action should be selected.

The template depends on mcp-agent, which describes itself as a simple, composable framework for building agents with MCP. The containing repository is published under the Apache License, Version 2.0.

How to run it

Requirements and keys

The template README lists these prerequisites:

  • Python 3.8+.
  • Node.js and npm for Playwright.
  • An OpenAI or Anthropic API key.

The README calls Node.js and npm a critical requirement and links to the Node.js download page. The template documentation does not state a separate project-install command, so do not substitute a guessed installation sequence. Check the current template README for dependency installation, then use the documented start command below.

The checked-in mcp_agent.config.yaml contains the relevant server and model settings. The YAML is shown here as an indented excerpt:

  mcp:
    servers:
      playwright:
        command: "npx"
        args: ["@playwright/mcp@latest"]
  openai:
    default_model: "gpt-4o-mini"

The configuration says secrets belong in a separate mcp_agent.secrets.yaml file that can be gitignored. It does not show the exact secret field to create, and the README does not explain how to select the Anthropic provider even though it lists an Anthropic API key as an alternative. Check the current template and mcp-agent documentation before adding either key; do not invent a field name or assume that changing the provider requires only a different secret.

Start the app

The exact start command given by the README is:

streamlit run main.py

That command starts the template, but it does not add an isolation argument or an action-policy layer. In our setup, we would review the Playwright MCP invocation before treating the process as ready to receive consequential work.

If the task requires clean browser state, use the isolated-context mode described in Playwright MCP’s current documentation. The documentation available here does not state the exact isolated-mode option, so copy the supported invocation from the current Playwright MCP README rather than guessing a flag.

What we observed

This page is based on reading the template’s code and documentation, and we have not run it.

The agent in main.py is bound to the MCP server named playwright. Its instruction tells it to take screenshots of page elements when useful, summarize web content in Markdown, follow multi-step browsing sequences, and return a status update when the commands are complete.

The same file enables conversation history and sets maxTokens to 10000. A nearby comment suggests turning history off to reduce the context passed to the model. That is a configuration choice visible in the code, not a measured capacity or success guarantee.

The documented flow does not contain a feature called action limits. It also does not define a visual pass condition that rejects a connection error, missing content, or an unexpected page. The screenshots and status update help the agent communicate, but neither should be treated as proof that the intended user-visible state was reached.

Check it worked

Start with a controlled page that cannot send a message, place an order, submit a final form, or alter shared data. A successful Streamlit launch only establishes that the application started. It does not establish that the MCP server launched cleanly, the browser reached the intended page, or the requested interaction occurred.

For each relevant interaction, use browser_snapshot as the basis for deciding what to do next. After the action, use browser_take_screenshot and inspect the image content. Check the expected heading, text, confirmation, or main image rather than confirming only that a screenshot file exists.

In our operating record, a screenshot was once accepted as evidence even though it showed a connection-refused page. The rule that followed was to verify evidence content, not file presence. We would apply the same rule here: read the final state where a user would see it, and do not accept the agent’s completion message when the rendered page shows an error or an earlier state.

For an isolated session, close the browser and confirm that the expected storage state is absent before treating the session boundary as verified. The completed result is not just a navigation sequence or a status message; it is the expected browser state plus readable visual evidence.

Where it breaks

Playwright MCP’s own documentation says that it is not a security boundary. Isolation, action restrictions, approval, and acceptance checks therefore need explicit ownership outside the browser tool.

Isolated browser state

The template launches npx with @playwright/mcp@latest and no extra arguments. Under the documented Playwright MCP defaults, that means a persistent profile rather than an isolated context. The Playwright MCP README says the persistent profile stores logged-in information.

That default is unsuitable when unrelated tasks must not inherit cookies, authenticated sessions, or other browser storage. Isolated contexts are the documented alternative. In that mode, each session begins with isolated storage, and closing the browser ends the session and loses its storage state.

If persistence is genuinely required, our operating rule would keep that persistence inside a task-specific boundary rather than sharing a logged-in profile across unrelated work. The browser documentation describes persistent and isolated modes, but the exact command needed to select the current isolated behavior is not established here. Check the live Playwright MCP documentation when configuring it.

There is also an unresolved browser-mode discrepancy. The template README says the application uses Playwright to control a headless browser. Playwright MCP’s documentation says the browser is headed by default and that --headless switches it to headless mode. We do not reconcile those statements; check the actual launch and record which mode is running before relying on screenshots or browser automation.

Action limits

The template’s natural-language browsing flow does not enforce a limit on what the agent may do. We would add action limits in the surrounding orchestration by defining which action classes are allowed, which outcomes terminate the task, and which actions require approval.

For example, a research task may permit navigation and reading while reserving form submission for explicit review. Outbound messages, payments, and final submissions to third parties should wait for the human owner’s approval. Approval relayed by another agent is not equivalent to approval from the owner.

The MCP tools specification says there should always be a human in the loop able to deny tool invocations. We would make that ability operational by presenting the intended action and target before a consequential browser step, then allowing the main session to continue only after the required approval.

An arbitrary tool-call count is not a substitute for an action boundary or a definition of done. The stop condition should describe the allowed outcome, while irreversible or public actions remain behind the approval gate.

Visual verification

browser_take_screenshot captures the current page, but its documentation explicitly says actions cannot be performed from the screenshot; Playwright directs the agent to use browser_snapshot for actions. This prevents a visual guess from becoming the next click target.

The template’s instruction to take screenshots when useful complements that interaction model. We would use the accessibility snapshot to locate and select the next action, then use the screenshot to answer a different question: does the page now look correct to a person?

A useful visual check examines content, not merely image dimensions or file creation. Confirm the expected heading or confirmation text, check that the main image loaded when it is part of the task, and make sure the page is not displaying an error. If the screenshot fails, isolate the failed interaction and recheck the live page state before repeating any submission.

Long sessions and recovery

The template enables conversation history with maxTokens set to 10000, while its comment recommends disabling history to reduce context. For long or stateful work, we would preserve a written handoff and start a fresh session when accumulated context is no longer useful. The code does not establish how much task complexity 10000 can handle.

After a crash or restart, our rule is to inspect the real state before acting again. Do not blindly resubmit a form or repeat an external action because the local process disappeared; the action may already have landed. Check the browser, application, and other live systems first, then continue from the verified state.