Skip to content
  1. Home
  2. Guides
  3. Guide

Why Claude Code Got Slow: A Post-Mortem of Context, Tools, and Retries

Claude Code can feel slow when a long conversation keeps expanding the request context, a tool result repeatedly refills it, or transient service errors trigger retries; the Claude Code cost guide explains why later turns and tool calls carry accumulated material. Treat the symptom as a condition to diagnose, not proof of one product-wide slowdown: match it to context, customizations, cache behavior, or service pressure before changing the setup.

What broke

A useful post-mortem starts with the signal Claude Code actually produced:

  • Context pressure: High CPU or memory usage is one documented reason to run /compact regularly and reduce context size. The same remedy can help when an old conversation has accumulated work that no longer belongs in the current task. See the Claude Code troubleshooting guide.
  • Compaction that cannot settle: Autocompact is thrashing means automatic compaction succeeded, but a file or tool output refilled the context several times in succession. Claude Code stops retrying rather than continue an unproductive loop that would waste API calls. The troubleshooting guide documents this failure mode.
  • A customization-dependent path: A plugin, MCP server, or hook may be contributing to the problem. Restarting with claude --safe-mode disables customizations for that session so you can test whether they are responsible, as explained in the Claude Code troubleshooting guide.
  • Retry or service pressure: Transient failures are retried up to 10 times with exponential backoff before an error appears. A repeated 529 Overloaded means temporary capacity pressure across all users, while Request timed out means the API missed the connection deadline; the default request timeout is 10 minutes. These cases are defined in the Claude Code error reference.

Why

Context. Claude Code sends the full conversation with every request. Each tool use creates another request carrying that batch of tool results, so a long Claude Code loop can keep carrying earlier decisions, file content, and command output even when the current step is small. The cost guide also gives two useful boundaries: /clear starts fresh when moving to unrelated work because stale context wastes tokens on later messages, and a scheduled task can fire while the session is idle while sending the full context again.

Cache. A pause immediately after changing models has a different cause. Each model has its own prompt cache, so switching with /model makes the next request read the entire conversation with no cache hits, even when its content has not changed. After a sufficiently long idle gap, the next request recomputes the input and re-establishes the cache, which can make the first turn after returning noticeably slower. Both behaviors come from the Claude Code prompt-caching guide.

Tools. Oversized MCP results have separate handling. When a result with no image content exceeds the output limit, Claude Code saves it to a file and puts the file path into the conversation instead of leaving the full result inline. Claude reads that file when it needs the content, as described in the Claude Code MCP guide. If the file is repeatedly read and immediately refills the context, the result becomes part of the compaction-thrashing pattern rather than evidence that compaction itself failed.

Retries. A visible error may arrive only after the automatic retry sequence has run. A 529 Overloaded is a temporary capacity condition and does not count against your quota. A timeout is different: the API failed to respond before the deadline, potentially during high load or while generating a very large response. The error reference distinguishes these signals, so they should not all be treated as local context problems.

Fix

Start with the narrowest command that matches the failure.

For context pressure or an oversized result that should not remain in the working conversation, use:

/compact
/compact keep only the plan and the diff
/clear

/compact summarizes the conversation and accepts optional focus instructions; the focused form above is the troubleshooting example for dropping large output while retaining the plan and diff. The Claude Code commands reference defines the command, while the troubleshooting guide recommends regular compaction for high CPU or memory use. Use /clear when switching to unrelated work rather than carrying stale context into every subsequent message.

If the problem is an oversized file, the documented recovery is to ask Claude to read a smaller section, such as a specific line range or function, instead of the whole file. You can also move large-file work to a subagent so it runs in a separate context window. These options are listed in the Claude Code troubleshooting guide; the documentation does not prescribe one universal subagent command for this recovery.

Then check the local installation and extension path:

/doctor

/doctor runs an automated check of the installation, settings, extensions, and context usage. It proposes fixes that it can apply after you confirm them, according to the troubleshooting guide.

If the slowdown appears only with your normal configuration, restart in safe mode:

claude --safe-mode

Safe mode disables all customizations for the session. Repeat the smallest failing action and compare the behavior. If the problem disappears, inspect the plugin, MCP server, and hook paths that the troubleshooting guide identifies as possible sources. That guide does not provide individual disable flags for each customization, so check the current Claude Code documentation before inventing a persistent configuration change.

For cache-related pauses, avoid unnecessary model switching when cache continuity matters. A request following /model has no cache hits for the previous model, while the first request after a long idle gap can be slower because it rebuilds the cache. The prompt-caching guide does not define a universal idle cutoff, so do not convert this behavior into a fixed timer rule.

Retry errors need different handling. A 529 Overloaded is temporary capacity pressure, not a local usage-limit event, and neither /compact nor /clear removes it. For Request timed out, avoid requesting a very large generated response when the task needs only a smaller result; this follows the output condition identified in the error reference, rather than relying on an undocumented timeout flag. The reference does not define another local command for persistent overload or timeout behavior.

Guard

Use the same diagnostic order on the next slowdown:

  • Run /doctor and review each proposed fix before confirmation. Record what it changes rather than accepting every change blindly.
  • Reproduce the smallest failing action under claude --safe-mode. If the symptom changes, investigate plugins, MCP servers, and hooks; if it does not, that weakens the customization hypothesis without proving the cache or service layer is responsible.
  • Match the exact signal. Treat Autocompact is thrashing as repeated context refilling, 529 Overloaded as temporary capacity pressure, and Request timed out as a missed deadline.
  • Before a long phase, use focused /compact instructions to retain the current plan and diff. Use /clear between unrelated tasks, avoid needless /model changes, and account for scheduled tasks that may send full context while the session is idle.
  • After changing the workflow, repeat the real task and inspect its output. Do not use a successful command exit, an agent summary, or the absence of an immediate error as proof that the underlying work succeeded.

In our own agent operations, we treat command completion as a signal rather than verification. We read the result back from where a user would see it, and we keep checks inline when we can run them ourselves. We also use subagents for noisy searches, but a fresh subagent knows only its brief and the parent sees only its summary, so that summary is not proof. This is our operating experience, not a measured Claude Code performance result.

If behavior persists outside these documented patterns, keep the exact error text and check the current troubleshooting guidance. Do not invent a flag, configuration key, retry count, or performance claim to fill the gap.

Sources