TL;DR: Agent scope control means deciding, before a run starts, what an AI agent may touch, which actions need approval, and what condition ends the run. A goal like "improve this codebase" isn't a stop condition; a checkable fact like "tests pass" is. Frameworks enforce part of this; you still have to write the rest into the prompt.
What Is Agent Scope Control?
Scope control is the answer to two questions, decided before the agent takes its first action rather than discovered afterward: what is this agent allowed to touch, and what has to be true for the run to be over. Everything else, which files it reads, which commands it runs, which tools it calls, follows from those two answers.
It happens in two separate places, and conflating them is where most "the agent went too far" stories start. The first is the runtime layer: the actual permission system in Claude Code, Codex, or Cursor, which decides whether an action needs your approval before it happens. The second is the prompt layer: the instructions you write that tell the model what "this task" includes and when it's finished. The runtime layer can stop a specific action. Only the prompt layer knows what the task actually is, which means it's the only layer that can tell a productive step from a pointless one.
Treat these as complementary, not interchangeable. A tight runtime sandbox with a vague prompt still produces an agent that burns its entire turn budget rewriting files it was never asked to touch, just inside a smaller box. A precise prompt running with no runtime limits at all is one bad tool call away from a bad day. Scope control means using both, and this piece is mostly about the one vendors can't write for you: the prompt.
That distinction also answers a question worth asking up front: is this about the agent being capable enough, or about it being bounded enough? Those are unrelated. A fully capable agent with no declared scope and no stop condition is exactly the one that finishes the task correctly and then keeps going, because nothing told it, concretely, that it was done.
Why Do Agents Drift Past the Task You Gave Them?
Three causes show up repeatedly, and none of them are the model "misbehaving."
The first is a goal instead of a check. "Refactor this module until it's clean" has no finish line a script could verify, so the model keeps finding one more thing worth touching. The second is tool access wider than the task needs. If an agent can write anywhere in a repository, it will eventually write somewhere you didn't mean, because nothing in its instructions distinguishes "the file I was asked about" from "a file I noticed along the way." The third is that the safety net you're counting on was never designed to be one. Guardrails built into a coding agent are described by their own vendors in hedged terms, not absolute ones, which matters if you're treating them as the actual stop condition instead of a backstop.
None of this requires a malicious prompt or a broken model. It requires an instruction that names a direction and never names an end.
What Exactly Counts as a Stop Condition?
A stop condition has to be checkable by something other than the model's own judgment of its progress. In practice, four types cover almost every agent task:
- Completion check — a fact someone else could verify without reading the transcript. A named file exists. A test suite exits 0. Every row in a list carries a status. If you can't write it as a pass/fail check, it isn't a stop condition yet, it's still a goal.
- Budget cap — a hard ceiling on turns, tool calls, wall-clock time, or spend, independent of whether the task feels done. This is the backstop for every stop condition you got wrong.
- Boundary trigger — an action that ends the run immediately regardless of progress: touching a path outside the declared scope, needing a credential the agent wasn't given, or reaching anything resembling production data.
- Escalation trigger — a condition that pauses and asks rather than silently continuing or silently stopping: an ambiguous instruction, conflicting file state, or an action that can't be undone.
Most agent prompts only specify the first type, if they specify any at all. The other three are what keep a wrong guess about "done" from turning into an expensive one.
A single vague instruction can trigger all four gaps at once. Take "update the docs to match the codebase," handed to an agent with no other qualifiers. It has no fact to check for "match" (no completion check), no reason to stay inside the docs folder if a code comment looks stale somewhere else (no boundary trigger), no reason to ask before restructuring a page instead of editing it (no escalation trigger), and no ceiling on how much of the repository it visits while it decides (no budget cap). Naming even one of the four narrows the run; naming none of them leaves the model to improvise all four, in whatever order is most convenient for calling the task finished.
How Do Claude Code, Codex, and Cursor Actually Enforce This?
Runtime enforcement is not standardized across tools, and describing one framework's defaults as if they applied everywhere is the single most common mistake in agent-safety content. Verified directly against each vendor's own docs on September 4, 2026:
| Feature | Claude Code | Codex | Cursor |
|---|---|---|---|
| Default file-write access | Read-only until you approve edits (Manual mode) | Read-only recommended outside version control; Auto (workspace-write) inside it | Can edit workspace files without approval, except configuration files |
| Documented turn or step cap | --max-turns flag (print mode); no limit by default | Not found on the vendor's security docs | Not found on the vendor's security docs |
| MCP servers need approval before first use | Yes — new MCP servers require trust verification | Not addressed on this page | Yes — every connection, then every individual tool call |
| Vendor's own framing of its guardrails | Reduce risk; "no system is completely immune to all attacks" | Reduces exposure; "treat web results as untrusted" | "best-effort guardrails rather than a hard security boundary" |
Two things in that table are easy to miss. First, Claude Code's own CLI reference documents --max-turns as opt-in and states plainly there is no limit by default, meaning the default for an unattended print-mode run is to keep going until something else stops it. Second, none of the three vendors describes its protections as sufficient on their own. Codex's docs warn that prompt injection can cause the agent to fetch and follow untrusted instructions, and that you should still treat web results as untrusted even with mitigations in place.
Where Does MCP Fit Into This?
If your agent connects to tools over MCP, the same scope question applies one layer further out: what can the tools do, not just the model. The protocol's specification says there should "always be a human in the loop with the ability to deny tool invocations", and that clients "consider tool annotations to be untrusted unless they come from trusted servers". It also tells clients to check a tool's results before handing them to the model, rather than trusting them by default. That's a protocol-level design choice, not a client feature, though how strictly it's enforced still varies: using MCP inside Claude Code walks through the connection and approval flow for that specific client.
The practical implication for scope control: an agent with three MCP tools attached has a larger scope than the same agent with none, even if its prompt never changed. Every tool you connect needs its own line in your scope statement, not an assumption that the model will only reach for what's relevant.
How Do You Write a Stop Condition Into Your Own Prompt?
Put scope and stop conditions in their own labeled block, separate from the task description, so nothing about them is ambiguous or buried in prose. The block below is deliberately narrow: it is the scope-and-stop layer only, not a full agent prompt with verification and output-format sections built in. If you want the fuller structure, AI agent prompt templates covers that.
SCOPE
- You may read and edit files inside: <path>
- You may not touch: <credentials, production config, anything outside SCOPE>
- You may call these tools without asking: <tool list>
- Ask before: <network access, deleting a file, anything not listed above>
STOP WHEN
- <a checkable completion fact — e.g. "the test suite exits 0 AND <file> contains <value>">
- OR you have used <N> tool calls / <N> minutes, whichever comes first
- OR you hit anything listed under "Ask before" — pause and report, do not proceed
IF BLOCKED
- State what you were trying to do and what stopped you
- Propose the smallest safe next step
- Do not retry the same blocked action more than once
Two habits make this hold up in practice. Write the completion fact as something a different person, or a script, could check without reading the run's reasoning. And put a budget cap in even when you're confident about the completion check, because your confidence is exactly what an ambiguous edge case will exploit. If your team already keeps a shared instructions file for coding agents, the AGENTS.md standard is the natural place to put boundary language that should apply to every run, not just one prompt.
What Should You Check Before You Let an Agent Run Unattended?
A short pass before you walk away from the terminal:
- Read the runtime's actual default for this session. Not what you remember from last month, the current mode. Defaults change between versions and between "starting fresh" and "resuming a session."
- Confirm the working-directory boundary matches your SCOPE block. If the tool can write to a parent directory with one extra approval click, know that before you grant it.
- List every attached tool and MCP server, and check whether each needed its own approval or was pre-approved from an earlier session.
- Set the budget cap in the runner, not just the prompt. A prompt-level budget is a request the model can misjudge; a runner-level flag like
--max-turnsfails hard. - Decide what "blocked" looks like before you start, so a stalled run reads as a report you can act on instead of a wall of unexplained tool calls.
- Read the SCOPE block back once, out loud. Any sentence in it that needs a follow-up question to resolve will get resolved by the model instead, in whichever direction finishes the run fastest, not the one you would have picked. If you want a more adversarial pass than reading it back, red-team the prompt the same way you would any other one: rephrase the SCOPE block, then check whether the boundary still holds.
What Actually Goes Wrong Without Scope Control?
The realistic failure modes are less dramatic than "the agent went rogue," and more expensive for exactly that reason. A run with no completion check keeps finding marginal changes to make, because the vendor docs are honest that nothing stops it by default in that mode. A run with tool access wider than the task edits a file adjacent to the one it was asked about, because nothing in the prompt said it couldn't. A run that fetches a page as part of its task follows an instruction embedded in that page, because, as Codex's own docs put it, prompt injection can cause the agent to fetch and follow untrusted instructions, and no vendor claims their mitigations remove that risk entirely.
None of these need a hostile prompt. They need an unbounded goal, tool access nobody scoped down, and no budget cap catching the difference. All three are fixable in the prompt, before the run starts, for less effort than reading the transcript afterward to figure out what happened.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.
Create An Account