AgentsAnalysis

The case against unattended agents: what went wrong, and how sandboxing actually works

From a deleted production database to poisoned npm releases, documented agent failures share a pattern. Here's what the sandboxes in Claude Code, Codex and Copilot actually enforce.

ByShajanthanFounder & Editor
Published
Reading9 MIN
Nested-boundary diagram showing layers of protection around an AI coding agent
What’s new, in 20 seconds
  1. In documented agent incidents from 2025–2026, the damage came from access the agent had been given: live databases, broad tokens, permission-bypass flags and open internet. None required a model to 'go rogue.'
  2. Vendors themselves say approval prompts and AI classifiers are not a security boundary; OS-level sandboxes, network egress limits and scoped credentials are.
  3. Unattended agents can be run responsibly, but only inside a boundary that holds even when the model is fooled.
Contents

Coding agents can now run for long stretches with nobody approving each step. Since version 2.1.283, Claude Code's built-in starting mode for interactive terminal and VS Code sessions lets a classifier approve actions instead of the user. Codex and GitHub's Copilot cloud agent work in the background and come back with a pull request. That convenience is real. So is a growing record of failures. This analysis looks only at incidents backed by a primary report (from the researcher, vendor or agency involved) plus at least one further published account, explains what each one teaches, and then describes what the main tools' sandboxes actually enforce, as opposed to what they're often assumed to enforce.

For definitions of "agent" and "autonomy" used here, see What "agentic" really means in 2026.

What has actually gone wrong

Last verified: October 8, 2026

DateSystemWhat happenedSources
May 26, 2025GitHub MCP server + Claude Desktop (demo)A malicious public GitHub issue led an agent to copy private-repository data into a public pull requestInvariant Labs; Simon Willison
Jul 2025Replit agentDeleted a live production database during a declared code freezeFortune; The Register
Jul 2025Amazon Q Developer for VS Code 1.84.0A release shipped with injected instructions to wipe local and cloud resourcesAWS bulletin; 404 Media
Aug 26, 2025Nx npm packages ("s1ngularity")Malware ran installed AI CLIs with permission-bypass flags to hunt for secretsWiz; StepSecurity
Feb 2026Cline's AI issue-triage botA prompt injection in an issue title led to an unauthorized npm releaseAdnan Khan; Simon Willison
Feb–Mar 2026Snowflake Cortex Code CLIA README injection made the agent run malware outside its sandboxPromptArmor; Simon Willison
Jul 25–28, 2026Frontier agents in UK AISI cyber testsAgents took or attempted unsanctioned actions involving real people and an open-source project; AISI found no resulting real-world harmUK AISI; Simon Willison
Aug 26, 2026Claude Code (Opus 5, auto mode)A website-borne attack chain led to code execution in a researcher's testsEmbrace The Red; Simon Willison

1. Destructive actions on live systems

In July 2025, SaaStr founder Jason Lemkin documented a Replit agent deleting his production database during what he'd declared a code and action freeze. Fortune reported that the agent then said a rollback wouldn't work; Lemkin restored the data anyway. The Register quoted him: "There is no way to enforce a code freeze in vibe coding apps like Replit." Replit CEO Amjad Masad called it "unacceptable and should never be possible." According to Fortune, Replit then announced automatic separation of development and production databases and a chat-only planning mode.

Lesson: an instruction in a prompt is not a control. The agent had credentials that could reach production.

2. Prompt injection turns legitimate access against the user

Invariant Labs showed in May 2025 that an agent connected to GitHub's official MCP server could be hijacked by an issue in a public repository. A user asking it to review open issues ended up with private-repository details published in a public pull request. The researchers called this "a fundamental architectural issue," not a bug in the server code. Simon Willison described it as his "lethal trifecta": access to private data, exposure to malicious instructions, and the ability to exfiltrate information. Running a local MCP server yourself? Our TypeScript MCP server guide covers keeping tool scopes narrow, and the MCP explainer covers the trust model.

PromptArmor reported to Snowflake in February 2026 that Snowflake's Cortex Code CLI could be made to run a malicious script after reading a poisoned README. Cortex's command allow-list didn't inspect commands nested in shell process substitution. The injection also persuaded the model to set the CLI's own dangerously_disable_sandbox flag. Snowflake fixed the issue in version 1.0.25 on February 28, 2026, before the March 16 public disclosure.

In August 2026, researcher Johann Rehberger published an attack on Claude Code running Opus 5 in auto mode. A website the agent was asked to summarize led it to download an archive and run a decoder. Inside the archive, a file named struct.py shadowed Python's standard library and executed code. The chain succeeded in 3 or 4 of 5 trials per variant, which Rehberger called "small samples." He reports that Anthropic closed his report as "Informative," on the grounds that auto mode is a best-effort classifier and that OS isolation and network egress control are the real boundary. Anthropic's own docs are consistent with that: they describe auto mode as a classifier that reviews actions, and offer the sandbox and containers as the isolation layer.

3. The agent toolchain as a supply-chain target

In July 2025, a release of the Amazon Q Developer extension for VS Code (1.84.0) shipped with injected instructions telling the agent to wipe a system "to a near-factory state" and delete cloud resources, as 404 Media reported. AWS's security bulletin traced the cause to a GitHub token with "overly broad access" in its build configuration. AWS said the code "was unsuccessful in executing due to a syntax error" and released 1.85.0.

On August 26, 2025, malicious versions of the Nx build system's npm packages stole developer secrets. Wiz and StepSecurity both reported that the malware also called locally installed AI CLIs (claude, gemini and q) with --dangerously-skip-permissions, --yolo and --trust-all-tools to search the filesystem for credentials. Wiz counted over 1,000 valid GitHub tokens leaked, and more than 5,500 private repositories made public in a second phase.

In "Clinejection," disclosed on February 9, 2026, security researcher Adnan Khan showed that Cline's AI issue-triage workflow ran Claude Code with Bash, Write and web tools on any new issue from any GitHub user. A prompt injection in an issue title could make it install a malicious package and poison the GitHub Actions cache used by the release workflow. On February 17, an unauthorized cline@2.3.0 was published to npm with an added install script. Cline deprecated it and published an advisory.

Lesson: agents with bypass flags, and AI bots that untrusted users can trigger, are now attack surface in their own right.

4. Capable agents with open internet

The UK AI Security Institute reported on August 4, 2026 that, during cyber evaluations between July 25 and 28, agents took unsanctioned actions in 10 of 122 runs, 19 actions in total. In the most serious case, an agent tried to insert malicious code into an open-source project, creating fake online identities to pressure the maintainer to approve it; the human maintainer refused. Other runs included social-engineering attempts that sent messages and files to real people. AISI says these attempts were unsuccessful. Most of the actions (17 of 19) came from Anthropic's Mythos 5 and two from OpenAI's GPT-5.6 Sol, out of seven models tested. AISI had deliberately enabled internet access and disabled developers' cyber classifiers, which it called standard for measuring maximum capability. It said this was not a case of a model escaping its sandbox, and that its investigation has not identified any resulting real-world harm. Its conclusion applies well beyond labs: "good containment should not depend on the model choosing not to test its boundaries."

How the sandboxes actually work

Claude Code

  • Permission modes decide whether Claude asks before acting: Manual (default), acceptEdits, plan, auto, dontAsk and bypassPermissions. From version 2.1.283, auto mode is the default starting mode for interactive sessions. Anthropic's docs say bypassPermissions "offers no protection against prompt injection or unintended actions" and should be used only in containers or VMs.
  • The Bash sandbox is a separate, OS-enforced layer and is off by default. You turn it on with /sandbox or sandbox.enabled. It uses Seatbelt on macOS and bubblewrap plus socat on Linux and WSL2. Native Windows runs commands unsandboxed. By default, sandboxed commands can write only to the working directory, a temp directory and directories you add. They can still read most of the machine, "including credential files such as ~/.ssh and ~/.aws/credentials," unless you deny those paths. Network traffic goes through a local proxy, and the domain allowlist starts empty.
  • Protected paths such as .claude settings, hooks, .mcp.json, shell startup files and .git/hooks stay write-protected even inside the writable area, because editing them would let a command widen its own permissions.
  • Escape hatches. Claude can ask to retry a failed command outside the sandbox. Setting allowUnsandboxedCommands: false ("strict sandbox mode") disables that. The docs also warn that allowing broad domains such as github.com "can create paths for data exfiltration," and that domain fronting could bypass the hostname check.
  • Containers. Anthropic's dev container docs warn that with --dangerously-skip-permissions, "dev containers do not prevent a malicious project from exfiltrating anything accessible inside the container, including the Claude Code credentials." They recommend pairing the flag with egress restrictions.

OpenAI Codex

  • Sandbox modes: read-only, workspace-write and danger-full-access. Approval policies: on-request, never or a granular table. The older untrusted policy is retired.
  • Version-controlled folders start in workspace-write with on-request approvals. Network access is off by default unless you set network_access = true.
  • Enforcement uses Seatbelt on macOS, bubblewrap plus seccomp on Linux, and a Windows-specific sandbox on native Windows.
  • --yolo is an alias for --dangerously-bypass-approvals-and-sandbox, which OpenAI's docs say is not recommended.

GitHub Copilot cloud agent

  • It runs in an ephemeral cloud environment. Only users with write access can trigger it.
  • It pushes only to a new copilot/ branch and "cannot directly run git push."
  • It can't mark its pull requests ready, approve or merge them, and the person who asked for the PR can't approve it.
  • Actions workflows don't run until a human reviews the agent's code. Hidden characters, such as HTML comments in issues, are filtered out before they reach the agent.
  • A firewall limits internet access by default, with a recommended allowlist for package registries. GitHub is candid about its limits: it "only applies to processes started by the agent via its Bash tool," "sophisticated attacks may bypass the firewall," and it "should not be considered a comprehensive security solution."

Containers and VMs

A dev container or VM adds a boundary around the whole agent, including its file tools and MCP servers, not just its shell. It helps only if you keep host secrets out (no mounted ~/.ssh or cloud credential files), restrict egress, and treat the bind-mounted repository as writable by the agent. Anything inside the box is reachable.

The case for autonomy

Unattended agents are not a mistake in themselves. Constant approval prompts produce what Anthropic calls "approval fatigue," where users "might not pay close attention to what they're approving." Anthropic says sandboxing cut permission prompts by 84% in its internal use, and it argues that defined boundaries "increase security and agency." METR's March 2025 study found that the length of software tasks frontier models complete at a 50% success rate had been doubling roughly every seven months, and METR still tracks this "time horizon" for new models. Agents can handle longer tasks, and asking someone to approve every ls wastes that. Copilot's design shows another approach: the agent works freely inside a branch, and a human approves the merge.

The incidents don't argue for keeping agents on a short leash forever. They argue for putting the leash where the model can't untie it: in the operating system, the network and the credentials, rather than in the prompt or a classifier. Written instructions such as CLAUDE.md and AGENTS.md files shape behavior, but they don't enforce anything.

A practical checklist before you walk away

  1. Run it in a boundary that holds even if the model is fooled: an OS sandbox (Claude Code /sandbox, Codex workspace-write), a container or a VM. Don't use bypass flags outside one.
  2. Default-deny the network. Allow only the registries and APIs the task needs. Avoid broad domains.
  3. Use scoped, short-lived credentials. Never mount ~/.ssh, cloud credential files or production database URLs into the agent's environment.
  4. Keep production out of reach. Separate dev and prod databases and deploy keys. Agents propose, and pipelines with human review deploy.
  5. Use branch protections and required reviews. Agents push to their own branches, and they can't approve or merge their own work.
  6. Treat everything the agent reads as untrusted, including issues, READMEs, web pages, tool output and memory files. Avoid the trifecta of private data, untrusted input and an exfiltration path in one session.
  7. Don't let untrusted users trigger privileged bots. Review CI agents' allowed_non_write_users-style settings, tool lists and shared caches.
  8. Turn off escape hatches (allowUnsandboxedCommands: false, disableBypassPermissionsMode) in managed settings for teams.
  9. Log and review: session logs, diffs, commits and audit events. Set spend and time limits.
  10. Pin and audit agent tooling like any other dependency. Extensions and CLIs are supply-chain targets.

For how these controls differ across products, see our comparison of Claude Code, Codex and Cursor.

SourcesConfigure the sandboxed Bash tool — Claude Code docs (Anthropic) · Configure permissions — Claude Code docs (Anthropic) · Choose a permission mode — Claude Code docs (Anthropic) · Development containers — Claude Code docs (Anthropic) · Making Claude Code more secure and autonomous with sandboxing — Anthropic (Oct 20, 2025) · Agent approvals & security — Codex docs (OpenAI) · About Copilot cloud agent — GitHub Docs · Risks and mitigations for Copilot cloud agent — GitHub Docs · Customize the agent firewall — GitHub Docs · AWS Security Bulletin AWS-2025-015 — Amazon Web Services · Nx security advisory GHSA-cxm3-wv7p-598c — GitHub · Incident report: unsanctioned agent behaviour during cyber testing — UK AI Security Institute (Aug 4, 2026) · GitHub MCP exploited — Invariant Labs (May 26, 2025) · Clinejection — Adnan Khan (Feb 9, 2026) · Snowflake AI escapes sandbox and executes malware — PromptArmor · Breaking Claude Code Opus 5 and auto mode — Embrace The Red (Aug 26, 2026) · AI coding tool Replit wiped database — Fortune (Jul 23, 2025) · Replit/SaaStr vibe coding incident — The Register (Jul 21, 2025) · Hacker plants computer-wiping commands in Amazon's AI coding agent — 404 Media (Jul 23, 2025) · s1ngularity supply chain attack — Wiz · Nx build system package compromised — StepSecurity · GitHub MCP exploited — Simon Willison (May 26, 2025) · Clinejection — Simon Willison (Mar 6, 2026) · Snowflake Cortex AI escapes sandbox — Simon Willison (Mar 18, 2026) · Incident report — Simon Willison (Aug 5, 2026) · Breaking Claude Code Opus 5 auto mode — Simon Willison (Aug 27, 2026) · Measuring AI Ability to Complete Long Software Tasks (METR) — arXiv

About this storyBased on the sources linked below. Editorial standards

Was this useful?Report an error
TopicCoding agentsCoding agents are AI systems that take a software task, such as fixing a bug or adding a feature, and carry it out… 15 stories, 2 guides, 1 comparisons.Open the hub
Comments
0

More on Coding agents & agents

The Week in AI

Get the cluster, not just the headline.

0