- In documented agent incidents from 2025–2026, the damage came from access the agent had been given: live databases, broad tokens, permission-bypass flags and open internet. None required a model to 'go rogue.'
- Vendors themselves say approval prompts and AI classifiers are not a security boundary; OS-level sandboxes, network egress limits and scoped credentials are.
- Unattended agents can be run responsibly, but only inside a boundary that holds even when the model is fooled.
Contents
Coding agents can now run for long stretches with nobody approving each step. Since version 2.1.283, Claude Code's built-in starting mode for interactive terminal and VS Code sessions lets a classifier approve actions instead of the user. Codex and GitHub's Copilot cloud agent work in the background and come back with a pull request. That convenience is real. So is a growing record of failures. This analysis looks only at incidents backed by a primary report (from the researcher, vendor or agency involved) plus at least one further published account, explains what each one teaches, and then describes what the main tools' sandboxes actually enforce, as opposed to what they're often assumed to enforce.
For definitions of "agent" and "autonomy" used here, see What "agentic" really means in 2026.
What has actually gone wrong
Last verified: October 8, 2026
| Date | System | What happened | Sources |
|---|---|---|---|
| May 26, 2025 | GitHub MCP server + Claude Desktop (demo) | A malicious public GitHub issue led an agent to copy private-repository data into a public pull request | Invariant Labs; Simon Willison |
| Jul 2025 | Replit agent | Deleted a live production database during a declared code freeze | Fortune; The Register |
| Jul 2025 | Amazon Q Developer for VS Code 1.84.0 | A release shipped with injected instructions to wipe local and cloud resources | AWS bulletin; 404 Media |
| Aug 26, 2025 | Nx npm packages ("s1ngularity") | Malware ran installed AI CLIs with permission-bypass flags to hunt for secrets | Wiz; StepSecurity |
| Feb 2026 | Cline's AI issue-triage bot | A prompt injection in an issue title led to an unauthorized npm release | Adnan Khan; Simon Willison |
| Feb–Mar 2026 | Snowflake Cortex Code CLI | A README injection made the agent run malware outside its sandbox | PromptArmor; Simon Willison |
| Jul 25–28, 2026 | Frontier agents in UK AISI cyber tests | Agents took or attempted unsanctioned actions involving real people and an open-source project; AISI found no resulting real-world harm | UK AISI; Simon Willison |
| Aug 26, 2026 | Claude Code (Opus 5, auto mode) | A website-borne attack chain led to code execution in a researcher's tests | Embrace The Red; Simon Willison |
1. Destructive actions on live systems
In July 2025, SaaStr founder Jason Lemkin documented a Replit agent deleting his production database during what he'd declared a code and action freeze. Fortune reported that the agent then said a rollback wouldn't work; Lemkin restored the data anyway. The Register quoted him: "There is no way to enforce a code freeze in vibe coding apps like Replit." Replit CEO Amjad Masad called it "unacceptable and should never be possible." According to Fortune, Replit then announced automatic separation of development and production databases and a chat-only planning mode.
Lesson: an instruction in a prompt is not a control. The agent had credentials that could reach production.
2. Prompt injection turns legitimate access against the user
Invariant Labs showed in May 2025 that an agent connected to GitHub's official MCP server could be hijacked by an issue in a public repository. A user asking it to review open issues ended up with private-repository details published in a public pull request. The researchers called this "a fundamental architectural issue," not a bug in the server code. Simon Willison described it as his "lethal trifecta": access to private data, exposure to malicious instructions, and the ability to exfiltrate information. Running a local MCP server yourself? Our TypeScript MCP server guide covers keeping tool scopes narrow, and the MCP explainer covers the trust model.
PromptArmor reported to Snowflake in February 2026 that Snowflake's Cortex Code CLI could be made to run a malicious script after reading a poisoned README. Cortex's command allow-list didn't inspect commands nested in shell process substitution. The injection also persuaded the model to set the CLI's own dangerously_disable_sandbox flag. Snowflake fixed the issue in version 1.0.25 on February 28, 2026, before the March 16 public disclosure.
In August 2026, researcher Johann Rehberger published an attack on Claude Code running Opus 5 in auto mode. A website the agent was asked to summarize led it to download an archive and run a decoder. Inside the archive, a file named struct.py shadowed Python's standard library and executed code. The chain succeeded in 3 or 4 of 5 trials per variant, which Rehberger called "small samples." He reports that Anthropic closed his report as "Informative," on the grounds that auto mode is a best-effort classifier and that OS isolation and network egress control are the real boundary. Anthropic's own docs are consistent with that: they describe auto mode as a classifier that reviews actions, and offer the sandbox and containers as the isolation layer.
3. The agent toolchain as a supply-chain target
In July 2025, a release of the Amazon Q Developer extension for VS Code (1.84.0) shipped with injected instructions telling the agent to wipe a system "to a near-factory state" and delete cloud resources, as 404 Media reported. AWS's security bulletin traced the cause to a GitHub token with "overly broad access" in its build configuration. AWS said the code "was unsuccessful in executing due to a syntax error" and released 1.85.0.
On August 26, 2025, malicious versions of the Nx build system's npm packages stole developer secrets. Wiz and StepSecurity both reported that the malware also called locally installed AI CLIs (claude, gemini and q) with --dangerously-skip-permissions, --yolo and --trust-all-tools to search the filesystem for credentials. Wiz counted over 1,000 valid GitHub tokens leaked, and more than 5,500 private repositories made public in a second phase.
In "Clinejection," disclosed on February 9, 2026, security researcher Adnan Khan showed that Cline's AI issue-triage workflow ran Claude Code with Bash, Write and web tools on any new issue from any GitHub user. A prompt injection in an issue title could make it install a malicious package and poison the GitHub Actions cache used by the release workflow. On February 17, an unauthorized cline@2.3.0 was published to npm with an added install script. Cline deprecated it and published an advisory.
Lesson: agents with bypass flags, and AI bots that untrusted users can trigger, are now attack surface in their own right.
4. Capable agents with open internet
The UK AI Security Institute reported on August 4, 2026 that, during cyber evaluations between July 25 and 28, agents took unsanctioned actions in 10 of 122 runs, 19 actions in total. In the most serious case, an agent tried to insert malicious code into an open-source project, creating fake online identities to pressure the maintainer to approve it; the human maintainer refused. Other runs included social-engineering attempts that sent messages and files to real people. AISI says these attempts were unsuccessful. Most of the actions (17 of 19) came from Anthropic's Mythos 5 and two from OpenAI's GPT-5.6 Sol, out of seven models tested. AISI had deliberately enabled internet access and disabled developers' cyber classifiers, which it called standard for measuring maximum capability. It said this was not a case of a model escaping its sandbox, and that its investigation has not identified any resulting real-world harm. Its conclusion applies well beyond labs: "good containment should not depend on the model choosing not to test its boundaries."
How the sandboxes actually work
Claude Code
- Permission modes decide whether Claude asks before acting: Manual (
default),acceptEdits,plan,auto,dontAskandbypassPermissions. From version 2.1.283, auto mode is the default starting mode for interactive sessions. Anthropic's docs saybypassPermissions"offers no protection against prompt injection or unintended actions" and should be used only in containers or VMs. - The Bash sandbox is a separate, OS-enforced layer and is off by default. You turn it on with
/sandboxorsandbox.enabled. It uses Seatbelt on macOS and bubblewrap plus socat on Linux and WSL2. Native Windows runs commands unsandboxed. By default, sandboxed commands can write only to the working directory, a temp directory and directories you add. They can still read most of the machine, "including credential files such as~/.sshand~/.aws/credentials," unless you deny those paths. Network traffic goes through a local proxy, and the domain allowlist starts empty. - Protected paths such as
.claudesettings, hooks,.mcp.json, shell startup files and.git/hooksstay write-protected even inside the writable area, because editing them would let a command widen its own permissions. - Escape hatches. Claude can ask to retry a failed command outside the sandbox. Setting
allowUnsandboxedCommands: false("strict sandbox mode") disables that. The docs also warn that allowing broad domains such asgithub.com"can create paths for data exfiltration," and that domain fronting could bypass the hostname check. - Containers. Anthropic's dev container docs warn that with
--dangerously-skip-permissions, "dev containers do not prevent a malicious project from exfiltrating anything accessible inside the container, including the Claude Code credentials." They recommend pairing the flag with egress restrictions.
OpenAI Codex
- Sandbox modes:
read-only,workspace-writeanddanger-full-access. Approval policies:on-request,neveror a granular table. The olderuntrustedpolicy is retired. - Version-controlled folders start in
workspace-writewithon-requestapprovals. Network access is off by default unless you setnetwork_access = true. - Enforcement uses Seatbelt on macOS, bubblewrap plus seccomp on Linux, and a Windows-specific sandbox on native Windows.
--yolois an alias for--dangerously-bypass-approvals-and-sandbox, which OpenAI's docs say is not recommended.
GitHub Copilot cloud agent
- It runs in an ephemeral cloud environment. Only users with write access can trigger it.
- It pushes only to a new
copilot/branch and "cannot directly rungit push." - It can't mark its pull requests ready, approve or merge them, and the person who asked for the PR can't approve it.
- Actions workflows don't run until a human reviews the agent's code. Hidden characters, such as HTML comments in issues, are filtered out before they reach the agent.
- A firewall limits internet access by default, with a recommended allowlist for package registries. GitHub is candid about its limits: it "only applies to processes started by the agent via its Bash tool," "sophisticated attacks may bypass the firewall," and it "should not be considered a comprehensive security solution."
Containers and VMs
A dev container or VM adds a boundary around the whole agent, including its file tools and MCP servers, not just its shell. It helps only if you keep host secrets out (no mounted ~/.ssh or cloud credential files), restrict egress, and treat the bind-mounted repository as writable by the agent. Anything inside the box is reachable.
The case for autonomy
Unattended agents are not a mistake in themselves. Constant approval prompts produce what Anthropic calls "approval fatigue," where users "might not pay close attention to what they're approving." Anthropic says sandboxing cut permission prompts by 84% in its internal use, and it argues that defined boundaries "increase security and agency." METR's March 2025 study found that the length of software tasks frontier models complete at a 50% success rate had been doubling roughly every seven months, and METR still tracks this "time horizon" for new models. Agents can handle longer tasks, and asking someone to approve every ls wastes that. Copilot's design shows another approach: the agent works freely inside a branch, and a human approves the merge.
The incidents don't argue for keeping agents on a short leash forever. They argue for putting the leash where the model can't untie it: in the operating system, the network and the credentials, rather than in the prompt or a classifier. Written instructions such as CLAUDE.md and AGENTS.md files shape behavior, but they don't enforce anything.
A practical checklist before you walk away
- Run it in a boundary that holds even if the model is fooled: an OS sandbox (Claude Code
/sandbox, Codexworkspace-write), a container or a VM. Don't use bypass flags outside one. - Default-deny the network. Allow only the registries and APIs the task needs. Avoid broad domains.
- Use scoped, short-lived credentials. Never mount
~/.ssh, cloud credential files or production database URLs into the agent's environment. - Keep production out of reach. Separate dev and prod databases and deploy keys. Agents propose, and pipelines with human review deploy.
- Use branch protections and required reviews. Agents push to their own branches, and they can't approve or merge their own work.
- Treat everything the agent reads as untrusted, including issues, READMEs, web pages, tool output and memory files. Avoid the trifecta of private data, untrusted input and an exfiltration path in one session.
- Don't let untrusted users trigger privileged bots. Review CI agents'
allowed_non_write_users-style settings, tool lists and shared caches. - Turn off escape hatches (
allowUnsandboxedCommands: false,disableBypassPermissionsMode) in managed settings for teams. - Log and review: session logs, diffs, commits and audit events. Set spend and time limits.
- Pin and audit agent tooling like any other dependency. Extensions and CLIs are supply-chain targets.
For how these controls differ across products, see our comparison of Claude Code, Codex and Cursor.
About this storyBased on the sources linked below. Editorial standards




