Blog
The TrapDoor Attack: How Poisoned Instruction Files Turn Your Coding Agent Against You
A sophisticated supply chain attack targets developer environments by planting invisible instructions in project-level AI configuration files like .cursorrules and CLAUDE.md.
· updated 1 October 2026
The TrapDoor supply chain attack is a sophisticated campaign that compromises developer environments by modifying project-level AI instruction files like .cursorrules and CLAUDE.md. After initial infection via malicious packages across npm, PyPI, and Crates.io, the malware uses hidden Unicode characters, invisible to human developers but parsed by LLMs, to instruct coding agents to secretly exfiltrate sensitive data. The poisoned files are a local persistence mechanism on the already-compromised machine, alongside git hooks, shell hooks, systemd and cron entries. A separate attempt to spread them through pull requests against open-source AI projects was caught by GitHub's hidden-Unicode detection and closed before any were merged.
The Mechanics of the TrapDoor Compromise
In May 2026, the TrapDoor campaign executed a synchronized strike across package registries including npm, PyPI, and Crates.io. Like many supply chain attacks, it began with malicious packages masquerading as legitimate dependencies. When a developer installed one of these packages (e.g. via npm install or pip install), a postinstall script executed with the developer's local privileges. The initial execution behaved predictably for a credential harvester, scraping environment variables and known credential paths like ~/.aws/credentials and ~/.ssh/id_rsa. Its main target was crypto wallet data for chains such as Solana, Sui and Aptos; the instruction-file poisoning was a newer technique layered on top.
However, TrapDoor introduced a novel secondary phase designed specifically for the era of AI-assisted development. After the initial credential theft, the malware actively sought out project-level AI instruction files, specifically targeting .cursorrules for Cursor users and CLAUDE.md for Claude Code environments. The script injected malicious instructions into these files, establishing a persistent foothold directly within the repository's configuration.
These injected instructions directed the AI coding agent to periodically bundle and transmit sensitive data under the guise of routine operations. The agent, assuming the instructions were legitimate project guidelines, would silently execute exfiltration routines during normal development tasks. Stolen data included AWS credentials, GitHub personal access tokens (ghp_*), SSH keys, and local .env variables, all funneled out through the agent's standard execution environment.
Exploiting LLMs with Hidden Unicode
The most insidious aspect of the TrapDoor attack was its method of hiding the malicious instructions. Attackers utilized a combination of zero-width joiners (U+200D), zero-width non-joiners (U+200C), and invisible Unicode characters like U+200B. To a human developer reviewing a pull request on GitHub or inspecting the .cursorrules file in an editor, the file appeared completely normal, perhaps with some extra trailing whitespace at the end of a line.
Large Language Models, however, process text differently. The underlying tokenizers (such as tiktoken or Anthropic's byte-pair encoders) tokenize and parse these hidden characters just as they would visible text. The malware encoded the exfiltration commands using these invisible characters, effectively delivering a prompt injection payload that the AI could read perfectly while remaining entirely hidden from human oversight.
This technique completely bypasses standard code review and diff inspection. A developer scrutinizing a commit would see no suspicious plain text commands. The malware effectively created a covert channel directly into the LLM's context window, weaponizing the agent's ability to interpret and act on project-level instructions without raising alarms during human inspection (see what a hostile instruction can actually reach).
A New Paradigm of Persistence and Propagation
Traditional supply chain attacks often rely on remaining resident in memory, modifying local system binaries, or hiding in deeply nested node_modules dependencies. TrapDoor represents a shift because it achieves persistence by modifying the source code repository itself, specifically targeting files that are explicitly trusted and tracked in version control.
In the observed campaign, TrapDoor wrote the poisoned .cursorrules and CLAUDE.md files into the compromised developer's own project directories as one of several local persistence mechanisms. A more ambitious attempt to plant the same files in shared repositories went through pull requests against six open-source AI projects, including LangChain and LlamaIndex. GitHub's hidden and bidirectional Unicode warnings flagged all six, and none were merged.
What made TrapDoor novel was not spreading through git, since the one attempt at that was caught, but treating .cursorrules and CLAUDE.md as just another persistence location next to git hooks, shell hooks, systemd and cron. Because these files are committed to share conventions across a team, a poisoned one that does get merged would reach every developer who pulls it, which is why review of instruction-file diffs matters.
The Instruction File as an Attack Surface
The rise of AI coding agents has introduced new, highly privileged configuration files into the standard repository structure. Files like .cursorrules, CLAUDE.md, and .github/copilot-instructions.md are designed to provide deep context and operational guidelines to LLMs. They are, by definition, executable instructions, even if written in natural language.
Agents treat the contents of these files with a high degree of trust, often prioritizing them over general system prompts. TrapDoor demonstrates that this trust boundary is actively being exploited. Attackers now recognize these instruction files as a viable attack surface: a place to inject commands that the agent will execute with the full privileges of the user running the editor or CLI.
This incident builds upon the precedent set by earlier campaigns like the Nx s1ngularity attack, which similarly targeted developer environments but relied on different mechanisms to exploit already-installed agents. TrapDoor refines this approach by targeting the configuration files themselves, highlighting a critical vulnerability in how AI context is managed and secured (compare this with how AI code leaks secrets at 2x human rate).
Architectural Approaches to Blast Radius Reduction
Mitigating attacks like TrapDoor requires architectural changes to how developer environments are structured. The core vulnerability is not just the execution of a malicious script, but the unrestricted access that script has to the developer's entire machine. Limiting the blast radius of any single compromised project is essential.
One approach is strict per-workspace credential scoping. In this model, an AI agent operating within a specific project repository only has access to the credentials explicitly provisioned for that project. It cannot read the developer's global ~/.aws/credentials or access environment variables belonging to other projects. This containment strategy ensures that a poisoned .cursorrules file in one repository cannot compromise the developer's broader infrastructure access.
Implementing isolated execution environments for AI agents is another critical layer. For example, Forkbench utilizes a Thread model combined with macOS Seatbelt sandboxing where agent actions are confined to specific filesystem and network boundaries. By architecturally separating the agent's execution context from the developer's global state, the potential damage of a supply chain compromise is significantly contained, preventing lateral movement across projects.
Understanding the Limits of Protection
While architectural isolation and credential scoping reduce the blast radius, they do not eliminate the fundamental risk of running untrusted code. No security tool can completely prevent a malicious postinstall script from executing if a developer intentionally installs a compromised package and that script runs with their user privileges.
If a postinstall script executes as your user, it can read any file your user can read, including ~/.ssh, local project .env files, and browser data. It can also modify any file your user can write to, which is how TrapDoor poisons the .cursorrules file in the first place. Sandboxing can restrict what the agent does afterward, but the initial compromise relies on the package manager executing code.
Security in this context requires a defense-in-depth approach. Vaults and secret managers protect the credentials they hold (see how credential brokerage works), and sandboxing protects against rogue agent actions, but developers must still exercise caution when installing dependencies. Relying solely on downstream containment without addressing upstream package integrity leaves a critical gap in the security posture.
Related: What a hostile instruction can actually reach, The Nx attack that used agents already installed, AI-assisted code leaks secrets at 2x human rate
Frequently asked
What is the TrapDoor supply chain attack?
The TrapDoor attack is a campaign that compromises developer environments by modifying AI instruction files like
.cursorrulesandCLAUDE.md. It uses hidden Unicode characters to secretly instruct coding agents to exfiltrate sensitive data, and keeps them as one of several persistence mechanisms on the infected machine. Its attempt to spread them through pull requests to open-source projects was caught before any were merged.Can a .cursorrules file be malicious?
Yes. Because agents treat
.cursorrulesas trusted instructions, attackers can inject commands into this file that instruct the AI to perform malicious actions, such as reading sensitive files like.envand transmitting them to external servers (see why AI code leaks secrets and MCP config risks).Can CLAUDE.md contain hidden instructions?
Yes. Attackers can use invisible Unicode characters and zero-width joiners (
U+200B,U+200D) inCLAUDE.md. While these characters are invisible to humans reviewing the file, LLM tokenizers parse them normally and will execute the hidden instructions.How do I check for hidden Unicode in my instruction files?
You can use hex editors, specialized text editor plugins, syntax tree inspection like Tree-sitter, or command-line tools like
cat -v,hexdump -C, orxxdto inspect the raw bytes of your configuration files for unexpected Unicode sequences.Does sandboxing protect against supply chain attacks on coding agents?
Sandboxing limits the blast radius by preventing a compromised agent from accessing your entire machine (see sandbox AI coding agents on macOS). However, it cannot stop a malicious package's
postinstallscript from running and accessing files if that script executes directly on your host machine with your user privileges, which is why isolating credentials in a secure vault is critical.