The FakeGit campaign documented by security firm Island in July 2026 maps the scale of a new attack surface: roughly 7,600 fake GitHub repositories, 6,600 fraudulent profiles, and more than 14 million downloads, with over 800 repositories impersonating AI skills and MCP servers to distribute the SmartLoader dropper and the StealC infostealer. Fake repositories are an old problem. The twist was who recommended them: Gemini and ChatGPT independently pointed users to the same malicious walmart-mcp repository, complete with installation instructions.
Attackers no longer need to deceive users directly. They can deceive the assistants users trust. Two architectural traits make this possible. First, agents process instructions and external content as text, so a malicious line hidden in a README or tool description can be obeyed rather than analyzed, the classic indirect prompt injection. Second, agents can act on what they read. Security researcher Simon Willison's 'lethal trifecta' combines access to valuable data, exposure to untrusted content, and the ability to communicate externally, turning hostile text into a potential breach.
Researchers have cataloged a growing playbook. AgentBaiting targets the assistant itself, as in FakeGit, where the agents were never compromised, they simply recommended software whose credibility had been manufactured. Tool poisoning hides instructions inside MCP tool descriptions, demonstrated by Invariant Labs, where a malicious calculator tool manipulated a trusted email connector into forwarding messages to an attacker.
Some malicious skills go further: a 2026 academic study of 98,380 registry skills confirmed 157 as malicious and found recurring instructions like 'Do Not Mention This to the User,' letting agents report tasks complete while quietly exfiltrating data. Rug-pull attacks exploit accumulated trust: Koi Security found the postmark-mcp connector stayed harmless through version 1.0.15, then added a hidden BCC that copied outgoing emails from roughly 300 organizations to an attacker domain. Other vectors include config changes outside the package, like the MCPoison vulnerability in Cursor patched in version 1.3, and repository-controlled commands in agentic development environments, with Claude Code vulnerabilities CVE-2025-59536 and CVE-2026-21852 patched by Anthropic.
The ClickFix pattern needs no injection at all: during the ClawHavoc campaign, attackers disguised malicious commands as installation steps in README and SKILL.md files across the OpenClaw ecosystem, where audits found 341 malicious skills out of 2,857, and Antiy CERT later tracked 1,184 malicious skills to just twelve accounts. Most striking is agents turning attacker: Anthropic's reported GTG-1002 espionage campaign connected penetration-testing tools to Claude Code through MCP, with the model autonomously executing 80 to 90 percent of tactical operations, attributed to a state-sponsored group.
Underpinning it all is fabricated reputation. GitHub stars sell for three to ten cents each, researchers identified six million suspicious stars across nearly 16,000 repositories, and one developer lost about $500,000 after trusting a malicious extension with inflated download counts. Established registries eventually hardened themselves with mandatory two-factor authentication and verified provenance, but AI skill marketplaces are growing far faster than their security infrastructure. As agents gain access to email, databases, and credentials, verifying what they trust matters as much as securing where they run.
Comments (0)
Log in to join the discussion
Log InNo comments yet