Tobias Kaisar and Aritra Dhar, researchers at Huawei's Computing System Labs in Zurich, built an automated attack that slips malicious skills past AI agent scanners up to 97% of the time S¹. The scanners they defeated, including NVIDIA's SkillSpector, pair static code checks with an LLM-based semantic judge, the exact two-layer defense vendors are building to stop malicious agent plugins S¹. Their preprint, posted on arXiv on 30 September, has not been peer-reviewed S¹.

My read: This is the first attack I've seen that specifically targets the two-stage scanner architecture rather than one layer in isolation. I'm skeptical of the 97% figure because it comes from a frozen detector, a setup that hands the attacker full knowledge of the defense. The 77% against a co-adaptive detector, where the scanner tries to adapt, is the number that matters, and it is still alarmingly high. The technique of moving payloads from code into natural language is simple enough that I expect it to work against most current scanners regardless of which model sits behind the judge.

The problem: agent skills are plug-ins anyone can write

Skills extend an AI agent's capabilities by injecting instructions and information into its context window, and they are widely used by agents such as OpenClaw and Claude Code S¹. Claude Code, Anthropic's terminal-based coding agent, has attracted more than 145,000 stars on GitHub P⁵, and its skill ecosystem lets third parties contribute capabilities the core team never reviewed.

Prior research has shown that third-party marketplaces already host malicious skills that give attackers direct influence over a victim's agent S¹. A separate study from Hong Kong University of Science and Technology, titled "Cloak and Detonate," found that LLM coding agents increasingly rely on third-party skills from public marketplaces that execute with the agent's full privileges P³. As we found when LLM security agents act but lack guardrails, the gap between what agents can do and what constrains them keeps widening.

How Pretext defeats a two-stage scanner

The emerging defense against malicious skills scans them before installation, pairing deterministic static checks with an LLM-based semantic judge, as implemented in NVIDIA's SkillSpector S¹. The static layer looks for suspicious code patterns. The LLM layer reads the skill's description and instructions to judge whether the stated purpose matches what the skill actually does.

Pretext defeats both stages with three moves S¹. First, it moves the malicious payload out of code and into natural language, leaving the static analysis with nothing to flag. Second, it frames the payload as the skill's legitimate purpose, so the LLM judge sees a coherent, benign-looking description. Third, it splits instructions across multiple files, keeping each piece small enough that the LLM stage stays below its blocking threshold.

The attack is white-box, meaning the attacker knows the detector's configuration and model. It iteratively crafts skills, testing each version against the scanner and refining until the skill passes detection while still delivering its payload and performing its stated benign task S¹.

The numbers that matter

Across three open-source models, Pretext achieved up to 97% success against a frozen detector and 77% against a co-adaptive detector S¹. A frozen detector is one the attacker has full knowledge of and the scanner does not change. A co-adaptive detector tries to adjust its defenses in response to the attack, a harder target.

Pretext attack success rate by detector type

Neither number is reassuring.

What would have to hold for this to survive review

The paper has not been peer-reviewed S¹. The experiments were conducted on three open-source models under controlled conditions, not against closed-source or proprietary systems S¹. The attack has not been demonstrated against Claude Code or OpenClaw specifically, only against the detection frameworks tested on the three open-source models S¹. NVIDIA's SkillSpector is referenced as an example of the two-stage architecture, but the paper does not claim SkillSpector is widely deployed in production S¹.

The authors state that their findings reveal major gaps in current skill scanners S¹. That claim rests on the three open-source models tested, and nobody outside the lab has checked the work yet.

Who should pay attention

Any team building or running an agent skill marketplace. If you operate a scanner that pairs static checks with an LLM judge, Pretext describes a specific bypass: payloads in natural language, split across files, framed as legitimate purpose. A security engineer at a coding-agent vendor could use the paper's three-technique breakdown as a checklist to test whether their scanner catches each move individually.

As we noted when OpenAI said coding agents reshape its own AI research, the shift toward agents that execute tasks autonomously makes the skill layer a growing attack surface. And as we saw when a new system spawned Claude Code agents with pre-loaded memory, the convenience of pre-configured agent capabilities is exactly what makes third-party skills attractive to both users and attackers.

The paper is available on arXiv at 2609.39607, posted 30 September 2026 S¹. No peer review date or conference submission has been announced.


Sources: S1 — Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents · P2 — Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents · P3 — Cloak and Detonate: Scanner Evasion and Dynamic Detection of Agent Ski · P4 — Cloak and Detonate: Scanner Evasion and Dynamic Detection of Agent Ski · P5 — anthropics/claude-code


Written from 5 sourced items, 4 of them primary.