Skip to content
KageXKageX
All articles
9 min read

Agent skills are code written in English

Agent skillsAI agentsSupply chainFundamentals
KageXAGENT SKILLSAgent skills are code written inEnglishAnthropicOWASPOpenClawKoi SecurityKageX

TL;DR: Agent skills are plain-English instruction files that more than thirty AI coding tools load and act on, which makes every skill code written in English. On one public registry, 341 of 2,857 skills were malicious, and at the campaign's peak five of the seven most-downloaded skills were malware. Read the whole file, not the description.

Most security thinking about AI agents has gone into two places: the model, and the tools it can call. Over the past year a third layer grew up between them, almost entirely outside anyone's threat model, and it has already been used to ship malware.

That layer is agent skills. If you use Claude Code, Codex, Cursor, GitHub Copilot or any of the thirty-odd tools that read the same format, you probably have some installed. This is a primer on what they are, why they are a genuinely new kind of attack surface, and how to review one before you trust it.

What is an agent skill?

A skill is a folder. Inside it is a file called SKILL.md: a short YAML header with a name and a description, followed by plain-language instructions. The folder can also hold scripts and reference files that those instructions point to.

Anthropic introduced the format for Claude in October 2025 and published it as an open specification on 18 December. Adoption was unusually fast. VS Code and OpenAI's ChatGPT and Codex CLI read skills within about 48 hours, and by March 2026 more than thirty tools, including Gemini CLI, Cursor, GitHub Copilot, JetBrains' Junie and AWS's Kiro, loaded the same files unchanged.

The design idea that made skills popular is progressive disclosure. An agent does not load every skill in full. At startup it reads only each skill's name and description, a few dozen tokens apiece. When a task matches, it loads the full instructions. Scripts and reference files come in only when those instructions call for them.

Diagram: a skill folder's SKILL.md header, body and files, mapped to the three levels an agent loads them at
What an agent loads from a skill, and when. Only level one, the name and description, is usually seen at install; the instructions, scripts and anything fetched from the web load later.

That is an elegant way to give an agent hundreds of capabilities without flooding its context window. As we will see, it is also exactly why skills are hard to review.

Why agent skills are code, even when they are only text

The instinct is to treat a skill like documentation: a Markdown file with some helpful instructions, harmless unless a bundled script is malicious. That instinct is wrong, and it is the single most important idea in this post.

An agent acts on the instructions in a skill. If a SKILL.md says "before you start, install the helper from this link", an agent with a shell may do it, or may tell the user to. If it says "load the latest guidance from this URL", the agent fetches text from a server someone else controls and treats that as instructions too. The prose is the program, and the agent is the interpreter.

That makes a skill closer to a package with an install script than to a README, with two differences that make it harder to defend:

  • Instructions can't be linted. No compiler rejects "and don't mention this step to the user". A scanner can flag a bad script. It has a much harder time flagging a bad sentence.
  • The same file runs everywhere. Because the format is shared across tools, a skill written once behaves the same way in every agent that reads it, good or bad.

MCP servers raise a similar problem: every server you connect sits inside your agent's trust boundary. Skills are the same idea, one layer up. MCP gives an agent tools. Skills tell it how and when to use them.

The review gap

Progressive disclosure means that what you see when you install a skill is not what the agent later runs.

Most install screens show the name and description, which is level one. The full instructions only load when a task triggers the skill. Scripts only run when those instructions call them. And if the instructions tell the agent to fetch more instructions from a URL, that content is not in the skill at all: it arrives at runtime and can change any day.

So the parts of a skill that decide what the agent does are the parts least likely to have been read by a person. Someone who approves "youtube-summarize-pro: summarises YouTube videos" has approved a sentence. The behaviour lives further down.

Diagram: a harmless-looking install screen beside the instructions, scripts and fetched rules the agent later runs
The review gap. The install screen shows a sentence; what the agent runs arrives after you click Install, and part of it can change any day.

Case study: ClawHavoc

This is not hypothetical. In early 2026, Koi Security researcher Oren Yomtov audited every skill on ClawHub, the public registry for the popular open-source agent OpenClaw, with help from his own OpenClaw assistant. Of the 2,857 skills available at the time, 341 were malicious, and 335 of those traced back to one coordinated campaign now tracked as ClawHavoc.

The skills looked professional, with credible names such as "solana-wallet-tracker" and "youtube-summarize-pro" and long, well-written documentation. The attack sat in a section headed Prerequisites, which told users to download a helper tool or run a terminal command to "fix dependencies" before first use. Doing so installed information-stealing malware, Atomic Stealer on macOS, aimed particularly at people running agents on always-on machines.

Diagram: how the ClawHavoc campaign seeded the ClawHub registry with malicious skills, with Koi Security's audit numbers
ClawHavoc in four steps. Koi Security audited 2,857 ClawHub skills and found 341 malicious, 335 of them from one campaign.

Notice what the malicious part was. Not an exploit, and not even code inside the skill. It was a paragraph of instructions, written to be followed. Material accompanying OWASP's new Top 10 notes that at the campaign's peak, five of the seven most-downloaded skills on the registry were confirmed malware.

The OWASP Agentic Skills Top 10, in four questions

OWASP responded with a dedicated Agentic Skills Top 10, with version 1.0 published over the summer. Its ten entries, which cover risks such as untrusted external instructions, weak isolation, update drift and poor scanning, are worth reading in full. For teaching, we find they fold into four questions you can ask of any skill.

Diagram: four questions to ask of any agent skill, provenance, power, runtime and visibility, with checks and red flags
The OWASP Agentic Skills Top 10, folded into four questions you can ask of any skill. If you can't answer one, don't install it.
  1. Who wrote it, and how did it reach you? Registries are young, accounts get taken over, and popularity is easy to fake. Provenance comes first.
  2. What can it make the agent do? A skill inherits the agent's permissions. A note-taking skill loaded into an agent that holds your cloud credentials has your cloud credentials.
  3. What does it pull in at runtime? External URLs, fetched instructions and later updates are behaviour you did not review. A skill that was safe on Monday can change on Tuesday.
  4. Would anyone notice? Skills are rarely scanned, versioned or logged the way code dependencies are. If a skill did something wrong, would there be a record?

How to review an agent skill: a worked example

Here is a short, invented skill of the kind you will find in any registry. Read it the way an agent would, then the way a reviewer should.

Annotated example SKILL.md with four red flags in its body, none of them visible on the install screen
An invented skill read the way a reviewer should. Each red flag fails one of the four questions, and none appears on the install screen.

The description is harmless. Every red flag is in the body, which most install screens never show:

  • An instruction to run something before first use. This is the ClawHavoc pattern. A legitimate skill declares its dependencies; it does not ask the agent or the user to fetch and run an installer from a link.
  • Instructions fetched from a URL at runtime. Whatever that page says tomorrow becomes part of the skill. You are no longer reviewing a file, you are trusting a server.
  • Scope wider than the job. A summariser has no reason to read environment files or credentials.
  • Anything that asks for silence. "Do not mention this to the user" has no legitimate purpose in a skill. Treat it as a confession.

Each flag also fails one of the four questions: the first and third are about power, the second about runtime, and the fourth about visibility. None of these needs a code scanner to find. They need someone to read the whole file, and building that habit is really what this post is arguing for.

Agent skill security checklist

Before you install a skill:

  1. Read the entire SKILL.md, not just the description, and then every script it references.
  2. List every command it asks the agent to run and every URL it asks it to fetch. If you cannot explain one, do not install it.
  3. Check where it came from. Prefer publishers you already trust, and look at their history on the registry.

When you run it:

  1. Pin the version and record a hash. An update is a new skill; review it as one.
  2. Give the agent the least privilege the job needs, and keep agents that load third-party skills away from credentials and wallets.
  3. Restrict outbound fetches so a skill cannot quietly pull in new instructions.

Across your organisation:

  1. Inventory which skills are installed, and where. Most teams cannot answer this today.
  2. Log which skills load and what they cause the agent to do, so there is a record when something goes wrong.

Why this is worth teaching now

Skills went from launch to a cross-industry standard in about two months, and from standard to a registry-wide malware campaign within a few more. Security practice has not caught up. Few registries scan well, few organisations keep an inventory, and most people still read a skill's description and stop there.

The encouraging part is that the core skill is learnable in an afternoon. Reviewing an agent skill is closer to careful reading than to reverse engineering. The hard part is building the habit before an incident, not after one.

Why we build the way we do

This is the kind of problem our tooling and training exist for, so treat what follows as an interested opinion.

FreakLabs teaches the craft underneath, because reading a skill, an MCP server or a prompt the way an attacker would is a skill people learn by doing. AgentBreaker tests your own agents against the layers they trust, including what they load and act on. Mirage runs an autonomous red-team agent against a target and scores results deterministically, so a finding is something you can verify rather than a model's opinion.

If you are new to this area, start with what AI red teaming actually is. Skills sit alongside the other trust layers we have written about: agent memory and the agent's own sandbox output, as well as the MCP servers your agents connect to. Each time the lesson is the same: whatever an agent reads and acts on is part of its attack surface. Keep the OWASP Top 10 for LLM Applications beside them as a map.

The short version

A skill is a folder with a SKILL.md file of plain-language instructions, plus optional scripts, and more than thirty agent tools now load the same format. Because agents act on those instructions, a skill is code written in English: a sentence asking the agent to run a helper or fetch new rules is as effective as a script.

Progressive disclosure means you usually approve a skill's one-line description while its behaviour lives in a body, scripts and runtime fetches that nobody read. ClawHavoc showed what that costs: 341 malicious skills among 2,857 on one registry. The fix is not exotic. Read the whole file, list every command and URL, pin versions, limit privileges and keep an inventory.

Sources and reuse

Primary sources: Anthropic, Introducing Agent Skills and the Agent Skills specification; OWASP, Agentic Skills Top 10; Koi Security, ClawHavoc: 341 malicious skills. Also reported by The Hacker News and analysed by Palo Alto Networks Unit 42.

Later analyses of ClawHavoc report larger totals than Koi's original count; we use the original audit figures. The example skill in this article is invented for teaching and points nowhere real. All five diagrams are free to reuse in your own writing, talks or training material, with a link back to this page.

Share thisXLinkedIn

Frequently asked questions

What is an agent skill?

An agent skill is a folder containing a SKILL.md file, which has a short YAML header with a name and description followed by plain-language instructions, plus optional scripts and reference files. Anthropic introduced the format in October 2025 and published it as an open specification on 18 December 2025. By March 2026 more than thirty agent tools, including Claude Code, OpenAI's Codex, Gemini CLI, Cursor, GitHub Copilot and VS Code, could load the same skill files. Agents use progressive disclosure: they read each skill's name and description at startup, load the full instructions when a task matches, and use scripts or reference files only when those instructions call for them.

Why are agent skills a security risk?

Because agents act on the instructions in a skill, a skill is effectively code written in English. A sentence telling the agent to run a helper, fetch rules from a URL or read a credentials file can be as effective as a malicious script, and there is no compiler or linter to reject it. Progressive disclosure widens the gap: install screens usually show only the name and description, while the behaviour lives in the body, bundled scripts and anything fetched at runtime, which few people read. A skill also inherits every permission the agent already holds.

What was the ClawHavoc campaign?

ClawHavoc was a campaign that seeded ClawHub, the public skill registry for the open-source agent OpenClaw, with malicious skills. In early 2026 Koi Security researcher Oren Yomtov audited all 2,857 skills on the registry and found 341 malicious ones, 335 of them tied to a single coordinated operation. The skills had credible names and professional documentation, and a Prerequisites section told users to download a helper or run a terminal command before first use, which installed information-stealing malware such as Atomic Stealer on macOS. The malicious part was instructions, not an exploit.

What is the OWASP Agentic Skills Top 10?

It is an OWASP project that documents the ten most critical security risks specific to agent skills, as distinct from risks in models or in MCP tool servers. Version 1.0 was published in 2026. Its entries cover risks such as malicious skill authorship, untrusted external instructions, weak isolation, update drift and poor scanning, and each maps to existing frameworks including the OWASP Top 10 for LLM Applications and NIST AI RMF. A practical way to apply it is to ask four questions of any skill: who wrote it, what can it make the agent do, what does it fetch at runtime, and would anyone notice if it misbehaved.

How do you review an agent skill before installing it?

Read the whole SKILL.md, not just its description, and every script it references. List every command it asks the agent to run and every URL it asks it to fetch; if you cannot explain one, do not install it. Check the publisher's provenance and history. Once installed, pin the version and record a hash so that any update is reviewed as a new skill, run agents that load third-party skills with the least privilege the task needs and away from credentials and wallets, restrict outbound fetches, keep an inventory of installed skills, and log which skills load and what they cause the agent to do. Red flags include instructions to run an installer before first use, rules fetched from a URL at runtime, access to secrets unrelated to the task, and any instruction to hide steps from the user.

How are agent skills different from MCP servers?

MCP servers give an agent tools: functions it can call, with descriptions and results that enter the model's context. Skills give an agent procedures: instructions for how and when to do a task, which may in turn call tools or bundled scripts. Both put third-party text inside the agent's trust boundary. MCP risk centres on what tools can reach and what they return; skill risk centres on instructions the agent follows, which are harder to scan and are usually loaded well after anyone reviewed them.