Writing a skill description that actually triggers

A skill description triggers reliably when it names the request the way a user says it, and when the task is one the agent cannot finish with a single obvious edit. Across 480 should-fire runs on four Claude models, my question-shaped skill fired 120 of 120 times. The skill whose prompts read as ordinary edits fired 65 of 120.

Where the numbers come from

This site has four project skills: one scaffolds a post, one publishes a draft, one revises a published post, and one reviews search performance. A 36-case trigger suite fires cold prompts at a fresh headless session and a script grades which skill the agent invokes. Each skill has six prompts that should fire it and three adjacent decoys that should not.

On 2026-09-04 the suite ran five times on each of five models, which is 900 case-runs. That post reads the results by model. This one reads the same files by prompt, because the prompt-level rates say which phrasings a description catches. Unless I say otherwise, a rate here is out of 20: five runs on each of the four larger models. The smallest model has its own section.

The four descriptions and how often each fired

skill description opens with words should-fire rate
reviewing “Use when reviewing how umar.codes performs in Google Search…” 39 120/120
scaffolding “Scaffold a new post… Use whenever creating any new essay, reference page, or log entry” 43 101/120
publishing “Publish a draft post… Use whenever flipping any post from draft to published” 42 97/120
revising “Revise an already-published post… Use for ANY change to a published post, however the request is phrased” 81 65/120

The longest description has the lowest rate. Length did not buy reliability. The shortest one, which is the only one that is nothing but a list of situations, did not miss once.

Prompts that name the state change fired 40 of 40

The publishing description says the skill is for “flipping any post from draft to published”. Two of its six prompts name that state change, and both were perfect:

prompt rate
“Flip testing-skill-triggers to published” 20/20
“Let’s take the crawler logging post out of draft” 20/20
“logging-ai-crawlers is done, get it onto the live site” 16/20
“The dogfooding-specdx draft is ready — publish it” 15/20
“Ship the URL state essay live” 13/20
“Make the Aug 10 essay public” 13/20

The four prompts that describe the outcome scored 57 of 80, and all 23 misses were silence: no skill fired and the agent started on the task itself. “Ship”, “live”, and “public” are not in the description. “Publish it” is, and it still lost five runs, so shared vocabulary alone does not explain the table. What the two perfect prompts share is that they name the thing the skill changes, a status field, in the same terms the description uses.

The scaffolding skill shows the same pattern. Its description starts with the verb “Scaffold” and lists “essay, reference page, or log entry”. “Scaffold something so I can start writing about DuckDB log analysis” fired 20 of 20 and “Add a reference page” fired 19. “I want to write a new essay” fired 11, and 8 of its 9 misses went to a general-purpose brainstorming skill from an installed plugin. That miss is competition, not silence. My description says what the skill does; the competing one claims the user’s intent, and “I want to write” is an intent.

Listing request shapes fixed one prompt and left three broken

The revising description started at 40 words: “Use for any meaningful edit to a published post”. The prompt “The ag-grid post’s canonical link changed — sort it out” fired nothing, because it describes a fact that changed and never says “edit”. I rewrote the description to 81 words. It now claims “ANY change to a published post, however the request is phrased” and lists the shapes: errors, typos, versions, links, canonical URLs, frontmatter, new sections. The canonical-link prompt went from 0 of 1 before the rewrite to 5 of 5 after it, and held 17 of 20 in September.

The same rewrite lists “adding or reworking sections” and “fixing errors”. These three prompts match those words and still failed:

prompt rate
“Fix the outdated version number in the sdx reference” 5/20
“Add a paragraph about exit codes to the testing claude skills post” 6/20
“Tweak the wording in the vibe coding essay’s opening paragraph” 6/20

All 43 misses were silence. A diagnostic run shows what happens instead: the agent lists the posts directory and opens the file. The request names a file and a small change, so the agent does the change. The description says why the skill matters, a revision entry on every edit, but nothing in the prompt tells the agent that a discipline applies. A description can claim a request. It cannot make an easy task look hard.

The description that asks nothing of the agent never missed

The reviewing description is the only one written entirely as trigger conditions. It starts “Use when”, names three situations, and ends with a phrase in quotation marks that a user would type: “how is the site doing”. Its six prompts fired 120 of 120.

I cannot separate the description from the prompts here. All six prompts are questions, such as “Has the sdx post picked up impressions since it went live?”, and a question has no file to open and no obvious command. The description may be good, or the task may only be hard to do without it. The suite as written cannot tell these apart, and I have not run the test that would: the same six prompts against a description written in the “what it does” style.

The quoted phrase had one measurable cost. The decoy “How many visits are we getting from AI crawlers?” is a data question that the skill does not answer, and it pulled the reviewing skill in 4 of 25 runs. It is the only decoy that failed on any model. A broad quoted phrase claims its neighbours too.

Qualifiers in the description kept 296 of 300 decoys clean

Boundary words did their job. The revising description says “already-published” and the decoy “Edit the draft body of logging-ai-crawlers” stayed clean 25 of 25. The publishing description says “post” and “draft”, and “Publish a new version of ag-grid-url-sync to npm” stayed clean 25 of 25. “Fix the typo in the sdx post title” never fired the scaffolding skill and went to the revising skill in 14 of 25 runs, which is the correct route.

Across five models the decoys passed 296 of 300. The words that say what a skill is not for were the most reliable part of every description.

On the smallest model no phrasing helped

The smallest model fired the right skill in 6 of 120 should-fire runs, and 0 of 30 on the reviewing skill that was perfect everywhere else. It kept 60 of 60 decoys clean. Its median case took 3.7 seconds, against 12.5 or more on the other models: it reads the prompt and acts. No description in this suite changed that, so I do not count it as evidence for or against any phrasing.

What I do now when I write a description

Four rules, each tied to a number above:

  1. Name the state the skill changes in the words a user types. “Draft to published” fired 40 of 40; “ship it live” fired 13 of 20.
  2. Write trigger situations, not a summary of the steps. The one description that is only situations fired 120 of 120, with the confound stated above.
  3. Add qualifiers for what the skill is not for. They cost a few words and held 296 of 300.
  4. Do not expect a description to win against an easy task. Three prompts that the description names explicitly fired 17 of 60.

Only the first half of the revising result is a controlled comparison: one description, rewritten once, on one prompt. Everything else is a rate observed across 36 prompts in a session that had 126 skills loaded, and the same suite moved 11 case-runs on an unchanged model in six weeks. The next test is a second rewrite of the revising description, measured against the 65 of 120 here with a same-day control.

Revisions

  1. Created; body written from the per-prompt rates in the six committed trigger-suite results files.
  2. Published on its slot; auto-drafted by the publishing run because no draft existed for the row.