skill-creator

skill-creator

Team Pick
AI & Agents
Anthropic
Draft, test, evaluate, rewrite: the loop for turning one of your own repeated processes into a skill that behaves the same way every time.
REPO
anthropics/skills
INSTALL
Claude Code plugin
NEEDS
A process you repeat
Last updated
August 4, 2026

Preview

skill-creator

What it is.

It pulls the process out of a conversation where you already did the work, drafts the skill, writes test prompts, runs them, helps you score the results, then rewrites from whatever failed.

A separate script tunes the description, which decides whether the skill fires on the right request at all. Your brief format and your QA gates are the things no public library will ever hold, and this is how they become repeatable.

What you get.

  • A drafted skill with a name, a trigger description and instructions
  • Test prompts and a run, so you see the skill behaving before you rely on it
  • Quantitative evals alongside your own read of the output
  • A review page for looking at the results side by side
  • A description improver that tunes when the skill fires
HOW TO USE IT

How to set it up.

1

Install the example-skills plugin from the anthropics/skills marketplace.

2

Start from a conversation where you already did the work. It extracts the tools, the sequence and your corrections from the history, which is faster than describing the process cold.

3

Answer the four setup questions: what it should enable, when it should trigger, the output format, and whether the output is objectively checkable.

4

Say yes to test cases when the output is verifiable, such as a file transform or a set sequence, and skip them for subjective work like writing style.

5

Run the test prompts and read the outputs yourself before you look at any score.

6

Rewrite from what the evaluation showed, then widen the test set and run it again.

Pricing Plans

Free. Apache 2.0 license.

Included in Anthropic's example-skills plugin. Checked against the repo in July 2026.

Use cases

Bottling a finished workflow

You just walked Claude through your brief format by hand, again. Start the skill from that conversation and it extracts the sequence and your corrections from the history, then drafts the skill for you to test.

Testing before trusting

The drafted skill handles a file transform, so the output is checkable. Say yes to test cases and run the prompts, then read the outputs yourself before looking at any score.

A skill that never fires

You wrote a skill and requests keep going past it. Run the description improver on it as a standalone step, since the description is what decides whether the skill triggers at all.

Second-round tightening

An existing skill works most of the time. Enter the loop at the evaluation stage and widen the test set from what failed, then rewrite against the misses instead of starting over.

Best for

A process only your team runs

Your brief format and your QA gates are the things no public library will ever contain, and this is how they become repeatable.

A skill that never triggers

The description improver exists for exactly this, and it is a separate script you can run on an existing skill.

Improving something you already wrote

You can enter the loop at the evaluation stage instead of starting from a blank draft.

Read the source

Published by Anthropic. Opens in a new tab.
Open the Tool

Questions about skill-creator

Do I need to know how to code?
Should I write test cases?
Can I improve an existing skill?
What makes a skill trigger reliably?

Questions about AI & Agents

What makes a workflow worth packaging?
How do you test a skill?
Strengths
  • Testing sits inside the loop, and the skill tells you when tests are worth writing.
  • It extracts intent from a conversation where the work already happened.
  • The description improver treats triggering as its own problem, which it is.
  • It adjusts how it explains itself depending on how technical you are, so a marketer gets a usable version.
Limitations
  • Evaluation runs cost tokens, and a wide test set at the end costs more than the draft did.
  • For subjective output the numbers help less, and the skill says so instead of pretending otherwise.
  • A demonstration skill, with Anthropic's caveat attached.
Skip this if
  • Skip it if you already have a loop for drafting and testing your own skills.

The team behind these plays.

We build inbound GTM engines for B2B software teams, and these are the plays we build from. Tell us the pipeline target and we'll show the plan under it.