Skip to main content

title: agentskills-tools description: Command line authoring, inspection, evaluation, and MCP publication diagnostics for Agent Skills.

Command line tools for authoring and validating Agent Skills.

Part of the Agent Skills SDK.

Install

pip install agentskills-tools

The serve command needs the MCP server, which is an optional extra so that validating skills in CI does not pull in mcp and pydantic:

pip install "agentskills-tools[serve]"

Commands

Every command takes either one skill folder or a folder of skill folders — the one containing SKILL.md, or the one containing directories that do.

Command What it does
agentskills init <name> Scaffold a skill that already validates.
agentskills validate <path> Check skills against the specification. Exits 1 on any error.
agentskills lint <path> Report what is legal but still costly.
agentskills inspect <path> Show what an agent would actually receive, and what it costs.
agentskills eval <path> Measure what difference a skill makes.
agentskills serve <path> Run an MCP server over a folder of skills.

init

agentskills init incident-response --path ./skills

Creates skills/incident-response/ with a valid SKILL.md and empty references/, scripts/, and assets/ directories. The template is validated before anything is written, so an unusable name is refused rather than scaffolded.

validate

agentskills validate ./skills
skills/incident-response
  ok
skills/broken-skill
  error   frontmatter-invalid-yaml (line 3): frontmatter is not valid YAML: mapping values are not allowed here

2 skills checked, 1 error, 0 warnings

Frontmatter is parsed by the CLI before the skill reaches the SDK's validator. The SDK's parser is deliberately forgiving — malformed YAML yields an empty mapping — which downstream reads as "no name, no description" and tells you nothing about the colon you missed.

lint

agentskills lint ./skills --strict --max-body-tokens 4000
Code Warning
missing-version No version, so consumers cannot pin the skill or detect drift.
description-too-long-for-catalog Catalog entries sit in context every turn.
missing-selection-metadata A long description with no when_to_use or when_not_to_use is usually smuggling conditions into prose.
body-over-token-budget Body is large enough that detail belongs in references/.
unreferenced-resource A file the body never mentions is a file no agent will load.

Warnings do not fail the command unless --strict is passed.

inspect

agentskills inspect ./skills/incident-response

Prints the metadata, the resource list, the catalog entry the agent sees on every turn, and the body it loads on demand — each with an estimated token cost, so you can see the price before shipping.

Native manifest inspection

agentskills inspect ./skills --native --format json
agentskills inspect ./skills --native --max-file-bytes 1048576

Native inspection reads every file into a bounded snapshot and reports canonical URIs, original-byte sizes and SHA-256 digests, unchanged frontmatter, and protocol requirements. It includes binary, hidden, and nonstandard supporting files. It does not require the MCP extra and does not pass canonical content through the legacy body parser. --native and --cost are mutually exclusive.

Each skill is limited to 512 files and 16 MiB. The additional per-file bound defaults to 16 MiB and can be lowered with --max-file-bytes. Invalid manifests, unsafe paths, content drift, and exceeded limits fail before any report is written. Like existing inspection failures, they exit with code 2.

The report scope is localSkillSnapshot. This is an offline publication check, not metadata-only client discovery, a running-server probe, or a check of an entire server's publication namespace. A valid snapshot does not grant permission to activate or execute a skill. Use serve --native --check for catalog-wide publication validation. These options are available starting in v0.6.0.

Token cost

agentskills inspect ./skills --cost
skills/incident-response  (incident-response)
  counted with tiktoken/cl100k_base
  catalog entry                         66  every turn
  body                                 439  on load
    Incident Response                   14
      When to Declare an Incident       50
      Roles                             70
      General Triage Steps             121
  references/escalation-policy.md      448  on demand
  assets/escalation-flowchart.mermaid  235  on demand
  per turn 66, per load 505, all resources 2,197

1 skill, 66 tokens charged every turn

The right-hand column is the point. A catalog entry is injected on every turn whether or not the skill is ever used; a body is charged once per load; a reference is charged only if the agent goes and reads it. Authors reliably get this backwards, trimming a body while ignoring a description that costs a hundred tokens a turn forever.

Sections do not nest — a heading owns its own text up to the next heading of any level — so the parts sum to the body exactly. Depth shows in the indent instead. A # inside a fenced code block is a shell comment, not a heading.

The splitter itself lives in agentskills-core (split_sections), which also serves the runtime get_skill_outline / get_skill_section tools. agentskills_tools.cost re-exports Section, split_sections and PREAMBLE_TITLE so this report and what an agent sees at runtime cannot disagree about where a section begins.

A resource that is not UTF-8 text reports its size in bytes and no token count, because an image has a size but not a token cost.

Flag Effect
--budget N Exit 1 when catalog entry plus body exceeds N tokens.
--turn-budget N Exit 1 when the catalog entry alone exceeds N tokens.
--tokenizer auto (default), tiktoken, or heuristic.

Two budgets rather than one, for the same reason: a single threshold is dominated by the body, so the per-turn cost stays invisible to exactly the gate meant to catch it.

Counting is exact when tiktoken is installed and a four-characters-per-token estimate otherwise. It is not a dependency here: it ships a compiled wheel and fetches its vocabulary over the network on first use, which is a poor trade for a tool whose main job is reading YAML in CI. Install it yourself if you want exact numbers.

Whichever counter ran is named in every report, and --tokenizer tiktoken refuses to fall back — a budget gate that quietly changes its arithmetic depending on what happens to be installed is worse than no gate. Pin it in CI and leave auto for the terminal.

lint --max-body-tokens keeps the estimate regardless, so its verdict never depends on the machine it ran on.

eval

A skill is a prompt, and nobody measures whether a given prompt makes an agent better. Authors ship on intuition, reviewers approve on prose quality, and editing a body can degrade task success with no signal anywhere.

Write cases beside the skill, in evals/ inside the skill folder:

# skills/incident-response/evals/triage.yaml
skill: incident-response      # optional; checked against the folder
judge_model: gpt-4o           # required if any case uses `judge`
cases:
  - name: declares-and-triages
    prompt: Checkout is returning 500s for a third of users.
    repeat: 3                 # models are not deterministic
    threshold: 0.67           # fraction of repeats that must pass
    expect:
      - contains: "Incident Commander"
      - not_contains: "I don't have access"
      - regex: "(?i)severity"
      - judge: "Tells the responder to assess severity before attempting a fix"

repeat defaults to 1 and threshold to 1.0. Every expectation must hold for a repeat to pass.

Eval files are checked by agentskills validate, with no model and no API key, so a broken case fails in CI beside the skill rather than the first time somebody pays to run it.

agentskills eval ./skills --model mypkg.evals:openai_client
incident-response (triage.yaml)
  pass declares-and-triages: with 100%, without 33%, delta +67%
  FAIL postmortem-window: with 0%, without 0%, delta +0%
         unmet contains: 48 hours
  suite delta +33% on gpt-4o

2 cases run, 1 failed, mean delta +33%

Every case runs twice: once with the skill's body in the system prompt, once without. Absolute pass rates mostly measure the underlying model, so the number that means anything is the difference. A skill whose cases pass equally well without it is not earning its tokens.

Bringing your own model

--model takes module:factory — a dotted path to a zero-argument callable returning a client. Nothing in this project depends on a provider SDK, and a ten-line adapter is a smaller ask than an opinion about which vendor you should install:

# mypkg/evals.py
from openai import AsyncOpenAI
from agentskills_tools.evals import ModelResponse


class OpenAIModel:
    model_id = "gpt-4o"

    def __init__(self) -> None:
        self._client = AsyncOpenAI()

    async def complete(self, *, system: str, prompt: str) -> ModelResponse:
        reply = await self._client.chat.completions.create(
            model=self.model_id,
            temperature=0,
            messages=[
                {"role": "system", "content": system},
                {"role": "user", "content": prompt},
            ],
        )
        return ModelResponse(reply.choices[0].message.content or "")


openai_client = OpenAIModel

model_id is part of every report and of the cache key, because a pass rate without the model that produced it is not a measurement. Set temperature to zero if your provider allows it; this side has no opinion it could enforce.

--judge names a second client for judge expectations and defaults to the model under test — the cheapest judge and the least independent one. When repeat is above 1, the report flags cases whose repeats disagreed, because a case that passes three times in five has measured sampling noise rather than a skill.

Cost

These calls hit real APIs and cost real money. eval is never part of pytest: it runs only when you invoke it, with credentials you supply. Completions are cached under .agentskills/eval-cache by model, system prompt, user prompt, and repeat index — so editing a skill re-buys its runs, while tightening an expectation re-grades the answers already bought. --no-cache turns that off; --cache-dir moves it.

serve

agentskills serve ./skills --transport stdio

Runs the MCP server over a folder of skills without hand-writing a config file. For anything beyond a single filesystem root — HTTP providers, per-skill options, environment placeholders — use agentskills-mcp-server with a server.json.

Native Skills publication and preflight are available in the v0.6 checkout:

agentskills serve ./skills --native --check
agentskills serve ./skills --native --transport stdio
agentskills serve ./skills --native --transport streamable-http

Native serving requires the server extra and mcp>=2.2,<3. It keeps the full canonical resources and does not register legacy tools. Without --native, the existing compatibility mode remains the default. Clients without the Skills extension should use that legacy mode.

--check builds the actual server and exits without starting a listener. Native preflight validates all captures together, including publication conflicts and aggregate limits. Defaults are 128 skills and 64 MiB of captured bytes, in addition to the per-skill limits above. Use --max-file-bytes to apply the same per-file bound as an earlier inspection. Use the config-driven server for other catalog limits, aliases, or HTTP providers. Legacy mode also supports --check.

The local HTTP endpoint is http://127.0.0.1:8000/mcp. Do not expose that listener directly to an untrusted network. Preflight does not exercise the chosen transport, authentication, or host behavior. See the MCP deployment boundaries before publishing a remote endpoint.

Exit codes

Code Meaning
0 Ran, found nothing wrong.
1 Ran, found errors — or warnings under --strict, or a cost over budget.
2 Could not run: bad path, missing extra, unwritable directory.

The distinction matters in CI: 1 means a skill is broken, 2 means the invocation is.

JSON output

validate, lint, inspect, and eval accept --format json. The schema is a published contract; schemaVersion is bumped only for a breaking change, and new fields are added rather than existing ones repurposed.

{
  "schemaVersion": 1,
  "command": "validate",
  "ok": false,
  "summary": { "skills": 2, "errors": 1, "warnings": 0 },
  "skills": [
    {
      "id": "broken-skill",
      "path": "skills/broken-skill",
      "ok": false,
      "findings": [
        {
          "severity": "error",
          "code": "frontmatter-invalid-yaml",
          "message": "frontmatter is not valid YAML: mapping values are not allowed here",
          "line": 3,
          "file": "skills/broken-skill/SKILL.md"
        }
      ]
    }
  ]
}

ok mirrors the exit code, so a consumer never has to re-derive the strictness rules. line is null unless the problem can be attributed to one line. file is the skill's SKILL.md unless the finding is about another file in the folder, such as an eval case file.

inspect --cost --format json reports each skill's perTurn, perLoad and onDemand totals, the sections and resources they were summed from, the overBudget messages, and the counter that produced the numbers — including whether it was exact. A consumer that charts these over time needs to know when the unit changed underneath it.

Continuous integration

The published action wraps validate and lint and annotates every finding on the pull request diff:

- uses: pratikxpanda/agentskills-sdk/actions/validate@v1
  with:
    path: ./skills
    fail-on-lint: false

To run it yourself, validate and lint also accept --format github, which emits workflow commands instead of a report:

agentskills validate ./skills --format github
::error file=skills/deploy/SKILL.md,line=3,title=frontmatter-invalid-yaml::frontmatter is not valid YAML

Anywhere else, the exit code is enough:

- run: pip install agentskills-tools
- run: agentskills validate ./skills

Logging

Pass -v to send the SDK's debug logs to stderr, leaving stdout parseable:

agentskills validate ./skills --format json -v > report.json

Security

Agent Skills are equivalent to executable code — skill content is injected into an LLM agent's context verbatim. Validating a skill does not make it safe to run. Only load skills from sources you trust.

See SECURITY.md.

License

MIT

Metadata

Release files for agentskills-tools 0.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for agentskills-tools 0.6.0
File Size Uploaded
agentskills_tools-0.6.0.tar.gz 37.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for agentskills-tools 0.6.0
File Interpreter ABI Platform
agentskills_tools-0.6.0-py3-none-any.whl Python 3 none any Details

Total release size: 79.1 kB

Release files / agentskills_tools-0.6.0.tar.gz

Download URL agentskills_tools-0.6.0.tar.gz
Size 37.9 kB
Tags Source
SHA-256 checksum
How to use checksums
a2b0bfd140e0212f1fb1252c00707157df9b115ae93783584c5f0525891cb3c9
BLAKE2b-256 checksum
How to use checksums
6433809800b9ca7b062c5bcf4c207a08116d8f9f07bfebf9c456e5e95230e72a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release files / agentskills_tools-0.6.0-py3-none-any.whl

Download URL agentskills_tools-0.6.0-py3-none-any.whl
Size 41.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a2341205c372c3db67b70bf84efa839b5a86cecbf71cd9404fe91455d3e76b34
BLAKE2b-256 checksum
How to use checksums
d41f91c4470008d1936862db345bdc25e0a7697cab8acc2df72b57db7abb1c72
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release history Release notifications | RSS feed

0.7.0

2 release files

This release

0.6.0 This release

2 release files

0.5.0

2 release files

0.4.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page