Skip to main content

skills-eval

skills-eval is a pre-release CLI for Claude Plugin repositories containing one or more Skills. It validates the publishable structure and runs configured security scanners before release.

Install

pipx install skills-eval

Use

# Check every Skill declared by the plugin.
skills-eval check .

# Check one Skill by name or directory.
skills-eval check . --skill wenqu-write

# Show the selected scope and format checks without running security scanners
# or writing a report.
skills-eval check . --dry-run

# Before a full platform release, run every enabled native validator.
# This validates only: it never publishes a Skill or package.
skills-eval check . --external

# On a pull request, run only the native validators that do not need a
# platform login. Repeat --external-target to select them explicitly.
skills-eval check . \
  --external-target claude-plugin \
  --external-target workbuddy \
  --external-target clawhub

A normal run prints a compact summary and writes skills-eval-report.md in the target repository. The report records the selected Skills, each format rule, enabled publishing targets, security scanner configuration, and every finding.

Checks

The portable checks cover the plugin manifest, declared Skills, SKILL.md and frontmatter, local file references, duplicate names or paths, and configured temporary files. The bundled Cisco AI Skill Scanner reviews each selected Skill directory for risky commands, prompt injection, secret exposure, network access, and persistence-related behavior.

Scanner findings are signals for review, not a guarantee that a Skill is safe.

Configuration

Create .skills-eval.json in the plugin root. JSON Schema support is available through the GitHub-hosted schema URL:

{
  "$schema": "https://raw.githubusercontent.com/gogoingai/skills-eval/main/src/skills_eval/schemas/skills-eval.schema.json",
  "schemaVersion": 1,
  "format": {
    "requiredRootFiles": ["README.md"],
    "requiredSkillFrontmatter": ["license"],
    "forbiddenPaths": [".DS_Store"],
    "referenceExtensions": [".md", ".txt"]
  },
  "release": {
    "versionFile": "VERSION",
    "requireVersionSemver": true,
    "changelogFile": "CHANGELOG.md",
    "changelogVersionHeading": "## {version}"
  },
  "publishing": {
    "targets": [
      {
        "name": "claude-plugin",
        "enabled": true,
        "options": { "skillDirectoryPrefix": "skills-" }
      },
      {
        "name": "clawhub",
        "enabled": true,
        "options": { "packageName": "@example/skills" }
      }
    ]
  },
  "report": {
    "language": "auto"
  },
  "security": {
    "sources": [
      {
        "name": "cisco",
        "enabled": true,
        "options": {
          "policy": "balanced",
          "useBehavioral": true
        }
      }
    ]
  }
}

Skills Eval has no project-specific profiles or defaults. Each repository owns its own .skills-eval.json: format and release express repository conventions, while each enabled publishing target contributes only that platform's static checks. release.assetReferences can optionally define an asset directory, a documentation directory, and the reference prefix that must link them. Target options hold project-specific identities where a platform does not supply one itself—for example, claude-plugin.skillDirectoryPrefix and clawhub.packageName. Unsupported or duplicate target names are configuration errors.

Security sources are a configured list so future scanners can be added without changing the command interface.

Native platform validation

The regular check is safe for local development and GitHub Actions. Before a platform release, add --external to run native checks for the enabled publishing targets:

  • claude-plugin: claude plugin validate .
  • workbuddy: codebuddy plugin validate .claude-plugin/marketplace.json
  • skillhub: skillhub publish <selected-skill> --dry-run
  • clawhub: clawhub package validate . --out <temporary directory>

openclaw currently has no separate native CLI validation. Use --external-target <name> repeatedly to select only configured and enabled targets; a selected target is shown in the terminal and Markdown report. Missing tools, login failures, and network errors are reported as an external publishing validation environment failure, never as a Skill security finding. Use CODEBUDDY_BIN, SKILLHUB_BIN, or CLAWHUB_BIN when a CLI is not on PATH. The SkillHub command always includes --dry-run; skills-eval check never invokes a publish command without that platform-provided safety flag.

report.language accepts auto (the default), zh, or en. In auto mode, Skills Eval reads the computer's preferred language: Chinese preferences render the report in Chinese; every other preference renders it in English.

Publishing

skills-eval publish is the explicit opt-in counterpart of check: it pushes the declared Skills and the plugin package to the enabled publishing platforms (currently clawhub and skillhub). Every run verifies the platform login, re-runs the full skills-eval check (including native validations for the selected targets) as a defensive gate, then publishes with rate-limit retries and a final summary.

# Preview the exact publish commands without logging in or publishing.
skills-eval publish . --dry-run

# Publish every Skill and the plugin package to all enabled targets.
skills-eval publish . --changelog "0.2.0 修复画图"

# Publish one Skill to one platform.
skills-eval publish . --target clawhub --skill wenqu-write --changelog "修复画图"

# Publish only the plugin package (clawhub), or only Skills.
skills-eval publish . --plugin-only
skills-eval publish . --skills-only

Credentials come from the environment—CLAWHUB_TOKEN or SKILLHUB_TOKEN—or from a pre-existing clawhub login / skillhub login session; tokens are never printed or embedded in reported commands. Platform identities live in the target options of .skills-eval.json: clawhub.owner and clawhub.packageName are required for ClawHub publishing, clawhub.sourceRepo pins the provenance repository (otherwise derived from the origin remote), and skillhub.host overrides the default API host. Provenance commits and refs come from GITHUB_SHA / GITHUB_REF_NAME when present.

Exit codes: 0 everything published; 1 a publish item or the defensive check gate failed; 2 a prerequisite (CLI missing, not logged in, unknown or disabled target) is unmet. --skip-check bypasses the gate for emergencies.

Release automation

Tagged releases (v*) build the package and publish with PyPI Trusted Publishing through GitHub Actions OIDC. The publish job uses the pypi environment and id-token: write; it does not use a PYPI_TOKEN.

GitHub Action

The repository also provides a reusable GitHub Action. It installs the selected published CLI version, runs the check, and uploads skills-eval-report.md as an artifact even when the check fails. For pull requests, it creates one marked comment and updates that same comment after each later push, so the result is visible without downloading the report.

name: Skills review

on:
  pull_request:
  push:
    branches: [main]

jobs:
  audit:
    runs-on: ubuntu-latest
    permissions:
      contents: read
      actions: read
      pull-requests: write
    steps:
      - uses: actions/checkout@v4
      - uses: gogoingai/skills-eval@v0.2.1
        with:
          path: .

pull_request runs when a PR opens and on every later push to its branch, so maintainers see an up-to-date Skills Eval 审查结果 comment as well as the result in the PR's Checks tab. The comment identifies the checked commit, completion time, and workflow run, and includes a link to download the full report; the same artifact is also available from the run in the repository's Actions page. The comment step uses the automatic GITHUB_TOKEN; no secret needs to be configured. On a PR from an external fork GitHub can deny comment write access; the audit and artifact still complete because commenting is non-fatal. Set comment: false to disable PR comments. The caller controls triggers; this Action never publishes a package, creates a tag, or changes repository files.

For a PR, set external-targets to the validators to run. The Action installs the corresponding CLIs on the runner, then records exactly these commands in the report. claude-plugin, workbuddy, and clawhub are local validations that need no credential:

      - uses: gogoingai/skills-eval@v0.2.1
        with:
          path: .
          external-targets: claude-plugin,workbuddy,clawhub

skillhub is a remote --dry-run and needs a login. Add it to external-targets and pass skillhub-token (a repository secret); the Action installs the SkillHub CLI, logs in with the token, and retries 429 responses automatically. The token is exposed only as an environment variable:

      - uses: gogoingai/skills-eval@v0.2.1
        with:
          path: .
          external-targets: claude-plugin,workbuddy,clawhub,skillhub
          skillhub-token: ${{ secrets.SKILLHUB_TOKEN }}

To run every enabled native validator without per-target selection, use external: true; that mode assumes the required tools and credentials are already available on the runner.

GitHub Action: publish

A separate composite Action at gogoingai/skills-eval/publish runs skills-eval publish—the check Action above keeps its "never publishes" guarantee. A typical tag-triggered release workflow:

name: Publish

on:
  push:
    tags: ["v*"]

jobs:
  publish:
    runs-on: ubuntu-latest
    environment: release
    concurrency: publish
    permissions:
      contents: read
    steps:
      - uses: actions/checkout@v4
      - uses: gogoingai/skills-eval/publish@v0.2.1
        with:
          targets: clawhub,skillhub
          changelog: ${{ github.ref_name }}
          clawhub-token: ${{ secrets.CLAWHUB_TOKEN }}
          skillhub-token: ${{ secrets.SKILLHUB_TOKEN }}

The Action installs the platform CLIs (Node 22 for clawhub, the official installer for skillhub), passes tokens only as environment variables, and uploads the publish log as an artifact. Set dry-run: "true" (for example from a workflow_dispatch input) to preview the planned commands without publishing. Use a GitHub environment with required reviewers on release when a human approval gate is desired; without reviewers the tag push publishes directly.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

skills_eval-0.2.1.tar.gz (62.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

skills_eval-0.2.1-py3-none-any.whl (47.0 kB view details)

Uploaded Python 3

File details

Details for the file skills_eval-0.2.1.tar.gz.

File metadata

  • Download URL: skills_eval-0.2.1.tar.gz
  • Upload date:
  • Size: 62.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for skills_eval-0.2.1.tar.gz
Algorithm Hash digest
SHA256 ef6fde2a9f59373c0d8313ef26d432f99fd2333f22cca4d54ce4d7dd9d6c20a2
MD5 ebc0ead9eeecf943310df6e7c12ae026
BLAKE2b-256 e2c87dd5aa8553c0c37820970f40d52c0fb4849ed657369af8d51e9ab1da6953

See more details on using hashes here.

Provenance

The following attestation bundles were made for skills_eval-0.2.1.tar.gz:

Publisher: release.yml on gogoingai/skills-eval

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file skills_eval-0.2.1-py3-none-any.whl.

File metadata

  • Download URL: skills_eval-0.2.1-py3-none-any.whl
  • Upload date:
  • Size: 47.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for skills_eval-0.2.1-py3-none-any.whl
Algorithm Hash digest
SHA256 0f12692ea1fd45590e2916c2d576913350c1198aa3e3871e92e4207aaa10b8b9
MD5 9c49071c5826ab93de4af5f60c7a01bb
BLAKE2b-256 b84cb9165a6fb60a0e8b1fc8d77c28522697a8b72ef2dec479dfb900e08713a2

See more details on using hashes here.

Provenance

The following attestation bundles were made for skills_eval-0.2.1-py3-none-any.whl:

Publisher: release.yml on gogoingai/skills-eval

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.3.3

2 files

0.3.2

2 files

0.3.1

2 files

0.3.0

2 files

This release

0.2.1 This release

2 files

0.2.0

2 files

0.1.13

2 files

0.1.12

2 files

0.1.11

2 files

0.1.10

2 files

0.1.9

2 files

0.1.8

2 files

0.1.7

2 files

0.1.6

2 files

0.1.5

2 files

0.1.4

2 files

0.1.3

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page