Skill Quality Lab
A release checklist and test runner for Agent Skills.
Skill Quality Lab is a local-first toolkit for auditing, testing, packaging, and installing
SKILL.md-based skills. It combines repeatable checks with a manual review rubric. Static results
and observed runtime behavior are reported separately.
It targets Codex, Claude, and other clients that follow the Agent Skills layout. It can also run structural preflight checks for MCP, OpenAPI, LangChain, and Semantic Kernel projects.
[!IMPORTANT] A score of 100 means that the configured deterministic checks found no errors or warnings. It is not a universal guarantee of security, activation behavior, or cross-platform compatibility.
Resumo em português
O Skill Quality Lab audita skills de agentes com verificações reproduzíveis de estrutura, sintaxe, segurança, dependências, ativação e portabilidade. Ele também cria pacotes ZIP determinísticos, instala skills com proteção contra sobrescrita e oferece adaptadores básicos para MCP, OpenAPI, LangChain e Semantic Kernel. Tudo funciona localmente por padrão; rede, instalação de dependências e chamadas a APIs exigem opções explícitas.
Why this exists
A skill may look correct while still containing a broken script, an unreachable reference, an overbroad activation description, a leaked token, or a destructive default. Typical linters only cover one piece of that problem.
Skill Quality Lab gives maintainers a single release workflow:
- Audit structure, metadata, routed resources, syntax, safety, and portability.
- Review semantic quality with explicit evidence.
- Exercise activation with realistic positive, negative, and boundary prompts.
- Resolve Python dependencies in a disposable environment.
- Compare before/after findings.
- Build a deterministic package and install it safely.
Highlights
| Capability | What it checks |
|---|---|
| Skill audit | Frontmatter, naming, structure, routed resources, portability, safety, and release readiness |
| Runtime validation | Python, JSON, YAML, TOML, Bash, PowerShell, JavaScript, and TypeScript |
| Security review | Credential patterns, destructive commands, risky Python APIs, redacted evidence, and reviewed suppressions |
| External scanners | Optional Gitleaks and TruffleHog integration when already installed |
| Dependency isolation | Temporary virtualenv, requirements installation, pip check, and smoke imports |
| Activation testing | Validated prompt suites, local command harnesses, and opt-in provider classifiers |
| Ecosystem preflight | MCP configuration, OpenAPI 3.x, LangChain, and Semantic Kernel |
| Comparison | Resolved, introduced, and persistent findings between two audits or directories |
| Distribution | Deterministic ZIP archives, SHA-256 checksums, dry-run installation, backup, and rollback |
Requirements
- Python 3.11 or newer
piporpipx
Optional checks use tools already available on PATH:
| Resource | Checker |
|---|---|
| Shell | bash -n |
| JavaScript | node --check |
| TypeScript | tsc --noEmit |
| PowerShell | PowerShell parser API |
| Extended secret scanning | Gitleaks or TruffleHog |
Missing optional tools are reported as not_assessed; they are never counted as successful
checks.
Installation
After the first release is published, install the command-line tool from PyPI:
python -m pip install skill-quality-lab
For an isolated global command, use pipx:
pipx install skill-quality-lab
With pipx, run audits through the global skill-quality-lab command so they use the isolated
environment that includes PyYAML. Direct execution of a bundled scripts/*.py file requires
PyYAML 6.x in that script's Python interpreter; install scripts/requirements.txt when needed.
Then install the bundled skill for Codex or Claude:
skill-quality-lab install --client codex
skill-quality-lab install --client claude
Preview the target without changing it:
skill-quality-lab install --client codex --dry-run
Codex uses $CODEX_HOME/skills or ~/.codex/skills. Claude uses
$CLAUDE_CONFIG_DIR/skills or ~/.claude/skills. Override either destination explicitly when
needed:
skill-quality-lab install --client codex --destination /path/to/skills
Existing installations are never overwritten unless --replace is supplied. Replaced and
uninstalled copies are moved to a backup outside the watched skills directory.
Check the package and client installations:
skill-quality-lab doctor
Quick start
Audit a skill from any directory:
skill-quality-lab audit /path/to/my-skill --profile portable
Use strict mode for a release gate and JSON for CI or other automation:
skill-quality-lab audit /path/to/my-skill \
--profile codex \
--strict \
--format json \
--output reports/my-skill.audit.json
Available profiles:
portable— client-neutralSKILL.mdchecks.codex— portable checks plus Codex-oriented metadata expectations.claude— portable checks plus Claude-oriented compatibility checks.
Every report includes a verdict, score formula, findings with evidence and remediation, runtime coverage, security capabilities, and known limits.
Core workflows
Compare an improvement
Capture reports before and after a change:
skill-quality-lab audit /path/to/my-skill \
--format json --output reports/before.audit.json
skill-quality-lab audit /path/to/my-skill \
--format json --output reports/after.audit.json
skill-quality-lab compare \
reports/before.audit.json reports/after.audit.json
The comparison separates resolved, introduced, and persistent findings instead of treating a score change as proof of improvement.
Run a security scan
The built-in scan is local and read-only:
skill-quality-lab security /path/to/my-skill
Use an installed external scanner explicitly:
skill-quality-lab security /path/to/my-skill --external available
Candidate secret values are redacted from reports. Inline suppressions use the following form and remain visible as review notes:
shutil.rmtree(staging) # skill-quality: allow destructive-api-call -- validated staging child
Only suppress a finding after verifying the resolved target, safeguards, and recovery path.
Check dependencies in isolation
Plan mode discovers requirements without changing the environment:
skill-quality-lab dependencies /path/to/my-skill
After approving network access and package build-code execution, create a disposable environment:
skill-quality-lab dependencies /path/to/my-skill \
--create-venv \
--import yaml
The environment is removed after installation, pip check, and the requested smoke imports.
Test activation
Create a suite with at least three direct positives, two indirect positives, three negatives, and two boundary cases. Validate it before execution:
skill-quality-lab validate-activation activation-suite.json
Run it through a local command harness:
skill-quality-lab activation activation-suite.json \
--skill-directory /path/to/my-skill \
--output activation-results.json \
--runner command \
--command python my_client_harness.py
The harness receives one JSON object per case on standard input and returns:
{
"activation": true,
"evidence": "The client discovered and loaded the skill."
}
Provider classifiers are also available for openai, anthropic, and gemini. They require an
explicit model, --allow-network, and the corresponding environment credential. Use --limit
for cost-bounded trials.
skill-quality-lab activation activation-suite.json \
--skill-directory /path/to/my-skill \
--output classified-results.json \
--runner openai \
--model YOUR_MODEL_ID \
--limit 3 \
--allow-network
Provider output is classification evidence. It does not prove that an actual client discovered or loaded the skill. Conditional boundary cases always require human adjudication.
Audit adjacent ecosystems
skill-quality-lab ecosystem /path/to/artifact --adapter openapi
skill-quality-lab ecosystem /path/to/artifact --adapter mcp
skill-quality-lab ecosystem /path/to/project --adapter langchain
skill-quality-lab ecosystem /path/to/project --adapter semantic-kernel
These adapters are structural preflight checks. They do not connect to servers, invoke endpoints, restore every dependency ecosystem, or prove production behavior.
Package and install
Create a deterministic archive only after a clean audit:
skill-quality-lab package /path/to/my-skill \
--output dist/my-skill.zip \
--checksum
Preview installation into an explicit skills directory:
python scripts/install_skill.py dist/my-skill.zip \
--destination /path/to/skills \
--dry-run
Remove --dry-run after reviewing the destination. Replacing an existing skill requires
--replace; the installer creates a backup and restores it if installation fails.
Scoring and verdicts
The structural score uses this formula:
score = max(0, 100 - 20 × errors - 7 × warnings)
| Verdict | Meaning |
|---|---|
ready |
No deterministic errors or warnings were found |
ready with warnings |
No blocking error, but material risks remain |
not ready |
One or more blocking errors were found |
Runtime pass percentage is reported separately. Semantic quality is assessed with the rubric in
references/rubric.md, using not assessed whenever evidence is missing.
Configuration
Add an optional .skill-quality.json to the skill being audited:
{
"max_skill_lines": 450,
"max_description_chars": 900,
"require_openai_yaml": true,
"scan_secrets": true,
"scan_destructive_commands": true
}
Release-gate weakening should always be an explicit maintainer decision. See
references/configuration.md for the supported schema and review
rules.
Project structure
skill-quality-lab/
├── .github/workflows/ # Cross-platform CI and trusted releases
├── agents/openai.yaml # Codex UI metadata
├── references/ # Detailed operational guidance
├── scripts/
│ ├── skill_quality_lab/ # Canonical Python package
│ └── *.py # Standalone compatibility commands
├── tests/test_skill_quality.py # Unit and integration tests
├── pyproject.toml # PyPI metadata and build configuration
├── README.md # Community documentation
├── SKILL.md # Skill workflow and activation contract
└── LICENSE # MIT
README.md and tests are repository assets; the deterministic skill packager intentionally keeps
them out of the runtime ZIP.
Development
Run the complete test suite:
python -m pip install -e .
python -m unittest discover -s tests -v
Build and validate the PyPI distributions:
python -m pip install build twine
python -m build
python -m twine check dist/*
Validate the skill metadata with the official skill-creator validator when available:
python /path/to/skill-creator/scripts/quick_validate.py .
Audit the project against every supported profile:
skill-quality-lab audit . --profile portable --strict
skill-quality-lab audit . --profile codex --strict
skill-quality-lab audit . --profile claude --strict
Releases are tag-driven. The package version is derived directly from an annotated vX.Y.Z Git
tag, eliminating a separate source version to update manually. GitHub Actions tests the tag on
Python 3.11–3.14 across Linux, Windows, and macOS, publishes to TestPyPI, and then
publishes to PyPI through Trusted Publishing. No long-lived PyPI token is stored in the repository.
Maintainer release setup
Before publishing, enable two-factor authentication on PyPI and TestPyPI and store the recovery
codes securely. Create the GitHub environments testpypi and pypi, and require manual approval
for pypi. Protect main with required CI checks and add a GitHub ruleset that restricts creation,
updates, and deletion of v* tags.
Register a Pending GitHub Publisher on both package indexes with these exact values:
| Field | PyPI | TestPyPI |
|---|---|---|
| Project | skill-quality-lab |
skill-quality-lab |
| Owner | arthur-paraibano |
arthur-paraibano |
| Repository | skill-quality-lab |
skill-quality-lab |
| Workflow | release.yml |
release.yml |
| Environment | pypi |
testpypi |
PyPI and TestPyPI use separate accounts and publisher settings. A pending publisher does not
reserve the project name, so publish the first release promptly after configuration. Do not add a
PYPI_TOKEN secret. After both publishers are configured, start from a clean, synchronized main
branch and create the next unused version tag only after every release change is committed:
git pull --ff-only
git status --short
git tag -a v0.2.0 -m "Release v0.2.0"
git push origin v0.2.0
The empty git status --short output is required. Never create a release tag before its changes
are committed and pushed to main. The release workflow independently verifies the versions
embedded in both the wheel and source distribution before either artifact reaches a package index,
and rejects tags whose commit is not contained in main. TestPyPI uploads are safe to rerun when
that index already contains one or both artifacts for the release version.
Before opening a contribution:
- Add a focused regression test for behavioral changes.
- Preserve local-first and read-only defaults.
- Keep network access, package execution, and destructive actions explicitly opt-in.
- Update the relevant file under
references/without bloatingSKILL.md. - Report what was actually executed and what remains unassessed.
Limitations
- Static and heuristic checks cannot prove that an artifact is safe.
- API classifiers do not prove real client discovery or instruction loading.
- Cross-platform claims require a CI matrix across the advertised systems and runtime versions.
- Optional checkers only contribute evidence when installed.
- Ecosystem adapters validate structure; they are not full protocol or production integration tests.
- Obfuscated secrets, dynamically assembled commands, and semantic business risks still require human review.
Contributing
Issues and pull requests are welcome. Small, evidence-backed changes with regression tests are preferred. If you add a checker, document its failure mode, unavailable-tool behavior, security boundary, and what a passing result does not prove.
License
Released under the MIT License.
Metadata
Release files for skill-quality-lab 0.1.4
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| skill_quality_lab-0.1.4.tar.gz | 54.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| skill_quality_lab-0.1.4-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 119.3 kB
Release files / skill_quality_lab-0.1.4.tar.gz
| Download URL | skill_quality_lab-0.1.4.tar.gz |
|---|---|
| Size | 54.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4187d26a421b25fd2bc3d73689dca9adc056c911a166b8f66a4dc6653ef4fd36
|
|
BLAKE2b-256 checksum How to use checksums |
ab45cce3501a5f91fd8d23251bd90ccb964080d22b657226e4ed4dca3aa7adcd
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 12, 2026.
Transparency logRelease files / skill_quality_lab-0.1.4-py3-none-any.whl
| Download URL | skill_quality_lab-0.1.4-py3-none-any.whl |
|---|---|
| Size | 65.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
f54e8cb4646221cffd83220d75e68dbec9e8216256a16c46137684653aed214c
|
|
BLAKE2b-256 checksum How to use checksums |
cad26704fccc077e0967790061194cd1cf5fa86e6ea68af4ed18c1a4e8b90691
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 12, 2026.
Transparency log