Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

acumen

Build, benchmark, and optimize agentic skills for your Python package.

Tests Documentation

Agentic skills, tool instructions writen in plain text, allow agents to use tools more succesfuly and efficient. However most python packages do not ship skills with them because developers have no easy way to build and benchmark skills for their tools. Acumen closes this gap. Point it at a Python package and a few evaluation tasks, and it drafts a skill, benchmarks it and improves it across a train/test split so the gains are generalizable.

Many good tools are unusable by coding agents because their maintainers have no way to write a skill for them — or, having written one, no way to tell whether it helps. acumen closes that loop: point it at a Python package and a few tasks, and it drafts a skill, benchmarks it against a no-skill baseline, and improves it across a train/test split so the gains are real generalization, not memorized answers.

  • acumen draft — write skills/v1 from the package's own source.
  • acumen bench — score a skill against a no-skill baseline, in a scrubbed sandbox where the skill is the only difference between arms.
  • acumen improve — refine the skill from its train results, then benchmark again.
  • acumen report — aggregate every run into one self-contained report.html: success rate per version, train vs. test. Bars are coloured by model, with a grey bar pooling all of them; pass --palette claude-opus-5=#3b7ea1 (repeatable) to recolour any of them.

You decide when to stop. Every version is benchmarked on both splits, and only train results reach the improver — so a widening train/test gap is a visible sign a skill is overfitting rather than genuinely helping.

Quickstart

# 1. Scaffold a starter config.yaml and tasks.yaml
acumen init

# 2. Fill in config.yaml (repo). Write tasks.yaml by hand, or generate it:
acumen tasks                     # mine the package for real analyses -> tasks.yaml

# 3. Then run the loop:
acumen bench --no-skill          # the baseline arm
acumen draft                     # generate skills/v1 from the package source, or write by hand
acumen bench --skill v1          # benchmark the skill against the baseline
acumen improve                   # generate skills/v2 from v1's train results, or write by hand
acumen bench --skill v2
acumen report                    # aggregate every run into report.html

# 4. Once a version proves out, ship it into the package itself:
acumen ship --skill v2           # add a <dist>-install-skills console script (PR, or local edit)

acumen ship packages the chosen skill version into the target: the package gains a <dist>-install-skills command that installs the skill into the skills directory of whichever agent the user names — --agent {claude,codex,agents,claude-science}, or an explicit --dest — so the package's own users get the guidance with one command, wherever they run their agent. The same bundle installs verbatim into every framework.

acumen tasks, acumen draft, and acumen improve each accept --feedback "…" to steer the agent with context it can't infer — which functionality to skip when generating tasks, what a skill should emphasise or fix. The guidance is added to the prompt without overriding the train/test isolation, and for draft/improve it is recorded in the version's meta.json and shown in the report. (Don't paste held-out test answers into improve feedback — that would defeat the split.)

draft, improve, tasks, and ship each drive a long autonomous agent. Every run writes a live logs/acumen-<command>-<datetime>.jsonl (one event per step, flushed as it goes — so you can watch progress by reading the file) and a rendered .html transcript. Add --stream to mirror the conversation to the terminal, or --log-dir to change where the logs land.

Getting started

Please refer to the documentation, in particular, the API documentation.

Installation

You need to have Python 3.12 or newer installed on your system. If you don't have Python installed, we recommend installing uv.

And to install the acumen skill that ships with the package into your agent's skills directory, run acumen-install-skills --agent {claude,codex,agents,claude-science} (or --dest <dir> to choose the directory yourself):

acumen-install-skills --agent claude

Release notes

See the changelog.

Contact

For questions and help requests, you can reach out in the scverse discourse. If you found a bug, please use the issue tracker.

Citation

t.b.a

Metadata

Release files for acumen 0.0.1.dev0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for acumen 0.0.1.dev0
File Size Uploaded
acumen-0.0.1.dev0.tar.gz 198.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for acumen 0.0.1.dev0
File Interpreter ABI Platform
acumen-0.0.1.dev0-py3-none-any.whl Python 3 none any Details

Total release size: 353.5 kB

Release files / acumen-0.0.1.dev0.tar.gz

Download URL acumen-0.0.1.dev0.tar.gz
Size 198.7 kB
Tags Source
SHA-256 checksum
How to use checksums
bd0c62aeba3bfc3cea01a36ba946a69c868f0cc3debbedabfc195da922942092
BLAKE2b-256 checksum
How to use checksums
b9eed1ba2f6b4913b45531d37471323d1be7ff22632b4dc155020d2ff9db90b8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 27, 2026.

Transparency log

Release files / acumen-0.0.1.dev0-py3-none-any.whl

Download URL acumen-0.0.1.dev0-py3-none-any.whl
Size 154.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5b70c2fec010ddec300d8f5b7408444e4b58b47306965274a91f1905a4e0f734
BLAKE2b-256 checksum
How to use checksums
86da2cb9881ad52189fac44067678de0b694035215fb04e44e96d5ac402d55d1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 27, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.0.1.dev0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page