PromptPilot - SLM-powered control plane for AI coding agents; routes, clarifies, compresses, and preserves context before invoking Codex/Claude-style agents
Project description
PromptPilot
Small-language-model control layer for AI coding agents.
PromptPilot helps Codex and Claude-style coding agents spend less context on avoidable work.
Long coding sessions often burn frontier-model tokens on ambiguous prompts, repeated session history, noisy tool output, and constraints that should have been made explicit before coding starts. PromptPilot adds a small language model (SLM) control layer in front of those agents to make that work clearer before the expensive model starts coding.
It clarifies unclear requests, rewrites prompts to preserve constraints, routes simple or risky requests appropriately, carries bounded session memory, and can compress noisy tool output through agent hooks.
The goal is not to replace the frontier model with a small language model. The SLM manages the workflow; the frontier model writes and debugs the code.
PromptPilot optimizes for semantic-preserving context control, not blind token reduction. A rewritten prompt may be longer than the original when that preserves constraints; the savings come from fewer ambiguous agent turns, bounded session replay, and compressed noisy context.
How PromptPilot fits
%%{init: {"flowchart": {"curve": "basis", "nodeSpacing": 48, "rankSpacing": 60}}}%%
flowchart LR
U([Developer request])
subgraph PP["PromptPilot control plane"]
direction LR
M[["Session memory<br/>bounded summaries"]]
C{{"SLM route<br/>clarify / answer / passthrough / act"}}
Q["Clarify<br/>ask first"]
A["Answer<br/>offer reply"]
D["Direct reply<br/>opt-in only"]
P["Passthrough<br/>raw prompt"]
R["Act<br/>safe rewrite"]
end
subgraph AG["Frontier coding agent"]
direction LR
F["Codex / Claude CLI"]
O["Code changes<br/>tests / summary"]
T["Tool output"]
end
subgraph HK["Optional hooks"]
H["Compress logs<br/>pytest / grep / diff"]
end
U --> M --> C
C -->|clarify| Q
C -->|answer| A
A -->|enabled| D
A -.->|otherwise| F
C -->|passthrough| P --> F
C -->|act| R --> F
F --> O
F --> T --> H --> F
C -. "hybrid" .-> API[("Metered SLM API")]
F -. "hybrid" .-> SUB[("Subscription CLI")]
classDef entry fill:#fff7ed,stroke:#fb923c,stroke-width:2px,color:#7c2d12;
classDef control fill:#eef2ff,stroke:#6366f1,stroke-width:2px,color:#312e81;
classDef route fill:#f5f3ff,stroke:#8b5cf6,stroke-width:2px,color:#4c1d95;
classDef agent fill:#ecfeff,stroke:#06b6d4,stroke-width:2px,color:#164e63;
classDef hook fill:#f0fdf4,stroke:#22c55e,stroke-width:2px,color:#14532d;
classDef infra fill:#f8fafc,stroke:#94a3b8,stroke-width:1.5px,color:#334155;
class U entry;
class M,Q,A,D,P,R control;
class C route;
class F,O,T agent;
class H hook;
class API,SUB infra;
For answer, PromptPilot skips the downstream coding agent only when direct SLM answering is enabled with --let-slm-answer or PROMPTPILOT_LET_SLM_ANSWER; otherwise the request continues to the agent. The diagram keeps node labels intentionally short so GitHub Mermaid previews do not clip long text.
Measured example (hybrid mode): in one 15-turn chain, ~24k input tokens of SLM work directed ~12.66M input tokens of agent work. The control layer was ~0.2% of the input-token footprint, and the bounded session ran the same multi-turn work on ~7.6x fewer input tokens than the tool's native --resume session. Hybrid mode can route the small control layer to metered API usage and the heavy coding-agent work to a subscription CLI. See docs/HYBRID_MODE.md and docs/BENCHMARKS.md. Single workload, not a guarantee.
First-time user? Start with QUICKSTART.md.
Full documentation: read the rendered PromptPilot GitHub Wiki.
Docs
Long-form documentation lives in docs/ (source of truth) and is mirrored to the PromptPilot GitHub Wiki by scripts/publish_wiki.sh. The wiki is the easiest place to browse the project.
- Start at the GitHub Wiki home, the docs index, or the Project Overview.
- Operational pages stay at the repo root: this README, QUICKSTART.md, SECURITY.md, CONTRIBUTING.md.
Install
PromptPilot wraps an existing coding agent CLI, so install and authenticate at least one agent first:
- Claude Code:
npm install -g @anthropic-ai/claude-code, thenclaude auth login --claudeai - Codex:
npm install -g @openai/codex, thencodex login
Then install PromptPilot. Use the extra that matches the small-model API path you want available:
pip install prpt[claude] # Claude/Anthropic SLM path
pip install prpt[codex] # Codex/OpenAI SLM path
pip install prpt[all] # both
Subscription CLI auth and API keys both work; hybrid mode can use an API key for the small control layer and a subscription CLI for the coding agent. [anthropic] / [openai] are kept as aliases for backward compatibility.
First run
cd /path/to/your/repo
prpt setup # one-time onboarding (checks + smoke test)
prpt "fix the flaky test in payments" # auto-detects claude or codex from PATH
prpt --dry-run "refactor auth, no API changes" # preview the optimized prompt
prpt --tool codex "add dark mode" # force a specific agent
prpt doctor # re-run setup checks if something breaks
prpt install-hook # optional: wire prompt/tool hooks into Claude Code
After many turns the session grows heavy:
prpt restart # checkpoint -> handoff.md -> bootstrap fresh
What a run looks like
$ prpt "the test in tests/test_auth.py::test_token_refresh is flaky on CI
but passes locally. keep the public API of TokenStore intact."
[promptpilot] session: carrying 0 prior turns
[promptpilot] route=act
[token stats] raw 248 → optimized 332 tokens (SLM call: $0.0021)
=== forwarding to claude-code ===
... agent works ...
✓ tests/test_auth.py::test_token_refresh now stable (3/3 CI retries)
The SLM expanded the raw 248-token prompt into a 332-token optimized version that pinned the failing test name and made the TokenStore API-stability constraint explicit before the coding agent saw it. Walkthrough in docs/TELEMETRY_AND_REPLAY.md.
For the full guide see QUICKSTART.md and prpt --help
(or prpt --advanced-help for internal/researcher flags).
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file prpt-0.2.0.tar.gz.
File metadata
- Download URL: prpt-0.2.0.tar.gz
- Upload date:
- Size: 128.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f87446610d812d619486e67c3b4f2040510356802a546fcc17b122c9352ebe8a
|
|
| MD5 |
0e96bad1a98d3e61a3af1d769aa54dc0
|
|
| BLAKE2b-256 |
f2624226c7590a3d002717bd206e2d454f25cc79f3a1d4dbe26275ef82f357d6
|
Provenance
The following attestation bundles were made for prpt-0.2.0.tar.gz:
Publisher:
publish.yml on steyangdot/PromptPilot
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
prpt-0.2.0.tar.gz -
Subject digest:
f87446610d812d619486e67c3b4f2040510356802a546fcc17b122c9352ebe8a - Sigstore transparency entry: 1701724545
- Sigstore integration time:
-
Permalink:
steyangdot/PromptPilot@c90aa554081bf472fb6aad5135a390fff19f2181 -
Branch / Tag:
refs/tags/v0.2.0 - Owner: https://github.com/steyangdot
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@c90aa554081bf472fb6aad5135a390fff19f2181 -
Trigger Event:
release
-
Statement type:
File details
Details for the file prpt-0.2.0-py3-none-any.whl.
File metadata
- Download URL: prpt-0.2.0-py3-none-any.whl
- Upload date:
- Size: 111.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
821e3043623d0d7b057c48a8db1029b2b0cf78c964f1d8b11740a10fe3fd2278
|
|
| MD5 |
5f775fc15bbe582743dd9c75886c5dd6
|
|
| BLAKE2b-256 |
28e5b97921c3a7e59bbb4e1f0ce2c19cdb7ba03844f857e702af33f764da19e2
|
Provenance
The following attestation bundles were made for prpt-0.2.0-py3-none-any.whl:
Publisher:
publish.yml on steyangdot/PromptPilot
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
prpt-0.2.0-py3-none-any.whl -
Subject digest:
821e3043623d0d7b057c48a8db1029b2b0cf78c964f1d8b11740a10fe3fd2278 - Sigstore transparency entry: 1701724574
- Sigstore integration time:
-
Permalink:
steyangdot/PromptPilot@c90aa554081bf472fb6aad5135a390fff19f2181 -
Branch / Tag:
refs/tags/v0.2.0 - Owner: https://github.com/steyangdot
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@c90aa554081bf472fb6aad5135a390fff19f2181 -
Trigger Event:
release
-
Statement type: