Claudio - Claude Intelligence Optimizer
A control plane for the claude CLI: it makes a request well-specified before it costs anything, and enforces policy on what the answer may do.
Claudio is not a chatbot, not a new model, and — as of 2.0.0 — not a token compressor. It sits in front of the claude CLI and owns the decisions that surround a call: whether the request is specified well enough to be worth paying for, which model answers it, how hard that model thinks, what it is allowed to touch, and what the whole thing is permitted to cost.
Why not token reduction? Because the measurement doesn't support it. A one-line claudio ask bills ~24,000 input tokens, of which your prompt is about one. The rest is Claude Code's own system prompt, your CLAUDE.md, the tool definitions, and prompt-cache traffic — fixed overhead claudio cannot touch from outside. Trimming whitespace from an attached file is rounding error against that. 2.0.0 removed the lossy compression stage entirely, and stopped claiming a percentage off your input.
What actually moves cost per outcome is policy, and that is what claudio owns:
| Pillar | What it does |
|---|---|
| Specification — pre-flight check, free and local | names what a request leaves unpinned before you pay for a guess |
Routing — model tier + --effort |
the cheapest model and thinking depth that fits the job |
| Accounting — billed spend ledger | what each request actually cost, and why it was expensive |
| Enforcement — per-call and per-day budgets, permission postures | hard ceilings and a bounded blast radius, not advice |
None of these transform your words. The specification check reports what's missing and sends your prompt through untouched — silently reshaping a request is the same mistake as the compression stage, one layer up. Attached files are sent faithfully, in full, never summarized.
The premise: Claude doesn't need to change, and neither do your words. The conditions around the call do.
Install
Option 1: pip (recommended)
pip install claudio-cli
For closer token counts (BPE via tiktoken, still an approximation of Claude's
tokenizer) used by --estimate, --verbose, and stats:
pip install "claudio-cli[tokens]"
Without it, claudio falls back to a chars-per-token heuristic and labels every estimate accordingly.
Option 2: Source
git clone https://github.com/GuillaumeYves/claudio.git
cd claudio
pip install -e .
Both options install a single command: claudio.
- Run bare
claudioto drop inside the tool — interactive session with tab-completion, history, and slash commands (same feel asclaude). - Run
claudio <subcommand> …for one-shot use in scripts, pipes, or CI (claudio build -r @file "…",claudio ask -q "…").
Usage pattern — same as
claude:cdinto your project first, then launchclaudio.@filecompletion,claudio-task.json, and the.claudio/cache/folder all resolve against the current working directory. To switch projects mid-session without exiting, use/cwd <path>.
Requires Python 3.10+. Everything is bundled by default — no extras to enable.
Post-install setup
After installing, run:
claudio setup
This will:
- Set your permission posture — how much Claude may do on its own when you
build(see Permissions) - Detect if
claudiois on your PATH and offer to add it automatically (recommended for speed -- typeclaudiofrom any directory instead ofpython -m claudio) - Verify Claude CLI is installed
- Create config directory at
~/.config/claudio/ - Install shell completions (Bash, Zsh, or PowerShell) for tab-completing commands, modes, flags, and
@filepaths
If you decline any step, it shows you the exact command to do it manually. The permission step also runs automatically the first time you launch the claudio REPL, and is re-runnable anytime via /setup.
Claudio calls the Claude CLI under the hood. Install it separately if you haven't already. You can use --dry-run on any command to see the optimized prompt without sending it.
Commands
Claudio has 3 core commands + stats and setup. Each core command takes a mode, optional file attachments, and a description.
claudio <command> <mode> [@file [-lines]] ... [description]
Argument order is strictly enforced: mode first, then files, then description. This creates a logical parsing flow that lets Claude process the request with zero ambiguity.
claudio build
Create or modify code.
| Mode | Short | Purpose |
|---|---|---|
-refactor |
-r |
Refactor existing code (preserve behavior, improve structure) |
-generate |
-g |
Generate new code from a description |
Examples:
# Refactor a specific function (lines 40-80)
claudio build -refactor @src/auth.py -40-80 "extract the validation logic into its own function"
# Refactor with multiple context files
claudio build -refactor @src/handler.py @src/models.py "consolidate duplicate error handling"
# Generate new code using existing files as reference
claudio build -generate @src/models/user.py "create a REST endpoint for user CRUD operations"
# Generate from scratch
claudio build -generate "python script that watches a directory for CSV changes and loads them into SQLite"
Build applies its changes to disk. Unlike ask (read-only), build grants
Claude its editing tools and applies the edits directly, then prints the
resulting git diff plus a one-line summary. How much it may do is governed by
your permission posture (default Edits only — apply edits,
no shell commands); change it via claudio setup / /setup, or export
CLAUDIO_BUILD_PERMISSION_MODE=default for a one-off preview-only run. Use
--dry-run to see the optimized prompt without calling Claude at all.
Refactor output: edits applied in place + a short summary of what changed. Generate output: code written to the target file(s) + a short summary.
claudio ask
Ask Claude a question.
| Mode | Short | Purpose |
|---|---|---|
-review |
-rv |
Code review (security, quality, bugs) |
-question |
-q |
General question (explain, how-to, architecture) |
-debug |
-d |
Debug an issue (root cause, fix, explanation) |
Examples:
# Code review
claudio ask -review @src/auth.py "check for security issues"
# Review specific lines
claudio ask -review @src/api/handler.py -120-180 "is this input validation sufficient"
# Ask a question with file context
claudio ask -question @src/pipeline/process.py "how does the filter stage work"
# Ask without files
claudio ask -question "what is the difference between asyncio.gather and asyncio.wait"
# Debug with error context
claudio ask -debug @logs/error.log -500-520 "why is the connection pool exhausting"
# Debug specific code
claudio ask -debug @src/db.py -30-45 "this query returns duplicates when it shouldn't"
Review output: issues ranked by severity with fixes. Question output: concise, direct answer. Debug output: root cause, fix (as diff), brief explanation.
ask is always read-only — it never writes to disk, whatever your permission posture. If a request actually needs file changes, Claude says so and claudio offers to re-run it in build mode (resuming the same session). See Permissions.
claudio run
Execute a multi-step task plan from claudio-task.json.
# Execute the plan (prompts for confirmation)
claudio run
# Execute with additional file context
claudio run @src/config.py @docs/api-spec.md
# Preview all prompts without executing
claudio run --dry-run
# Run as a single agentic session with tool access (Read/Grep/Glob)
claudio run --agentic
claudio run always reads from claudio-task.json in the current directory. It validates the file, warns about missing fields, and asks for confirmation before executing.
Execution modes:
| Mode | How it runs | When to use |
|---|---|---|
| serial (default) | Oneclaude --print per task |
Tasks are independent; you want clean per-task output |
--agentic |
One session for the whole plan, withRead/Grep/Glob tools |
Tasks share reasoning, or Claude should discover files itself |
Agentic mode saves tokens on multi-task plans (no per-task prompt re-ingest) and lets Claude carry insight from task 1 into task 2.
Task file format:
{
"name": "Audit authentication module",
"tasks": [
{
"name": "Review auth middleware",
"prompt": "Review this middleware for security vulnerabilities",
"context": "This handles JWT validation for all API routes",
"intent": "review",
"constraints": ["Focus on token expiry handling", "Check for injection vectors"],
"output_format": "Severity-ranked list with fix suggestions"
},
{
"name": "Generate test cases",
"prompt": "Generate unit tests for the auth middleware edge cases",
"intent": "generate"
}
]
}
| Field | Required | Description |
|---|---|---|
name |
Yes (plan + each task) | Human-readable identifier |
tasks |
Yes | Array of task objects |
prompt |
Yes (per task) | What Claude should do |
context |
No | Additional input/context |
intent |
No | general, debug, refactor, generate, review |
constraints |
No | Array of requirements for the output |
output_format |
No | Expected output structure |
A template is provided at claudio-task.template.json.
claudio setup
Post-install configuration. Re-runnable anytime; the permission step also runs automatically the first time you launch the claudio REPL.
claudio setup
Sets your permission posture, checks PATH, installs shell completions, and verifies the Claude CLI. Inside the REPL, /setup re-runs just the permission picker.
Interactive Mode (claudio)
Running claudio drops you inside the tool — same feel as claude. Type commands directly at the prompt. No prefix required for anything.
claudio
█▀▀ █ ▄▀█ █ █ █▀▄ █ █▀█
█▄▄ █▄▄ █▀█ █▄█ █▄▀ █ █▄█
✻ Claudio v1.5.0
/help for commands · @ to reference files · Ctrl-D to exit
cwd: ~/Documents/Perso/claudio
claudio> ask -review @src/auth.py "any token-replay risk?"
claudio [ask -review]> what about input validation?
claudio [ask -review]> /mode build -r
claudio [build -r]> @main.py extract the duplicated try/except
Every command you'd run as claudio build … works as just build … inside the session. @ triggers live tab-completion from the current directory (.git, node_modules, __pycache__, and other noise are filtered out).
Sticky mode + files
After your first explicit ask -review @auth.py "...", the prompt becomes claudio [ask -review]> and follow-up lines that don't start with a command inherit both the mode and the @file set:
claudio [ask -review]> what about input validation?
claudio [ask -review]> any token-replay risk?
- Switch modes with a fresh command (
ask -q ...,build -r ...) or pin one without sending via/mode ask -review. - Adding new
@filetokens replaces the prior file set; bare prompts re-attach the previous ones. - First bare prompt before any mode is pinned defaults to
ask -qand prints a one-time hint nudging you to specify intent — the right mode means the right filter / output budget applies (review preserves comments where refactor strips them). /freshwipes sticky state along with the conversation.
Markdown rendering
Responses are rendered through a markdown → ANSI converter when stdout is a TTY: headers get cyan accent, bold is bold, italic is italic, inline code is green, fenced code blocks are dim, lists get cyan bullets, blockquotes get a dim rail, [links](url) show as cyan label + dim URL. Output stays plain when piped, in --json mode, or with NO_COLOR=1 / CLAUDIO_NO_COLOR=1. Streaming preserves styling — partial deltas are line-buffered so spans like **bold** never break mid-chunk.
Tool activity
While Claude uses tools (Read, Edit, Grep, Bash, …) the activity surfaces as either a live spinner update (⠋ claudio is reading auth.py (3.2s)) before text streams, or a dim stderr breadcrumb (↳ claudio is reading auth.py) once the response is mid-flight. Nothing pollutes the response itself.
Slash commands:
| Command | Purpose |
|---|---|
/help |
Show available commands |
/model NAME |
Pin a model for the session (haiku, sonnet, opus, fable). /model auto resets. |
/mode CMD MODE |
Pin a sticky mode (/mode ask -review). /mode alone shows current; /mode none clears. |
/setup |
Configure the permission posture (what Claude may do on its own) |
/cwd [PATH] |
Show or change the working directory |
/clear |
Clear the screen |
/fresh |
Start a new conversation (drops Claude's memory + sticky state) |
/session |
Print the current session id |
/undo |
Revert the files changed by the last build (restores pre-build content; removes files it created) |
/stats |
Shortcut forclaudio stats inside the REPL |
/exit | /quit |
Exit (Ctrl-D also works) |
History is stored at ~/.claudio/repl_history (or $CLAUDIO_HOME/repl_history). Errors in one command never kill the session — you stay inside until you exit.
The REPL requires a TTY. When stdin is piped or redirected, bare claudio exits with a hint pointing you at the one-shot form (claudio <subcommand> …).
Shell Completions
claudio setup installs completions automatically. To install manually:
Bash (add to ~/.bashrc):
eval "$(claudio --completions bash)"
Zsh (add to ~/.zshrc):
eval "$(claudio --completions zsh)"
PowerShell (add to $PROFILE):
claudio --completions powershell | Invoke-Expression
What you get:
claudio <TAB> -> build ask run stats setup
claudio build -<TAB> -> -refactor -r -generate -g
claudio ask -<TAB> -> -review -rv -question -q -debug -d
claudio ask -d @src/<TAB> -> @src/main.py @src/auth.py ...
Response Cache
Claudio caches responses locally. Same prompt = instant result, zero tokens spent.
- Location:
.claudio/cache/in your workspace (auto-gitignored) - Key: SHA-256 hash of the final optimized prompt
- TTL: 1 hour (expired entries are auto-cleaned)
- Scope: Per-workspace, because the same
@file.pyin different projects contains different code
Cache hits show a [cache hit] indicator:
[claudio] [cache hit] Returning cached response
Bypass cache for a single request:
claudio ask -question --no-cache @src/main.py "explain this"
Clear all cached responses:
claudio stats --reset
The cache is deterministic: if the file hasn't changed and your prompt is the same, you get the same answer. If you edit the file and run again, the prompt changes (different file contents) so you get a fresh response automatically.
Specification Check
Before a request is sent, claudio inspects it locally and for free for the things that actually degrade an answer — not poor wording, which Claude handles well, but missing specification: which file, which lines, what must not change, what "done" looks like.
$ claudio build -r "simplify"
[claudio:warn] under-specified request (sending anyway):
- no constraint on what the change must preserve
try: add what must not change, e.g. "...without altering the public API"
- only 1 word(s) of description
try: one more clause about the goal usually beats a longer file
(--strict-spec turns these into a refusal; --no-spec-check silences them)
Three rules keep it honest:
- It never rewrites your prompt. It reports the gap and sends your words through unchanged. A layer that silently reshapes intent is the same mistake as the lossy compression stage removed in 2.0.0.
- It costs nothing. Pure local inspection — no model call. A check that needed its own call could never pay for itself on a cheap request.
- It advises by default.
--strict-spec(orstrict_specin config) promotes findings to a refusal;--no-spec-checksilences them. Nobody is blocked by surprise.
A well-specified request produces no output at all — the check is designed around its false-positive rate, because a check that cries wolf gets ignored and then protects nothing.
Budgets
Two ceilings, deliberately different in scope:
| Setting | Scope | Behaviour |
|---|---|---|
--max-budget-usd N |
one call | the CLI halts mid-run rather than exceed it |
daily_budget_usd / CLAUDIO_DAILY_BUDGET_USD |
one day | refuses when spent; clamps each call to what's left |
The clamp is what makes the daily cap real. Claudio only regains control after a call finishes, so a check alone could never stop a single call from blowing the day's budget — instead, the remaining allowance is handed to the CLI as that call's own ceiling:
$ CLAUDIO_DAILY_BUDGET_USD=0.10 claudio ask -q "..."
[claudio:warn] daily budget: $0.0864 left of $0.10 — capping this call there
$ CLAUDIO_DAILY_BUDGET_USD=0.01 claudio ask -q "..."
[claudio:error] daily budget reached: $0.0136 of $0.01 spent today.
An explicit --max-budget-usd is respected when it's tighter than the daily
remainder, and can never raise the ceiling above it. A daily cap is only
meaningful because the ledger records billed figures (below) — enforcing it
against estimates would feel like protection while letting real spend through.
Cost Tracking
Every request is logged with the figures the claude CLI itself reports —
what Anthropic actually charged, prompt-cache traffic included. Claudio falls
back to local estimates only when the CLI reports nothing (a plain-text
buffered run). claudio stats labels which basis it is showing.
Why this matters: a local estimate can only see the prompt claudio composed.
The real request also carries Claude Code's system prompt, your CLAUDE.md,
every tool definition, and cache reads/writes. A trivial one-line ask
estimates at ~1 token but bills ~24,000 — so estimates read low by orders of
magnitude, and claudio no longer reports them as if they were the bill.
View your usage with:
claudio stats
Claudio Usage Stats
Period Requests Tokens In Cost Cache Hits
------------ --------- ----------- ---------- -----------
Today 8 3,200 $0.0340 2
This week 23 12,500 $0.1520 7
All time 91 48,000 $0.5800 19
By Command:
Command Requests Tokens In Cost
---------------------- --------- ----------- ----------
ask -review 12 8,000 $0.1200
build -refactor 15 6,200 $0.0900
ask -debug 8 4,800 $0.0700
...
Cache hit rate: 21% (19 of 91 requests)
JSON output for scripts/dashboards:
claudio stats --json
Reset all data (also clears cache):
claudio stats --reset
Usage data is stored at ~/.config/claudio/usage.json (global, persists across workspaces).
File Attachments
Attach up to 10 files from your workspace using @path:
claudio build -refactor @src/main.py "simplify error handling"
Add a line range immediately after any @file:
@src/main.py -10-25 # lines 10 through 25
@src/main.py -42 # line 42 only
@logs/error.log -500-520 # lines 500 through 520
Multiple files:
claudio ask -review @src/auth.py -30-60 @src/middleware.py @tests/test_auth.py "is the auth flow correct"
Each file gets its own line range. Claude receives only the lines that matter -- not entire files.
PowerShell note
PowerShell treats @name as the splatting operator and silently drops the token when $name isn't a defined variable — so claudio build -r @src/main.py "fix" becomes claudio build -r "fix" with no warning. Two safe alternatives on PowerShell:
claudio build -r '@src/main.py' "fix" # quote the @-token
claudio build -r -f src/main.py "fix" # use -f / --file instead
claudio build -r --file src/main.py -10-25 "fix"
-f / --file is rewritten to @<path> internally, so line ranges and all other features work identically. Bash, zsh, cmd, and the claudio REPL are unaffected.
Argument Order
Arguments must follow this order:
claudio <command> <mode> [@file [-lines]] ... [description]
1 2 3 4
| Position | What | Examples |
|---|---|---|
| 1 | Command | build, ask, run |
| 2 | Mode flag | -refactor, -review, -debug |
| 3 | Files + lines | @file.py -10-25 @other.py |
| 4 | Description | "your prompt text here" |
Wrong order = error:
claudio build "text" -refactor @file.py # ERROR: mode after description
claudio build -refactor "text" @file.py # ERROR: file after description
claudio build @file.py -refactor "text" # ERROR: mode after file
Why strict order? It eliminates parsing ambiguity. When Claude receives the structured prompt, every field is in a predictable position. No tokens wasted on disambiguation.
Global Flags
Global flags can go anywhere after the command:
claudio build -refactor --dry-run @file.py "simplify"
claudio ask -debug --verbose @log.txt "what happened"
claudio run --json
| Flag | Description |
|---|---|
--dry-run |
Print the optimized prompt without calling Claude |
--estimate |
Print token count + projected input cost, then exit without calling Claude |
--no-cache |
Bypass response cache for this request |
--verbose |
Show token count, model, and pipeline metadata |
--json |
Output results as structured JSON |
--model NAME |
Override model (haiku / sonnet / opus / fable or full ID) |
--effort LEVEL |
Thinking effort: low / medium / high / xhigh / max |
--max-budget-usd N |
Hard spend ceiling for this call; the run stops rather than exceed it |
--strict-spec |
Refuse an under-specified request instead of warning |
--no-spec-check |
Skip the pre-flight specification check |
--session-id UUID |
Start a session with a fixed ID (reusable via--resume) |
--resume UUID |
Resume an existing Claude session (warm prompt cache) |
--feedback |
Let Claude request missing context; auto-retry once with expanded range |
--agentic |
(claudio run only) Execute the plan in one agentic session with tool access |
-v, --version |
Print version |
-h, --help |
Show help |
Auto-Context
Claudio pulls three sources of context into every call automatically — no flag, no setup beyond dropping files in your repo. All three live in the cacheable prompt prefix, so you pay for them once per session.
Project preamble
If either of these files exists at the workspace root, its content is wrapped in a <project> tag at the very front of the prompt:
.claudio/project.md— claudio-specific tighter preamble (overrides CLAUDE.md when both present)CLAUDE.md— Claude Code's existing project memory file
Combined output capped at 2 KB. Disable: CLAUDIO_NO_PREAMBLE=1.
Stack detection
Claudio reads manifest files and emits a one-line stack summary alongside the preamble:
| Manifest | Detected |
|---|---|
pyproject.toml (PEP 621 or Poetry) |
Python version, project name, deps, framework |
requirements.txt |
Python deps, framework |
package.json |
JavaScript/TypeScript, node version, deps, framework |
Cargo.toml |
Rust edition, crate name |
go.mod |
Go version, module |
Identifies 12 common frameworks (Django, Flask, FastAPI, Next.js, React, Vue, Svelte, Express, NestJS, …). Disable: CLAUDIO_NO_STACK_DETECT=1.
Git changes
When the cwd is a git repo and the pipeline intent is review, debug, or refactor, a <changes> block is auto-included containing:
git diff HEAD --stat --patch(uncommitted work)git diff <base>...HEADagainstorigin/main,origin/master,main, ormaster— first ref that resolves (committed branch work)
Capped at 6 KB. <changes> sits in the volatile tail (right before <task>) since it shifts on every edit. Disable: CLAUDIO_NO_GIT_CONTEXT=1.
Net effect: ask claudio to review your in-progress work and Claude sees what changed, not the whole file with no signal about which lines moved.
How It Works
Every input goes through a three-stage pipeline:
Input --> Filter --> Prompt --> Claude
File bodies pass through faithfully — the only content-reducing stage is noise filtering below (whitespace, boilerplate, log dedup). Nothing is summarized or replaced with a map.
1. Filter (intent-aware)
Removes content that wastes tokens:
Always:
- Trailing whitespace from every line (~2-5% savings)
- License/copyright headers (legal boilerplate, not code)
- Shebang lines
- Consecutive blank lines collapsed to one
- Log deduplication (normalizes timestamps/UUIDs, shows repeat counts)
- Low-signal log lines (health checks, separators)
Only when intent is refactor:
- Strips comments (full-line and inline)
- Strips docstrings (triple-quote blocks)
- Rationale: a pure refactor is about structure, not stated intent.
Comments and docstrings are preserved for review and debug because that's exactly when they matter most — TODO markers, known-issue notes, and assertion docs are often where the bug lives. For question, all docs survive too since they may be what's being asked about.
No compression stage. Through 1.x, files over 300 lines were replaced with a lossy structural map. That was removed in 2.0.0: it silently traded fidelity for tokens on files you thought you'd attached in full. If a file is too large or costly to send whole, attach a narrower line range (
@file -START-END) — an explicit choice you make, not a summary claudio makes for you.
2. Prompt (XML-tagged, cache-aligned, zero duplication)
Builds a minimal prompt using XML tags instead of markdown, with the stable sections first and the variable tail last:
<project>
[from CLAUDE.md]
Stack: Django 4.2 + Postgres. Tests in pytest.
[stack]
Python >=3.10 (pyproject.toml) — project: myapp; framework: Django; deps: django, celery, ...
</project>
<rules>
- Preserve behavior
- Output unified diff
- One-line reason per change
</rules>
<format>diff with explanation</format>
<context>
<file path="auth.py" role="target" lines="40-80">
def validate_token(token):
...
</file>
</context>
<changes>
[uncommitted (git diff HEAD)]
diff --git a/auth.py ...
</changes>
<task>Refactor: extract validation logic</task>
Why this order? Anthropic's prompt cache keys on the prefix. <project>, <rules>, <format>, and <context> rarely change between back-to-back calls on the same file; <changes> shifts on every edit, and <task> shifts on every prompt. Putting them last means follow-up calls hit the 5-minute prompt cache instead of paying full ingest cost. Combine with --session-id / --resume for iterative workflows.
Why XML over markdown? Claude parses XML tags natively (it's the same format used for tool use). XML tags cost ~2 tokens each vs ~4-6 for ## Header + newlines. Across a session of 50 requests, this saves ~200 tokens of pure formatting overhead.
Zero duplication: The user's description appears exactly once in <task>. File contents appear exactly once in <context>.
No default padding: Question-mode (claudio ask -question) produces just <task>your question</task> -- no constraints, no format instructions, no boilerplate. Claude doesn't need to be told "be concise" on a simple question.
3. Execute
Sends to Claude CLI via claude --print. In --dry-run mode, prints the prompt instead.
While the call is in flight, a stderr spinner (| asking sonnet (2.3s)) shows progress with an elapsed-time counter. It stays silent when stderr is not a TTY (piped / CI / scripts) so captured output is never polluted, and it degrades to ASCII on legacy Windows consoles that can't render Unicode.
Flags plumbed to the Claude CLI when set:
--model→claude --model(auto-routed by intent + input size if unset; see below)--session-id/--resume→ session continuity across calls (warm prompt cache)--agentic(claudio run) → adds--allowedTools Read,Grep,Globfor agentic executionclaudio build→ adds--permission-mode <…>resolved from your permission posture (defaultacceptEdits) so the edits actually land on disk; mutating builds bypass the response cache since edits are side effects
Model Routing
When --model is not set, Claudio picks the cheapest model that fits the task:
| Intent | Input size | Model |
|---|---|---|
| question / general | < 2k tokens | haiku |
| review / refactor / debug | > 8k tokens | opus |
| any | > 20k tokens | opus |
| everything else | — | sonnet |
Override with --model haiku|sonnet|opus (or a full model ID) on any command. --verbose prints the resolved model.
Feedback Channel (--feedback)
--feedback opens a two-way channel. Claude can respond with one of two signals when something is off, and claudio honours it with a single auto-retry (no loops — max one retry per call).
<need-context> — data missing
When the attached line range is too narrow (e.g. a helper defined just outside the lines you pinned with @file -START-END), Claude requests specific line ranges:
<need-context file="PATH" lines="START-END" reason="..."/>
Multiple ranges in one response are allowed (back-to-back tags) and all are expanded in the single retry. Example:
claudio ask -review --feedback @src/auth.py "is the token flow correct?"
# -> Claude replies:
# <need-context file="auth.py" lines="120-180" reason="need validate_refresh body"/>
# <need-context file="auth.py" lines="40-60" reason="need helper used at line 145"/>
# -> Claudio re-runs once with both ranges included
<need-clarification> — task ambiguous
When the request itself is unclear (not the data), Claude asks back:
<need-clarification question="..."/>
In an interactive shell, claudio prompts you inline:
[claudio] Claude needs clarification: rename to camelCase or snake_case?
clarify > snake_case
…and resubmits with your answer appended. In non-interactive shells the question prints to stderr with a hint to re-run.
Both signals are gated behind --feedback and mutually exclusive in a single response.
Sessions (--session-id / --resume)
For iterative work, reuse a Claude session so the prompt cache stays warm and context carries across calls:
ID=$(uuidgen)
claudio ask -question --session-id $ID @src/pipeline.py "walk me through this"
claudio ask -question --resume $ID "now explain the model routing"
claudio ask -question --resume $ID "why XML over JSON in the prompt?"
--resume bypasses the local response cache (the point is to get a fresh Claude turn, not a stored echo).
Where the savings come from
Since 2.0.0 claudio sends files faithfully, so savings do not come from shrinking your code — and, measured honestly, they never mostly did. The fixed harness overhead on every call (system prompt, CLAUDE.md, tool definitions, cache traffic) dwarfs anything claudio can trim from your input. What is left are levers that change whether and how a call happens:
| Lever | What it saves | Applies to |
|---|---|---|
| Model routing | Ahaiku answer is ~5× cheaper than opus; trivial asks never pay opus rates |
every request, sized/intent-routed |
| Response cache | An identical prompt returns instantly forzero tokens | every repeat (great in iterative sessions) |
| Effort | --effort low on the same model costs less than dropping a tier, and often answers as well |
opt-in, per call or via config |
| Specification | A request that had to be re-asked cost double; the pre-flight check is free and catches it first | every request (advisory by default) |
| Noise filtering | Trailing whitespace, blank-line runs, license headers, log dedup; onrefactor, comments/docstrings too |
every request (single-digit % typically; more on comment-heavy refactors) |
Two measured points on the Claudio codebase itself (tiktoken BPE):
| Command | File size | Input | Sent | Note |
|---|---|---|---|---|
claudio ask -rv @repl.py "security" |
963 lines | 11,709 | 11,701 | sent in full — faithful, ~0% shaved |
claudio build -r @prompt.py "simplify" |
174 lines | 2,448 | 935 | 62% —refactor strips comments/docstrings |
If a file is genuinely too large or costly to send whole, attach a narrower range yourself (@file -START-END) — an explicit call, not a summary claudio makes behind your back.
Use --verbose to see the pre-flight estimate and, once the call returns,
what it actually billed:
[claudio] ~11,701 tokens (tiktoken BPE, est. $0.0585)
[claudio] model: opus
[claudio] billed: 36,412 in (19,189 cached) / 1,204 out | $0.1483 | claude-opus-5
The first line is claudio's pre-flight guess at the prompt it built; the last
is the CLI's own accounting of the whole request. Expect them to differ — the
gap is the system prompt, CLAUDE.md, tool definitions and cache traffic that
no local estimate can see.
Permissions
Claudio drives Claude headless (claude --print), where there's no mid-run popup to approve an edit — the decision has to be made up front. A single permission posture controls how much Claude may do on its own when you build:
| Posture | Claude may… | Under the hood |
|---|---|---|
| Autonomous | edit filesand run shell commands | --permission-mode bypassPermissions |
| Edits only (default) | edit files, but not run shell commands | --permission-mode acceptEdits |
| Confirm first | edit files after you approve once per build | acceptEdits + a Y/n gate |
| Preview only | nothing — just describe the change | --permission-mode plan |
Because nothing pauses mid-run, Confirm first is a single coarse Y/n gate before a build starts, not a per-edit prompt.
Preview only uses Claude Code's plan mode. Earlier versions passed no mode
at all and relied on headless --print auto-denying mutations — which worked,
but only as a side effect: Claude still attempted edits and had them denied,
so you paid for tool calls that could never land. In plan mode Claude plans
instead of editing, and nothing is attempted in the first place.
Set it with the first-run wizard (auto-runs the first time you launch claudio), or anytime via /setup in the REPL or claudio setup:
How much can Claude do on its own?
1) Autonomous - apply edits AND run shell commands automatically
2) Edits only - apply edits automatically, never run commands (default)
3) Confirm first - ask once before each build applies, then apply
4) Preview only - never apply - just print the diff
The choice is saved as permission_posture in ~/.config/claudio/config.json. A CLAUDIO_BUILD_PERMISSION_MODE env var still overrides the resolved --permission-mode for one-off runs.
Read-only modes escalate instead of failing
ask, review, question, and debug are always read-only and ignore the posture entirely. But rather than silently attempting an edit the headless CLI would deny, Claude signals when a request actually needs build mode — and claudio offers to switch:
claudio [ask -review]> add a ROADMAP.md from your analysis
⚠ This needs build mode: writing a new file is a mutation, not a review.
Re-run in build mode now? [Y/n]
On Y, claudio re-runs the request in build mode, resuming the same session so the analysis it just produced carries straight into the build.
Configuration
Claudio looks for config at ~/.config/claudio/config.json:
{
"claude_binary": "claude",
"default_model": "sonnet",
"max_input_tokens": 32000,
"output_format": "text",
"verbose": false,
"permission_posture": "edits",
"effort": null,
"max_budget_usd": null,
"daily_budget_usd": null,
"spec_check": true,
"strict_spec": false,
"partial_messages": true
}
All fields are optional. Defaults are used for anything not specified. permission_posture is one of autonomous, edits, confirm, or preview (see Permissions) — normally set by the wizard rather than by hand.
| Key | Meaning |
|---|---|
effort |
Thinking depth: low / medium / high / xhigh / max; null = CLI default |
max_budget_usd |
Hard ceiling for a single call; null = uncapped |
daily_budget_usd |
Cumulative ceiling per day; null = uncapped (see Budgets) |
spec_check |
Run the pre-flight specification check |
strict_spec |
Treat specification findings as a refusal rather than a warning |
partial_messages |
Request token-by-token streaming deltas |
Every one of these also has an env override (CLAUDIO_EFFORT, CLAUDIO_DAILY_BUDGET_USD, CLAUDIO_NO_PARTIAL, …) so a one-off run never needs a config edit.
Piping and Composability
Results go to stdout, info/warnings to stderr:
# JSON output piped to jq
claudio --json ask -question @api.py "list the public functions" | jq '.result'
# Use in scripts
if claudio --dry-run --verbose build -refactor @src/ 2>&1 | grep -q "WARNING"; then
echo "Input too large, consider narrowing scope"
fi
Project Structure
claudio/
cli.py Command router (subcommand dispatch)
repl.py Unified `claudio` entry point (REPL + subcommand dispatch)
cache.py Response cache (SHA-256 keyed, TTL-based)
usage.py Cost and usage tracking
config.py Configuration management
budget.py Cumulative spend enforcement (daily cap + per-call clamp)
spec_check.py Pre-flight specification check (local, free, non-rewriting)
executor.py Claude CLI integration (streaming + retry + spinner)
session_files.py Per-session file-hash tracking for unchanged markers
build_snapshot.py Pre-build file snapshots for /undo rollback
commands/
build.py claudio build (-refactor, -generate)
ask.py claudio ask (-review, -question, -debug)
run.py claudio run (claudio-task.json)
run_prompt.py Shared execution with cache + tracking + feedback channels
stats.py claudio stats (usage dashboard)
setup.py claudio setup (permission posture, PATH, completions, verification)
completions/
bash.py Bash completion generator
zsh.py Zsh completion generator
powershell.py PowerShell completion generator
pipeline/
filter.py Intent-aware noise filtering
prompt.py XML-tagged prompt construction
process.py Pipeline orchestrator (faithful — no compression)
utils/
args.py @file parser with strict order enforcement
project_context.py Project preamble discovery (CLAUDE.md + .claudio/project.md)
stack_detect.py Stack detection from manifest files
git_context.py Auto git-diff context for review/debug/refactor
tokens.py Token estimation (tiktoken when available)
output.py Output formatting (with markdown rendering)
files.py File reading and ingestion
colors.py TTY-aware ANSI color helpers
markdown.py Markdown -> ANSI renderer (streaming + buffered)
spinner.py TTY-only stderr progress spinner
update_check.py Background PyPI version check
model_router.py Cheapest-model-that-fits routing
tests/ pytest suite (336 tests across pipeline, REPL, context, executor)
Releases
Claudio uses tag-based releases. Each GitHub release auto-publishes to PyPI.
# Update to latest
pip install --upgrade claudio-cli
# Install a specific version
pip install claudio-cli==0.2.0
Metadata
Release files for claudio-cli 2.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| claudio_cli-2.0.0.tar.gz | 183.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| claudio_cli-2.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 314.2 kB
Release files / claudio_cli-2.0.0.tar.gz
| Download URL | claudio_cli-2.0.0.tar.gz |
|---|---|
| Size | 183.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
2b9e23d43d4015bbdabe652cee203574c0716c018e35a44ae992c9c0e11f0d7d
|
|
BLAKE2b-256 checksum How to use checksums |
7f90d2fa1295b3908bf08333ddc5423b4398461eeb622dbbacc2f73da6044826
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.
Transparency logRelease files / claudio_cli-2.0.0-py3-none-any.whl
| Download URL | claudio_cli-2.0.0-py3-none-any.whl |
|---|---|
| Size | 130.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
1807b64867f6a46bd943c204296bf6cd81aa4940260ec3be2930e06475a65dfe
|
|
BLAKE2b-256 checksum How to use checksums |
4431df22c4ce06b99a48377b8fe21dacece2b79b0421166aebf12473bae65c59
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.
Transparency log