Skip to main content

claudio logo

Claudio - Claude Intelligence Optimizer

Release PyPI Python License

A control plane for the claude CLI: it makes a request well-specified before it costs anything, and enforces policy on what the answer may do.

Claudio is not a chatbot, not a new model, and — as of 2.0.0 — not a token compressor. It sits in front of the claude CLI and owns the decisions that surround a call: whether the request is specified well enough to be worth paying for, which model answers it, how hard that model thinks, what it is allowed to touch, and what the whole thing is permitted to cost.

Why not token reduction? Because the measurement doesn't support it. A one-line claudio ask bills ~24,000 input tokens, of which your prompt is about one. The rest is Claude Code's own system prompt, your CLAUDE.md, the tool definitions, and prompt-cache traffic — fixed overhead claudio cannot touch from outside. Trimming whitespace from an attached file is rounding error against that. 2.0.0 removed the lossy compression stage entirely, and stopped claiming a percentage off your input.

What actually moves cost per outcome is policy, and that is what claudio owns:

Pillar What it does
Specification — pre-flight check, free and local names what a request leaves unpinned before you pay for a guess
Routing — model tier + --effort the cheapest model and thinking depth that fits the job
Accounting — billed spend ledger what each request actually cost, and why it was expensive
Enforcement — per-call and per-day budgets, permission postures hard ceilings and a bounded blast radius, not advice

None of these transform your words. The specification check reports what's missing and sends your prompt through untouched — silently reshaping a request is the same mistake as the compression stage, one layer up. Attached files are sent faithfully, in full, never summarized.

The premise: Claude doesn't need to change, and neither do your words. The conditions around the call do.


Install

pip install claudio-cli

For closer token counts (BPE via tiktoken, still an approximation of Claude's tokenizer) used by --estimate, --verbose, and stats:

pip install "claudio-cli[tokens]"

Without it, claudio falls back to a chars-per-token heuristic and labels every estimate accordingly.

Option 2: Source

git clone https://github.com/GuillaumeYves/claudio.git
cd claudio
pip install -e .

Both options install a single command: claudio.

  • Run bare claudio to drop inside the tool — interactive session with tab-completion, history, and slash commands (same feel as claude).
  • Run claudio <subcommand> … for one-shot use in scripts, pipes, or CI (claudio build -r @file "…", claudio ask -q "…").

Usage pattern — same as claude: cd into your project first, then launch claudio. @file completion, claudio-task.json, and the .claudio/cache/ folder all resolve against the current working directory. To switch projects mid-session without exiting, use /cwd <path>.

Requires Python 3.10+. Everything is bundled by default — no extras to enable.

Post-install setup

After installing, run:

claudio setup

This will:

  • Set your permission posture — how much Claude may do on its own when you build (see Permissions)
  • Detect if claudio is on your PATH and offer to add it automatically (recommended for speed -- type claudio from any directory instead of python -m claudio)
  • Verify Claude CLI is installed
  • Create config directory at ~/.config/claudio/
  • Install shell completions (Bash, Zsh, or PowerShell) for tab-completing commands, modes, flags, and @file paths

If you decline any step, it shows you the exact command to do it manually. The permission step also runs automatically the first time you launch the claudio REPL, and is re-runnable anytime via /setup.

Claudio calls the Claude CLI under the hood. Install it separately if you haven't already. You can use --dry-run on any command to see the optimized prompt without sending it.


Commands

Claudio has 3 core commands + stats and setup. Each core command takes a mode, optional file attachments, and a description.

claudio <command> <mode> [@file [-lines]] ... [description]

Argument order is strictly enforced: mode first, then files, then description. This creates a logical parsing flow that lets Claude process the request with zero ambiguity.


claudio build

Create or modify code.

Mode Short Purpose
-refactor -r Refactor existing code (preserve behavior, improve structure)
-generate -g Generate new code from a description

Examples:

# Refactor a specific function (lines 40-80)
claudio build -refactor @src/auth.py -40-80 "extract the validation logic into its own function"

# Refactor with multiple context files
claudio build -refactor @src/handler.py @src/models.py "consolidate duplicate error handling"

# Generate new code using existing files as reference
claudio build -generate @src/models/user.py "create a REST endpoint for user CRUD operations"

# Generate from scratch
claudio build -generate "python script that watches a directory for CSV changes and loads them into SQLite"

Build applies its changes to disk. Unlike ask (read-only), build grants Claude its editing tools and applies the edits directly, then prints the resulting git diff plus a one-line summary. How much it may do is governed by your permission posture (default Edits only — apply edits, no shell commands); change it via claudio setup / /setup, or export CLAUDIO_BUILD_PERMISSION_MODE=default for a one-off preview-only run. Use --dry-run to see the optimized prompt without calling Claude at all.

Refactor output: edits applied in place + a short summary of what changed. Generate output: code written to the target file(s) + a short summary.


claudio ask

Ask Claude a question.

Mode Short Purpose
-review -rv Code review (security, quality, bugs)
-question -q General question (explain, how-to, architecture)
-debug -d Debug an issue (root cause, fix, explanation)

Examples:

# Code review
claudio ask -review @src/auth.py "check for security issues"

# Review specific lines
claudio ask -review @src/api/handler.py -120-180 "is this input validation sufficient"

# Ask a question with file context
claudio ask -question @src/pipeline/process.py "how does the filter stage work"

# Ask without files
claudio ask -question "what is the difference between asyncio.gather and asyncio.wait"

# Debug with error context
claudio ask -debug @logs/error.log -500-520 "why is the connection pool exhausting"

# Debug specific code
claudio ask -debug @src/db.py -30-45 "this query returns duplicates when it shouldn't"

Review output: issues ranked by severity with fixes. Question output: concise, direct answer. Debug output: root cause, fix (as diff), brief explanation.

ask is always read-only — it never writes to disk, whatever your permission posture. If a request actually needs file changes, Claude says so and claudio offers to re-run it in build mode (resuming the same session). See Permissions.


claudio run

Execute a multi-step task plan from claudio-task.json.

# Execute the plan (prompts for confirmation)
claudio run

# Execute with additional file context
claudio run @src/config.py @docs/api-spec.md

# Preview all prompts without executing
claudio run --dry-run

# Run as a single agentic session with tool access (Read/Grep/Glob)
claudio run --agentic

claudio run always reads from claudio-task.json in the current directory. It validates the file, warns about missing fields, and asks for confirmation before executing.

Execution modes:

Mode How it runs When to use
serial (default) Oneclaude --print per task Tasks are independent; you want clean per-task output
--agentic One session for the whole plan, withRead/Grep/Glob tools Tasks share reasoning, or Claude should discover files itself

Agentic mode saves tokens on multi-task plans (no per-task prompt re-ingest) and lets Claude carry insight from task 1 into task 2.

Task file format:

{
  "name": "Audit authentication module",
  "tasks": [
    {
      "name": "Review auth middleware",
      "prompt": "Review this middleware for security vulnerabilities",
      "context": "This handles JWT validation for all API routes",
      "intent": "review",
      "constraints": ["Focus on token expiry handling", "Check for injection vectors"],
      "output_format": "Severity-ranked list with fix suggestions"
    },
    {
      "name": "Generate test cases",
      "prompt": "Generate unit tests for the auth middleware edge cases",
      "intent": "generate"
    }
  ]
}
Field Required Description
name Yes (plan + each task) Human-readable identifier
tasks Yes Array of task objects
prompt Yes (per task) What Claude should do
context No Additional input/context
intent No general, debug, refactor, generate, review
constraints No Array of requirements for the output
output_format No Expected output structure

A template is provided at claudio-task.template.json.


claudio setup

Post-install configuration. Re-runnable anytime; the permission step also runs automatically the first time you launch the claudio REPL.

claudio setup

Sets your permission posture, checks PATH, installs shell completions, and verifies the Claude CLI. Inside the REPL, /setup re-runs just the permission picker.


Interactive Mode (claudio)

Running claudio drops you inside the tool — same feel as claude. Type commands directly at the prompt. No prefix required for anything.

claudio
  █▀▀ █   ▄▀█ █ █ █▀▄ █ █▀█
  █▄▄ █▄▄ █▀█ █▄█ █▄▀ █ █▄█

  ✻ Claudio v1.5.0

  /help for commands  ·  @ to reference files  ·  Ctrl-D to exit
  cwd: ~/Documents/Perso/claudio

claudio> ask -review @src/auth.py "any token-replay risk?"
claudio [ask -review]> what about input validation?
claudio [ask -review]> /mode build -r
claudio [build -r]> @main.py extract the duplicated try/except

Every command you'd run as claudio build … works as just build … inside the session. @ triggers live tab-completion from the current directory (.git, node_modules, __pycache__, and other noise are filtered out).

Sticky mode + files

After your first explicit ask -review @auth.py "...", the prompt becomes claudio [ask -review]> and follow-up lines that don't start with a command inherit both the mode and the @file set:

claudio [ask -review]> what about input validation?
claudio [ask -review]> any token-replay risk?
  • Switch modes with a fresh command (ask -q ..., build -r ...) or pin one without sending via /mode ask -review.
  • Adding new @file tokens replaces the prior file set; bare prompts re-attach the previous ones.
  • First bare prompt before any mode is pinned defaults to ask -q and prints a one-time hint nudging you to specify intent — the right mode means the right filter / output budget applies (review preserves comments where refactor strips them).
  • /fresh wipes sticky state along with the conversation.

Markdown rendering

Responses are rendered through a markdown → ANSI converter when stdout is a TTY: headers get cyan accent, bold is bold, italic is italic, inline code is green, fenced code blocks are dim, lists get cyan bullets, blockquotes get a dim rail, [links](url) show as cyan label + dim URL. Output stays plain when piped, in --json mode, or with NO_COLOR=1 / CLAUDIO_NO_COLOR=1. Streaming preserves styling — partial deltas are line-buffered so spans like **bold** never break mid-chunk.

Tool activity

While Claude uses tools (Read, Edit, Grep, Bash, …) the activity surfaces as either a live spinner update (⠋ claudio is reading auth.py (3.2s)) before text streams, or a dim stderr breadcrumb (↳ claudio is reading auth.py) once the response is mid-flight. Nothing pollutes the response itself.

Slash commands:

Command Purpose
/help Show available commands
/model NAME Pin a model for the session (haiku, sonnet, opus, fable). /model auto resets.
/mode CMD MODE Pin a sticky mode (/mode ask -review). /mode alone shows current; /mode none clears.
/setup Configure the permission posture (what Claude may do on its own)
/cwd [PATH] Show or change the working directory
/clear Clear the screen
/fresh Start a new conversation (drops Claude's memory + sticky state)
/session Print the current session id
/undo Revert the files changed by the last build (restores pre-build content; removes files it created)
/stats Shortcut forclaudio stats inside the REPL
/exit | /quit Exit (Ctrl-D also works)

History is stored at ~/.claudio/repl_history (or $CLAUDIO_HOME/repl_history). Errors in one command never kill the session — you stay inside until you exit.

The REPL requires a TTY. When stdin is piped or redirected, bare claudio exits with a hint pointing you at the one-shot form (claudio <subcommand> …).


Shell Completions

claudio setup installs completions automatically. To install manually:

Bash (add to ~/.bashrc):

eval "$(claudio --completions bash)"

Zsh (add to ~/.zshrc):

eval "$(claudio --completions zsh)"

PowerShell (add to $PROFILE):

claudio --completions powershell | Invoke-Expression

What you get:

claudio <TAB>              -> build  ask  run  stats  setup
claudio build -<TAB>       -> -refactor  -r  -generate  -g
claudio ask -<TAB>         -> -review  -rv  -question  -q  -debug  -d
claudio ask -d @src/<TAB>  -> @src/main.py  @src/auth.py  ...

Response Cache

Claudio caches responses locally. Same prompt = instant result, zero tokens spent.

  • Location: .claudio/cache/ in your workspace (auto-gitignored)
  • Key: SHA-256 hash of the final optimized prompt
  • TTL: 1 hour (expired entries are auto-cleaned)
  • Scope: Per-workspace, because the same @file.py in different projects contains different code

Cache hits show a [cache hit] indicator:

[claudio] [cache hit] Returning cached response

Bypass cache for a single request:

claudio ask -question --no-cache @src/main.py "explain this"

Clear all cached responses:

claudio stats --reset

The cache is deterministic: if the file hasn't changed and your prompt is the same, you get the same answer. If you edit the file and run again, the prompt changes (different file contents) so you get a fresh response automatically.


Specification Check

Before a request is sent, claudio inspects it locally and for free for the things that actually degrade an answer — not poor wording, which Claude handles well, but missing specification: which file, which lines, what must not change, what "done" looks like.

$ claudio build -r "simplify"
[claudio:warn] under-specified request (sending anyway):
  - no constraint on what the change must preserve
    try: add what must not change, e.g. "...without altering the public API"
  - only 1 word(s) of description
    try: one more clause about the goal usually beats a longer file
  (--strict-spec turns these into a refusal; --no-spec-check silences them)

Three rules keep it honest:

  1. It never rewrites your prompt. It reports the gap and sends your words through unchanged. A layer that silently reshapes intent is the same mistake as the lossy compression stage removed in 2.0.0.
  2. It costs nothing. Pure local inspection — no model call. A check that needed its own call could never pay for itself on a cheap request.
  3. It advises by default. --strict-spec (or strict_spec in config) promotes findings to a refusal; --no-spec-check silences them. Nobody is blocked by surprise.

A well-specified request produces no output at all — the check is designed around its false-positive rate, because a check that cries wolf gets ignored and then protects nothing.


Budgets

Two ceilings, deliberately different in scope:

Setting Scope Behaviour
--max-budget-usd N one call the CLI halts mid-run rather than exceed it
daily_budget_usd / CLAUDIO_DAILY_BUDGET_USD one day refuses when spent; clamps each call to what's left

The clamp is what makes the daily cap real. Claudio only regains control after a call finishes, so a check alone could never stop a single call from blowing the day's budget — instead, the remaining allowance is handed to the CLI as that call's own ceiling:

$ CLAUDIO_DAILY_BUDGET_USD=0.10 claudio ask -q "..."
[claudio:warn] daily budget: $0.0864 left of $0.10 — capping this call there

$ CLAUDIO_DAILY_BUDGET_USD=0.01 claudio ask -q "..."
[claudio:error] daily budget reached: $0.0136 of $0.01 spent today.

An explicit --max-budget-usd is respected when it's tighter than the daily remainder, and can never raise the ceiling above it. A daily cap is only meaningful because the ledger records billed figures (below) — enforcing it against estimates would feel like protection while letting real spend through.


Cost Tracking

Every request is logged with the figures the claude CLI itself reports — what Anthropic actually charged, prompt-cache traffic included. Claudio falls back to local estimates only when the CLI reports nothing (a plain-text buffered run). claudio stats labels which basis it is showing.

Why this matters: a local estimate can only see the prompt claudio composed. The real request also carries Claude Code's system prompt, your CLAUDE.md, every tool definition, and cache reads/writes. A trivial one-line ask estimates at ~1 token but bills ~24,000 — so estimates read low by orders of magnitude, and claudio no longer reports them as if they were the bill.

View your usage with:

claudio stats
Claudio Usage Stats

  Period        Requests   Tokens In       Cost  Cache Hits
  ------------ --------- ----------- ---------- -----------
  Today                8       3,200    $0.0340           2
  This week           23      12,500    $0.1520           7
  All time            91      48,000    $0.5800          19

  By Command:
  Command                 Requests   Tokens In       Cost
  ---------------------- --------- ----------- ----------
  ask -review                   12       8,000    $0.1200
  build -refactor               15       6,200    $0.0900
  ask -debug                     8       4,800    $0.0700
  ...

  Cache hit rate: 21% (19 of 91 requests)

JSON output for scripts/dashboards:

claudio stats --json

Reset all data (also clears cache):

claudio stats --reset

Usage data is stored at ~/.config/claudio/usage.json (global, persists across workspaces).


File Attachments

Attach up to 10 files from your workspace using @path:

claudio build -refactor @src/main.py "simplify error handling"

Add a line range immediately after any @file:

@src/main.py -10-25          # lines 10 through 25
@src/main.py -42             # line 42 only
@logs/error.log -500-520     # lines 500 through 520

Multiple files:

claudio ask -review @src/auth.py -30-60 @src/middleware.py @tests/test_auth.py "is the auth flow correct"

Each file gets its own line range. Claude receives only the lines that matter -- not entire files.

PowerShell note

PowerShell treats @name as the splatting operator and silently drops the token when $name isn't a defined variable — so claudio build -r @src/main.py "fix" becomes claudio build -r "fix" with no warning. Two safe alternatives on PowerShell:

claudio build -r '@src/main.py' "fix"          # quote the @-token
claudio build -r -f src/main.py "fix"          # use -f / --file instead
claudio build -r --file src/main.py -10-25 "fix"

-f / --file is rewritten to @<path> internally, so line ranges and all other features work identically. Bash, zsh, cmd, and the claudio REPL are unaffected.


Argument Order

Arguments must follow this order:

claudio <command> <mode> [@file [-lines]] ... [description]
     1          2          3                    4
Position What Examples
1 Command build, ask, run
2 Mode flag -refactor, -review, -debug
3 Files + lines @file.py -10-25 @other.py
4 Description "your prompt text here"

Wrong order = error:

claudio build "text" -refactor @file.py      # ERROR: mode after description
claudio build -refactor "text" @file.py      # ERROR: file after description
claudio build @file.py -refactor "text"      # ERROR: mode after file

Why strict order? It eliminates parsing ambiguity. When Claude receives the structured prompt, every field is in a predictable position. No tokens wasted on disambiguation.


Global Flags

Global flags can go anywhere after the command:

claudio build -refactor --dry-run @file.py "simplify"
claudio ask -debug --verbose @log.txt "what happened"
claudio run --json
Flag Description
--dry-run Print the optimized prompt without calling Claude
--estimate Print token count + projected input cost, then exit without calling Claude
--no-cache Bypass response cache for this request
--verbose Show token count, model, and pipeline metadata
--json Output results as structured JSON
--model NAME Override model (haiku / sonnet / opus / fable or full ID)
--effort LEVEL Thinking effort: low / medium / high / xhigh / max
--max-budget-usd N Hard spend ceiling for this call; the run stops rather than exceed it
--strict-spec Refuse an under-specified request instead of warning
--no-spec-check Skip the pre-flight specification check
--session-id UUID Start a session with a fixed ID (reusable via--resume)
--resume UUID Resume an existing Claude session (warm prompt cache)
--feedback Let Claude request missing context; auto-retry once with expanded range
--agentic (claudio run only) Execute the plan in one agentic session with tool access
-v, --version Print version
-h, --help Show help

Auto-Context

Claudio pulls three sources of context into every call automatically — no flag, no setup beyond dropping files in your repo. All three live in the cacheable prompt prefix, so you pay for them once per session.

Project preamble

If either of these files exists at the workspace root, its content is wrapped in a <project> tag at the very front of the prompt:

  • .claudio/project.md — claudio-specific tighter preamble (overrides CLAUDE.md when both present)
  • CLAUDE.md — Claude Code's existing project memory file

Combined output capped at 2 KB. Disable: CLAUDIO_NO_PREAMBLE=1.

Stack detection

Claudio reads manifest files and emits a one-line stack summary alongside the preamble:

Manifest Detected
pyproject.toml (PEP 621 or Poetry) Python version, project name, deps, framework
requirements.txt Python deps, framework
package.json JavaScript/TypeScript, node version, deps, framework
Cargo.toml Rust edition, crate name
go.mod Go version, module

Identifies 12 common frameworks (Django, Flask, FastAPI, Next.js, React, Vue, Svelte, Express, NestJS, …). Disable: CLAUDIO_NO_STACK_DETECT=1.

Git changes

When the cwd is a git repo and the pipeline intent is review, debug, or refactor, a <changes> block is auto-included containing:

  • git diff HEAD --stat --patch (uncommitted work)
  • git diff <base>...HEAD against origin/main, origin/master, main, or master — first ref that resolves (committed branch work)

Capped at 6 KB. <changes> sits in the volatile tail (right before <task>) since it shifts on every edit. Disable: CLAUDIO_NO_GIT_CONTEXT=1.

Net effect: ask claudio to review your in-progress work and Claude sees what changed, not the whole file with no signal about which lines moved.


How It Works

Every input goes through a three-stage pipeline:

Input --> Filter --> Prompt --> Claude

File bodies pass through faithfully — the only content-reducing stage is noise filtering below (whitespace, boilerplate, log dedup). Nothing is summarized or replaced with a map.

1. Filter (intent-aware)

Removes content that wastes tokens:

Always:

  • Trailing whitespace from every line (~2-5% savings)
  • License/copyright headers (legal boilerplate, not code)
  • Shebang lines
  • Consecutive blank lines collapsed to one
  • Log deduplication (normalizes timestamps/UUIDs, shows repeat counts)
  • Low-signal log lines (health checks, separators)

Only when intent is refactor:

  • Strips comments (full-line and inline)
  • Strips docstrings (triple-quote blocks)
  • Rationale: a pure refactor is about structure, not stated intent.

Comments and docstrings are preserved for review and debug because that's exactly when they matter most — TODO markers, known-issue notes, and assertion docs are often where the bug lives. For question, all docs survive too since they may be what's being asked about.

No compression stage. Through 1.x, files over 300 lines were replaced with a lossy structural map. That was removed in 2.0.0: it silently traded fidelity for tokens on files you thought you'd attached in full. If a file is too large or costly to send whole, attach a narrower line range (@file -START-END) — an explicit choice you make, not a summary claudio makes for you.

2. Prompt (XML-tagged, cache-aligned, zero duplication)

Builds a minimal prompt using XML tags instead of markdown, with the stable sections first and the variable tail last:

<project>
[from CLAUDE.md]
Stack: Django 4.2 + Postgres. Tests in pytest.
[stack]
Python >=3.10 (pyproject.toml) — project: myapp; framework: Django; deps: django, celery, ...
</project>
<rules>
- Preserve behavior
- Output unified diff
- One-line reason per change
</rules>
<format>diff with explanation</format>
<context>
<file path="auth.py" role="target" lines="40-80">
def validate_token(token):
    ...
</file>
</context>
<changes>
[uncommitted (git diff HEAD)]
diff --git a/auth.py ...
</changes>
<task>Refactor: extract validation logic</task>

Why this order? Anthropic's prompt cache keys on the prefix. <project>, <rules>, <format>, and <context> rarely change between back-to-back calls on the same file; <changes> shifts on every edit, and <task> shifts on every prompt. Putting them last means follow-up calls hit the 5-minute prompt cache instead of paying full ingest cost. Combine with --session-id / --resume for iterative workflows.

Why XML over markdown? Claude parses XML tags natively (it's the same format used for tool use). XML tags cost ~2 tokens each vs ~4-6 for ## Header + newlines. Across a session of 50 requests, this saves ~200 tokens of pure formatting overhead.

Zero duplication: The user's description appears exactly once in <task>. File contents appear exactly once in <context>.

No default padding: Question-mode (claudio ask -question) produces just <task>your question</task> -- no constraints, no format instructions, no boilerplate. Claude doesn't need to be told "be concise" on a simple question.

3. Execute

Sends to Claude CLI via claude --print. In --dry-run mode, prints the prompt instead.

While the call is in flight, a stderr spinner (| asking sonnet (2.3s)) shows progress with an elapsed-time counter. It stays silent when stderr is not a TTY (piped / CI / scripts) so captured output is never polluted, and it degrades to ASCII on legacy Windows consoles that can't render Unicode.

Flags plumbed to the Claude CLI when set:

  • --model → claude --model (auto-routed by intent + input size if unset; see below)
  • --session-id / --resume → session continuity across calls (warm prompt cache)
  • --agentic (claudio run) → adds --allowedTools Read,Grep,Glob for agentic execution
  • claudio build → adds --permission-mode <…> resolved from your permission posture (default acceptEdits) so the edits actually land on disk; mutating builds bypass the response cache since edits are side effects

Model Routing

When --model is not set, Claudio picks the cheapest model that fits the task:

Intent Input size Model
question / general < 2k tokens haiku
review / refactor / debug > 8k tokens opus
any > 20k tokens opus
everything else — sonnet

Override with --model haiku|sonnet|opus (or a full model ID) on any command. --verbose prints the resolved model.


Feedback Channel (--feedback)

--feedback opens a two-way channel. Claude can respond with one of two signals when something is off, and claudio honours it with a single auto-retry (no loops — max one retry per call).

<need-context> — data missing

When the attached line range is too narrow (e.g. a helper defined just outside the lines you pinned with @file -START-END), Claude requests specific line ranges:

<need-context file="PATH" lines="START-END" reason="..."/>

Multiple ranges in one response are allowed (back-to-back tags) and all are expanded in the single retry. Example:

claudio ask -review --feedback @src/auth.py "is the token flow correct?"
# -> Claude replies:
#    <need-context file="auth.py" lines="120-180" reason="need validate_refresh body"/>
#    <need-context file="auth.py" lines="40-60"  reason="need helper used at line 145"/>
# -> Claudio re-runs once with both ranges included

<need-clarification> — task ambiguous

When the request itself is unclear (not the data), Claude asks back:

<need-clarification question="..."/>

In an interactive shell, claudio prompts you inline:

[claudio] Claude needs clarification: rename to camelCase or snake_case?
clarify > snake_case

…and resubmits with your answer appended. In non-interactive shells the question prints to stderr with a hint to re-run.

Both signals are gated behind --feedback and mutually exclusive in a single response.


Sessions (--session-id / --resume)

For iterative work, reuse a Claude session so the prompt cache stays warm and context carries across calls:

ID=$(uuidgen)
claudio ask -question --session-id $ID @src/pipeline.py "walk me through this"
claudio ask -question --resume $ID "now explain the model routing"
claudio ask -question --resume $ID "why XML over JSON in the prompt?"

--resume bypasses the local response cache (the point is to get a fresh Claude turn, not a stored echo).


Where the savings come from

Since 2.0.0 claudio sends files faithfully, so savings do not come from shrinking your code — and, measured honestly, they never mostly did. The fixed harness overhead on every call (system prompt, CLAUDE.md, tool definitions, cache traffic) dwarfs anything claudio can trim from your input. What is left are levers that change whether and how a call happens:

Lever What it saves Applies to
Model routing Ahaiku answer is ~5× cheaper than opus; trivial asks never pay opus rates every request, sized/intent-routed
Response cache An identical prompt returns instantly forzero tokens every repeat (great in iterative sessions)
Effort --effort low on the same model costs less than dropping a tier, and often answers as well opt-in, per call or via config
Specification A request that had to be re-asked cost double; the pre-flight check is free and catches it first every request (advisory by default)
Noise filtering Trailing whitespace, blank-line runs, license headers, log dedup; onrefactor, comments/docstrings too every request (single-digit % typically; more on comment-heavy refactors)

Two measured points on the Claudio codebase itself (tiktoken BPE):

Command File size Input Sent Note
claudio ask -rv @repl.py "security" 963 lines 11,709 11,701 sent in full — faithful, ~0% shaved
claudio build -r @prompt.py "simplify" 174 lines 2,448 935 62% —refactor strips comments/docstrings

If a file is genuinely too large or costly to send whole, attach a narrower range yourself (@file -START-END) — an explicit call, not a summary claudio makes behind your back.

Use --verbose to see the pre-flight estimate and, once the call returns, what it actually billed:

[claudio] ~11,701 tokens (tiktoken BPE, est. $0.0585)
[claudio] model: opus
[claudio] billed: 36,412 in (19,189 cached) / 1,204 out | $0.1483 | claude-opus-5

The first line is claudio's pre-flight guess at the prompt it built; the last is the CLI's own accounting of the whole request. Expect them to differ — the gap is the system prompt, CLAUDE.md, tool definitions and cache traffic that no local estimate can see.


Permissions

Claudio drives Claude headless (claude --print), where there's no mid-run popup to approve an edit — the decision has to be made up front. A single permission posture controls how much Claude may do on its own when you build:

Posture Claude may… Under the hood
Autonomous edit filesand run shell commands --permission-mode bypassPermissions
Edits only (default) edit files, but not run shell commands --permission-mode acceptEdits
Confirm first edit files after you approve once per build acceptEdits + a Y/n gate
Preview only nothing — just describe the change --permission-mode plan

Because nothing pauses mid-run, Confirm first is a single coarse Y/n gate before a build starts, not a per-edit prompt.

Preview only uses Claude Code's plan mode. Earlier versions passed no mode at all and relied on headless --print auto-denying mutations — which worked, but only as a side effect: Claude still attempted edits and had them denied, so you paid for tool calls that could never land. In plan mode Claude plans instead of editing, and nothing is attempted in the first place.

Set it with the first-run wizard (auto-runs the first time you launch claudio), or anytime via /setup in the REPL or claudio setup:

How much can Claude do on its own?
  1) Autonomous    - apply edits AND run shell commands automatically
  2) Edits only    - apply edits automatically, never run commands  (default)
  3) Confirm first - ask once before each build applies, then apply
  4) Preview only  - never apply - just print the diff

The choice is saved as permission_posture in ~/.config/claudio/config.json. A CLAUDIO_BUILD_PERMISSION_MODE env var still overrides the resolved --permission-mode for one-off runs.

Read-only modes escalate instead of failing

ask, review, question, and debug are always read-only and ignore the posture entirely. But rather than silently attempting an edit the headless CLI would deny, Claude signals when a request actually needs build mode — and claudio offers to switch:

claudio [ask -review]> add a ROADMAP.md from your analysis
  ⚠ This needs build mode: writing a new file is a mutation, not a review.
  Re-run in build mode now? [Y/n]

On Y, claudio re-runs the request in build mode, resuming the same session so the analysis it just produced carries straight into the build.


Configuration

Claudio looks for config at ~/.config/claudio/config.json:

{
  "claude_binary": "claude",
  "default_model": "sonnet",
  "max_input_tokens": 32000,
  "output_format": "text",
  "verbose": false,
  "permission_posture": "edits",

  "effort": null,
  "max_budget_usd": null,
  "daily_budget_usd": null,
  "spec_check": true,
  "strict_spec": false,
  "partial_messages": true
}

All fields are optional. Defaults are used for anything not specified. permission_posture is one of autonomous, edits, confirm, or preview (see Permissions) — normally set by the wizard rather than by hand.

Key Meaning
effort Thinking depth: low / medium / high / xhigh / max; null = CLI default
max_budget_usd Hard ceiling for a single call; null = uncapped
daily_budget_usd Cumulative ceiling per day; null = uncapped (see Budgets)
spec_check Run the pre-flight specification check
strict_spec Treat specification findings as a refusal rather than a warning
partial_messages Request token-by-token streaming deltas

Every one of these also has an env override (CLAUDIO_EFFORT, CLAUDIO_DAILY_BUDGET_USD, CLAUDIO_NO_PARTIAL, …) so a one-off run never needs a config edit.


Piping and Composability

Results go to stdout, info/warnings to stderr:

# JSON output piped to jq
claudio --json ask -question @api.py "list the public functions" | jq '.result'

# Use in scripts
if claudio --dry-run --verbose build -refactor @src/ 2>&1 | grep -q "WARNING"; then
  echo "Input too large, consider narrowing scope"
fi

Project Structure

claudio/
  cli.py                 Command router (subcommand dispatch)
  repl.py                Unified `claudio` entry point (REPL + subcommand dispatch)
  cache.py               Response cache (SHA-256 keyed, TTL-based)
  usage.py               Cost and usage tracking
  config.py              Configuration management
  budget.py              Cumulative spend enforcement (daily cap + per-call clamp)
  spec_check.py          Pre-flight specification check (local, free, non-rewriting)
  executor.py            Claude CLI integration (streaming + retry + spinner)
  session_files.py       Per-session file-hash tracking for unchanged markers
  build_snapshot.py      Pre-build file snapshots for /undo rollback
  commands/
    build.py             claudio build (-refactor, -generate)
    ask.py               claudio ask (-review, -question, -debug)
    run.py               claudio run (claudio-task.json)
    run_prompt.py        Shared execution with cache + tracking + feedback channels
    stats.py             claudio stats (usage dashboard)
    setup.py             claudio setup (permission posture, PATH, completions, verification)
  completions/
    bash.py              Bash completion generator
    zsh.py               Zsh completion generator
    powershell.py        PowerShell completion generator
  pipeline/
    filter.py            Intent-aware noise filtering
    prompt.py            XML-tagged prompt construction
    process.py           Pipeline orchestrator (faithful — no compression)
  utils/
    args.py              @file parser with strict order enforcement
    project_context.py   Project preamble discovery (CLAUDE.md + .claudio/project.md)
    stack_detect.py      Stack detection from manifest files
    git_context.py       Auto git-diff context for review/debug/refactor
    tokens.py            Token estimation (tiktoken when available)
    output.py            Output formatting (with markdown rendering)
    files.py             File reading and ingestion
    colors.py            TTY-aware ANSI color helpers
    markdown.py          Markdown -> ANSI renderer (streaming + buffered)
    spinner.py           TTY-only stderr progress spinner
    update_check.py      Background PyPI version check
    model_router.py      Cheapest-model-that-fits routing
tests/                   pytest suite (336 tests across pipeline, REPL, context, executor)

Releases

Claudio uses tag-based releases. Each GitHub release auto-publishes to PyPI.

# Update to latest
pip install --upgrade claudio-cli

# Install a specific version
pip install claudio-cli==0.2.0

Metadata

Release files for claudio-cli 2.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for claudio-cli 2.0.0
File Size Uploaded
claudio_cli-2.0.0.tar.gz 183.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for claudio-cli 2.0.0
File Interpreter ABI Platform
claudio_cli-2.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 314.2 kB

Release files / claudio_cli-2.0.0.tar.gz

Download URL claudio_cli-2.0.0.tar.gz
Size 183.7 kB
Tags Source
SHA-256 checksum
How to use checksums
2b9e23d43d4015bbdabe652cee203574c0716c018e35a44ae992c9c0e11f0d7d
BLAKE2b-256 checksum
How to use checksums
7f90d2fa1295b3908bf08333ddc5423b4398461eeb622dbbacc2f73da6044826
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.

Transparency log

Release files / claudio_cli-2.0.0-py3-none-any.whl

Download URL claudio_cli-2.0.0-py3-none-any.whl
Size 130.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1807b64867f6a46bd943c204296bf6cd81aa4940260ec3be2930e06475a65dfe
BLAKE2b-256 checksum
How to use checksums
4431df22c4ce06b99a48377b8fe21dacece2b79b0421166aebf12473bae65c59
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

2.0.0 This release

2 release files

1.5.3

2 release files

1.5.2

2 release files

1.5.1

2 release files

1.5.0

2 release files

1.4.0

2 release files

1.3.1

2 release files

1.3.0

2 release files

1.2.0

2 release files

1.1.0

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page