Ratchet
A small agentic coding harness for Claude that you can read in one sitting.
Ratchet runs in your terminal. You give it a task, and Claude reads code, searches the project, edits files and runs commands in a loop until the task is done. Edits and commands wait for your approval unless you say otherwise.
The name describes the design: the conversation only moves forward. History is append-only. Ratchet never rewrites, prunes or reorders what it has sent. Two things depend on that: the prompt cache, which only matches an unchanged prefix, and Claude's thinking blocks, which stay valid only while the conversation before them is unchanged.
Install
Needs Python 3.10+ on Linux or macOS. ~/.local/bin must be on your PATH
(it is on most distributions). Uninstall with uv tool uninstall ratchet-harness.
uv tool install ratchet-harness # or: pipx install ratchet-harness; puts `ratchet` on your PATH
ratchet --set-key anthropic # or: export ANTHROPIC_API_KEY=sk-ant-...
Working on Ratchet itself? From a clone, uv tool install --editable . makes
source edits apply immediately.
Saving keys
Instead of exporting keys in every shell, save them once:
ratchet --set-key kimi # prompts without echoing; also anthropic, openrouter, glm, custom
pass show zai | ratchet --set-key glm # or pipe it in from a password manager
Inside a session, /key [provider] does the same and takes effect immediately.
Keys go to ~/.config/ratchet/keys.env (or $XDG_CONFIG_HOME/ratchet/keys.env),
in a directory only you can read. It's a plain NAME=value file, so you can
also edit it by hand, and any setting Ratchet reads from the environment works
there too: RATCHET_PROVIDER=glm, ZAI_BASE_URL=..., RATCHET_EFFORT=medium.
Variables exported in your shell take precedence over the file.
Using OpenRouter
export OPENROUTER_API_KEY=sk-or-...
ratchet --provider openrouter # anthropic/claude-opus-5.5 by default
ratchet --provider openrouter --model anthropic/claude-sonnet-5.5
export RATCHET_PROVIDER=openrouter # make it the default
Ratchet sends OpenRouter the same Anthropic Messages requests it sends Anthropic,
through OpenRouter's Anthropic-compatible endpoint (https://openrouter.ai/api,
or $OPENROUTER_BASE_URL), using the official SDK. Model names use OpenRouter's
slugs (anthropic/claude-opus-5.5). Differences from going direct:
- Auth: your key goes out as
Authorization: Bearer. Even ifANTHROPIC_API_KEYis set, it is never sent to OpenRouter. - Sticky routing: every request carries the session ID as OpenRouter's
session_id, so a session stays with one upstream provider. Switching providers mid-session would lose prompt-cache hits and could invalidate thinking blocks. - Tool input streaming:
eager_input_streamingis off, because OpenRouter's schema doesn't include it. Tool inputs arrive whole instead of streaming in. - Refusal fallback: Anthropic's
fallbacks: "default"isn't sent, because OpenRouter'sfallbacksfield means something different. - Unchanged: compaction, caching, thinking and effort work as they do direct.
Sessions remember their provider, so ratchet --resume picks it up without the
flag. Other Claude models work the same way.
Bring your own key: Kimi, GLM, or any Anthropic-compatible API
Moonshot (Kimi) and Z.ai (GLM) both serve the Anthropic Messages API, so Ratchet talks to them directly with your own key, without going through OpenRouter.
export MOONSHOT_API_KEY=sk-... # or KIMI_API_KEY
ratchet --provider kimi # kimi-k2.5 by default
ratchet --provider kimi --model kimi-k2-thinking
export ZAI_API_KEY=... # or ZHIPUAI_API_KEY / GLM_API_KEY
ratchet --provider glm # glm-4.7 by default
ratchet --provider glm --model glm-4.6
| Provider | Key | Default endpoint | Override with |
|---|---|---|---|
kimi (alias moonshot) |
$MOONSHOT_API_KEY |
https://api.moonshot.ai/anthropic |
$MOONSHOT_BASE_URL (China: https://api.moonshot.cn/anthropic) |
glm (aliases zai, zhipu) |
$ZAI_API_KEY |
https://api.z.ai/api/anthropic |
$ZAI_BASE_URL (China: https://open.bigmodel.cn/api/anthropic) |
custom |
$RATCHET_CUSTOM_API_KEY |
none: set $RATCHET_CUSTOM_BASE_URL |
model from $RATCHET_CUSTOM_MODEL or --model |
custom takes any other Anthropic-compatible endpoint (DeepSeek, MiniMax, a
self-hosted gateway, and so on). Give the API root: the SDK appends
/v1/messages. The key goes out as Authorization: Bearer, and
ANTHROPIC_API_KEY is never sent to these providers. Use /provider to switch
in a session (it starts a new conversation), and /model to pick from that
provider's models or type any ID it serves.
Ratchet only sends Claude-only features to Claude models:
- Thinking: non-Claude models get classic
thinking: {"type": "enabled", "budget_tokens": N}.--effortpicks the budget: 2k (low), 8k (medium), 16k (high), 32k (xhigh), or everything below--max-tokens(max). Adaptive thinking andoutput_config.effortgo to Claude only. - Not sent: automatic conversation caching (these providers cache prefixes
on their own), refusal fallback,
eager_input_streaming, and compaction. Long sessions on these models therefore aren't compacted. Watch/context. - Cost: Ratchet only knows Claude prices, so the status line shows
cost n/a. Check your provider's console.
Use
ratchet # interactive session in the current directory
ratchet "fix the failing test in tests/test_api.py" # start with a task, stay interactive
ratchet -p "add type hints to utils.py" # run one task to completion, then exit
ratchet -p -y "run the tests and fix what fails" # same, without approval prompts
ratchet --mode plan "explain how auth works" # can look, can't touch
ratchet --resume # pick up the latest session for this directory
The interface
In a terminal:
- Streaming output. Claude's answers stream in as rendered markdown, each
reply marked with
◆, above a status line that says what's happening (◜ Writing 12s · 1.2k tokens out · esc to interrupt). - Thinking. Summaries show dimmed behind a
┆gutter;/thinkinghides them. - Tool calls. Each one shows as
▸ Read src/app.py, with its result in a│gutter underneath (✗marks a failed call). Edits show a red/green diff with the file's line numbers, and shell commands show the first lines of their output. - Approvals. A
? Edit fileor? Run commandheader shows the diff or command, then a menu:yallow once /aallow that kind for the session /ndeny and say what to do instead. Your note goes to Claude with the denial. - Status line. Under the input it shows the mode, model, effort, how full the context window is, and the session cost.
| Key | |
|---|---|
/ |
commands, with completion as you type |
!command |
run a shell command yourself; its output goes along with your next message |
@path |
mention a file (Tab completes paths) |
? |
list shortcuts |
| Shift+Tab | cycle mode: default → accept edits → plan → auto |
| Esc | interrupt Claude (Ctrl-C works too) |
| Ctrl-C | clear the input; twice on an empty line to exit |
\ + Enter, Esc then Enter, Ctrl-J |
new line |
| Command | |
|---|---|
/model [name] |
switch model: a picker, or a name like sonnet, opus 5, fable, or a full ID |
/effort [level] |
low · medium · high · xhigh · max |
/mode [mode] |
default · accept-edits · plan · auto |
/provider [name] |
anthropic, openrouter, kimi, glm or custom (starts a new conversation) |
/key [provider] |
save a provider's API key to ~/.config/ratchet/keys.env |
/thinking [on|off] |
show or hide thinking summaries |
/status |
provider, model, effort, mode, session, and what's always allowed |
/usage, /cost |
tokens and estimated cost |
/context |
how full the context window is |
/clear, /new |
clear the screen and start a new conversation |
/resume [id] |
pick up an earlier conversation in this workspace |
/init |
have Claude write an AGENTS.md for the project |
/help, /exit |
Model and effort changes apply to the next request and keep the conversation. Switching models means the prompt cache starts fresh.
Modes.
defaultasks before every edit and command.accept-editsedits files without asking, but still asks before commands.planis read-only. Claude explores and proposes a plan, and nothing changes until you leave plan mode. Entering and leaving plan mode is announced to Claude in your next message, so the conversation stays append-only.autoruns everything without asking. Use it only where mistakes are cheap.
When output isn't a terminal (pipes, -p), Ratchet falls back to plain text.
| Option | Default | |
|---|---|---|
--provider |
anthropic (or $RATCHET_PROVIDER) |
anthropic, openrouter, kimi, glm or custom |
--model |
claude-opus-5-5 (or $RATCHET_MODEL); anthropic/claude-opus-5.5 on OpenRouter, kimi-k2.5 on Kimi, glm-4.7 on GLM |
Claude: any model with adaptive thinking (Opus 4.6+, Sonnet 4.6+, Fable). Other providers: any model ID they serve |
--effort |
high (or $RATCHET_EFFORT) |
low · medium · high · xhigh · max |
--mode |
default |
-y = --mode auto, --read-only = --mode plan |
--max-tokens |
64000 | Output cap per response, thinking included |
--max-steps |
200 | Model calls allowed per message you send |
--no-fallback |
on | Server-side refusal fallback (see below) |
--no-compaction |
on | Server-side context compaction (see below) |
--hide-thinking |
shown | Hide Claude's thinking summaries |
--root |
. |
Workspace root |
Ratchet adds RATCHET.md, AGENTS.md or CLAUDE.md from the workspace root
to the system prompt, if any exist.
Tools
| Tool | Needs approval | Notes |
|---|---|---|
read_file |
no | Line-numbered, pages through large files |
glob |
no | Newest first, skips node_modules, .venv, .git and similar |
grep |
no | Python regex, respects .gitignore |
edit_file |
yes | Exact, unique string replacement; keeps CRLF files CRLF |
write_file |
yes | Creates or overwrites whole files |
bash |
yes | Runs in the workspace root, with a timeout that kills the whole process group |
skill |
no | Loads one of Ratchet's built-in skills (see below) |
File tools can't reach outside the workspace root (symlinks included).
edit_file and write_file refuse to change a file Claude hasn't read in this
session, or one that changed on disk since, so Claude can't overwrite work it
hasn't seen. bash is not confined, so it asks first unless you pass -y.
Only use -y where a bad command can't do real damage.
When Claude asks for several read-only tools at once they run in parallel. Anything that changes state runs one at a time, in order.
Skills
Skills are guides for particular kinds of work. Each one is a
src/ratchet/skills/<name>/SKILL.md file, with name and description in
frontmatter. The system prompt lists only names and descriptions. When a task
matches, Claude loads the full text with the skill tool, so it arrives as a
tool result and the frozen prompt stays small. To add a skill, add a directory
with its own SKILL.md.
| Skill | |
|---|---|
frontend-design |
UI work: layout, type, color, states, motion, accessibility and a review checklist, drawn from Apple's HIG, Material Design and WCAG |
How the loop works
All of it is in src/ratchet/agent.py.
- Streaming and thinking. Requests stream, with adaptive thinking and
display: "summarized", so you see Claude's reasoning summaries and progress notes as they arrive rather than a long silence. - Caching. The system prompt and the tool list are fixed when the session
starts and carry a cache breakpoint. Top-level automatic caching covers the
growing conversation.
/usageshows the cache reads. - Stop reasons.
tool_useruns the tools and continues.end_turnwaits for you.max_tokensnever runs a tool whose input may have been cut off; Claude gets an error result and can retry in smaller pieces.refusalruns nothing and discards the partial turn.pause_turnresumes. - Tool input validation. Tool inputs stream as Claude writes them
(
eager_input_streaming), so the API doesn't validate them. Ratchet checks every input against the tool's schema before running it, and re-sends the request if the SDK couldn't parse the JSON at all. - No dangling calls. Every
tool_usegets atool_resultin the same append, even if you interrupt or deny it. If Claude never answered a message, Ratchet removes it instead, so a retry doesn't send it twice. - Refusal fallback. On models that support it, requests carry
fallbacks: "default". If a safety classifier declines a request, the API retries it on a fallback model chosen for that refusal category. If a model switches mid-answer, Ratchet drops the declined part's tool calls and thinking before sending the turn back, and never runs those calls. - Compaction. Long sessions use server-side compaction (
compact_20260112). The API summarizes older context, and Ratchet keeps the compaction blocks in history as the API requires. - Sessions are saved after every step to
~/.local/share/ratchet/sessions/, outside your repo, together with the frozen system prompt, so a resumed session sends exactly the same prefix.
Development
uv venv && uv pip install -e ".[dev]"
.venv/bin/python -m pytest
The tests run the real anthropic client against a scripted Messages API over
a mock HTTP transport that serves real SSE. They cover the SDK's streaming and
accumulation code as well as Ratchet's, and they check the exact request
bodies. The checks include byte-identical prefixes across turns, thinking
signatures echoed unchanged, and a valid tool_use/tool_result pairing
after every interrupt, denial, truncation, refusal and fallback.
Layout
src/ratchet/
agent.py the loop: streaming, stop reasons, tool execution, permissions
tools.py tool schemas, input validation, implementations
prompt.py system prompt (built once per session)
skills.py skill discovery; skills/<name>/SKILL.md holds the guides
session.py save / resume
keys.py saved API keys and settings (~/.config/ratchet/keys.env)
models.py providers, model catalogs, provider model IDs, permission modes
tui.py interactive output: streaming markdown, tool lines, diffs, approval menus
repl.py interactive input: slash commands, completion, status line, shortcuts
ui.py plain-text output for -p and pipes
cli.py arguments, provider clients, one-shot mode
Metadata
Release files for ratchet-harness 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ratchet_harness-0.1.0.tar.gz | 86.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| ratchet_harness-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 155.9 kB
Release files / ratchet_harness-0.1.0.tar.gz
| Download URL | ratchet_harness-0.1.0.tar.gz |
|---|---|
| Size | 86.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
3730e00d3281bac635f1d988fbb99a77f2fd06ccde5e25d693ffa5bba9bf80f8
|
|
BLAKE2b-256 checksum How to use checksums |
861fadba6bed4696fa2c1c60757eff84594e2d129f0be5f3bdf79a1b2f29a069
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"22.04","id":"jammy","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|
Release files / ratchet_harness-0.1.0-py3-none-any.whl
| Download URL | ratchet_harness-0.1.0-py3-none-any.whl |
|---|---|
| Size | 69.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
914d2035676668732586d9a5dbc3bc18202bd3f33eebcb83dd56efcb3b773636
|
|
BLAKE2b-256 checksum How to use checksums |
ded5aa5bca40877d4f325f6b5f0f4b495652e8bf6e9ae172ff36102aaacd307f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"22.04","id":"jammy","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|