PW Agent 🧠
CLI coding assistant powered by your Ollama GPUs via PastaWater.
Install
The recommended way to install pw-agent is using pipx to keep it isolated from your other Python packages:
pipx install pw-agent
Alternatively, you can use standard pip: pip install pw-agent
Usage
pw-agent
First run guides you through setup — paste your API token, pick a GPU, start chatting.
Features
- Interactive REPL with real-time streaming and a premium dashboard status bar.
- Plan vs Build Modes: Use
/planfor read-only analysis and/buildfor execution. - Context Discovery: Automatically finds
PW_AGENT.mdfor project-specific rules. - Tab Autocomplete for commands, file paths, and GPU slots.
- Session Control: Fresh sessions by default; use
-cto resume where you left off. - File Injection:
/add file.pyor@file.py— inject files into the LLM's context. - Batch Processing: Model can run multiple tool calls in a single turn.
- GPU Fleet Control:
/modelsto view GPUs and/use Nto switch connections or slots. - AI Commits:
/committo generate and apply git commit messages based on your diff. - Safety First:
-yflag for auto-approve; otherwise, every file edit requires confirmation.
Connect
- Cloud mode: Use your PastaWater API token to access your remote fleet.
- Direct mode: Point at a local Ollama instance (
--brain http://localhost:11434).
Get your token at pastawater.io/settings
Model compatibility (tool calling)
Agentic tool use needs both a capable model AND an Ollama whose tool-call parser tolerates that model's output drift.
| Model | Ollama | Tool calling | Notes |
|---|---|---|---|
qwen3-coder:30b |
>= 0.31.2, pw-agent >= 1.52.0 | verified | needs native tools mode (below); ~20 GB resident on a single 24 GB card |
qwen3-coder:30b |
>= 0.31.2, pw-agent <= 1.51.x | broken | every tool-requiring prompt dies on turn 0 with [Empty response from model], exit 3 |
qwen3-coder:30b |
0.21.x | broken | intermittent qwen tool call parsing failed: EOF — session degrades to plain chat |
llama3.1:8b |
any recent | works | weaker coder; fine for pipeline text tasks |
Native tools mode
Ollama >= 0.31 ships built-in renderer/parser pairs for some model families
(template selection ... selected=renderer_parser renderer=qwen3-coder). For
those models the server intercepts every <tool_call> tag the model emits
and parses it with that family's native grammar. pw-agent's textual protocol
puts JSON inside <tool_call>, which is not that grammar, so the server-side
parser dies with qwen tool call parsing failed: EOF, discards the whole
assistant message, and answers /api/chat with {"error":"EOF"}.
From 1.52.0 pw-agent sends Ollama's native tools schemas for these models,
drops the textual protocol from the system prompt, and reads structured
message.tool_calls back. Two safety nets:
- Any model that hits a server-side parse failure is flagged automatically and the turn is replayed with native tools — no failed run, just a slower first turn.
PW_NATIVE_TOOLS=1forces it on,PW_NATIVE_TOOLS=0forces it off (the off case now reports the parse failure as a named error rather than an empty response).
Side effect: the system prompt drops from ~4.7 KB to ~1.7 KB for these models, since the renderer injects the tool definitions itself.
Failure signals
A session that ends without executing any tool due to parse failure/stall
emits {"type":"result","subtype":"degraded","degraded_reason":..., "is_error":true} and exits with code 3; a missing model or dead endpoint
fails preflight with the installed-model list and exits with code 2.
Ollama-level errors (including tool-parse failures) are surfaced verbatim as
[Error: Ollama: ...] instead of an empty response.
Only one large model fits a 24 GB card at a time — requesting a second large tag while one is resident forces CPU offload.
Debugging a silent run
--debug (or PW_DEBUG=1) dumps the model name, native-tools decision,
num_ctx, every request message, the full assembled tool schema, and the raw
model completion. All of it goes to stderr, so --output-format stream-json on stdout stays machine-parseable:
pw-agent --instance 0 --yes --debug \
--output-format stream-json --print "..." 2>debug.log
--model is optional: omit it and pw-agent uses whatever the instance
currently has resident (brain's last_known_chat_model, else the loaded
model, else /api/tags). Pinning a tag the slot isn't serving is only
useful when you want Ollama to swap.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file pw_agent-1.52.0.tar.gz.
File metadata
- Download URL: pw_agent-1.52.0.tar.gz
- Upload date:
- Size: 128.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.12.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
578fda0b24ef6980a052b031be8065c7f18226a4477d808b18b9460e38011cf9
|
|
| MD5 |
26f14ce3f31705ecc73fb30f5061a422
|
|
| BLAKE2b-256 |
d68cd5dc00977fc9a81559401f1e76e2feb5506ccf4ee3fce46e9d34d1ccb3ba
|
File details
Details for the file pw_agent-1.52.0-py3-none-any.whl.
File metadata
- Download URL: pw_agent-1.52.0-py3-none-any.whl
- Upload date:
- Size: 134.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.12.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
df17c41929be520a47af87fdb5d24d942da16ec4d9dcd36619939141251ef5f9
|
|
| MD5 |
5e11c6ea009cda983fb31aecbc25ecc2
|
|
| BLAKE2b-256 |
8cf0086ecff48386c0b1dfe21b7a6a5020d574db2708eb73df368efa96da0294
|