PW Agent 🧠
CLI coding assistant powered by your Ollama GPUs via PastaWater.
Install
The recommended way to install pw-agent is using pipx to keep it isolated from your other Python packages:
pipx install pw-agent
Alternatively, you can use standard pip: pip install pw-agent
Usage
pw-agent
First run guides you through setup — paste your API token, pick a GPU, start chatting.
Features
- Interactive REPL with real-time streaming and a premium dashboard status bar.
- Plan vs Build Modes: Use
/planfor read-only analysis and/buildfor execution. - Context Discovery: Automatically finds
PW_AGENT.mdfor project-specific rules. - Tab Autocomplete for commands, file paths, and GPU slots.
- Session Control: Fresh sessions by default; use
-cto resume where you left off. - File Injection:
/add file.pyor@file.py— inject files into the LLM's context. - Batch Processing: Model can run multiple tool calls in a single turn.
- GPU Fleet Control:
/modelsto view GPUs and/use Nto switch connections or slots. - AI Commits:
/committo generate and apply git commit messages based on your diff. - Safety First:
-yflag for auto-approve; otherwise, every file edit requires confirmation.
Connect
- Cloud mode: Use your PastaWater API token to access your remote fleet.
- Direct mode: Point at a local Ollama instance (
--brain http://localhost:11434).
Get your token at pastawater.io/settings
Model compatibility (tool calling)
Agentic tool use needs both a capable model AND an Ollama whose tool-call parser tolerates that model's output drift.
| Model | Ollama | Tool calling | Notes |
|---|---|---|---|
qwen3-coder:30b |
>= 0.31.2, pw-agent >= 1.52.0 | verified | needs native tools mode (below); ~20 GB resident on a single 24 GB card |
qwen3-coder:30b |
>= 0.31.2, pw-agent <= 1.51.x | broken | every tool-requiring prompt dies on turn 0 with [Empty response from model], exit 3 |
qwen3-coder:30b |
0.21.x | broken | intermittent qwen tool call parsing failed: EOF — session degrades to plain chat |
llama3.1:8b |
any recent | works | weaker coder; fine for pipeline text tasks |
laguna-xs-2.1 |
>= 0.32.13, pw-agent >= 1.54.0 | verified | native tools; write+read acceptance passed on a 4090 in 24s (turns:2, exit 0), SWE-bench 70.9 |
muse-glimmer:30b |
>= 0.32.8 (CUDA), pw-agent >= 1.54.0 | verified | native tools; write+read acceptance passed on a 4090 in 35s cold (turns:2, exit 0), 16.6GB resident. Multimodal (image input) |
Native tools mode
Ollama >= 0.31 ships built-in renderer/parser pairs for some model families
(template selection ... selected=renderer_parser renderer=qwen3-coder). For
those models the server intercepts every <tool_call> tag the model emits
and parses it with that family's native grammar. pw-agent's textual protocol
puts JSON inside <tool_call>, which is not that grammar, so the server-side
parser dies with qwen tool call parsing failed: EOF, discards the whole
assistant message, and answers /api/chat with {"error":"EOF"}.
From 1.52.0 pw-agent sends Ollama's native tools schemas for these models,
drops the textual protocol from the system prompt, and reads structured
message.tool_calls back. Two safety nets:
- Any model that hits a server-side parse failure is flagged automatically and the turn is replayed with native tools — no failed run, just a slower first turn.
PW_NATIVE_TOOLS=1forces it on,PW_NATIVE_TOOLS=0forces it off (the off case now reports the parse failure as a named error rather than an empty response).
Side effect: the system prompt drops from ~4.7 KB to ~1.7 KB for these models, since the renderer injects the tool definitions itself.
Failure signals
A session that ends without executing any tool due to parse failure/stall
emits {"type":"result","subtype":"degraded","degraded_reason":..., "is_error":true} and exits with code 3; a missing model or dead endpoint
fails preflight with the installed-model list and exits with code 2.
Ollama-level errors (including tool-parse failures) are surfaced verbatim as
[Error: Ollama: ...] instead of an empty response.
Only one large model fits a 24 GB card at a time — requesting a second large tag while one is resident forces CPU offload.
Debugging a silent run
--debug (or PW_DEBUG=1) dumps the model name, native-tools decision,
num_ctx, every request message, the full assembled tool schema, and the raw
model completion. All of it goes to stderr, so --output-format stream-json on stdout stays machine-parseable:
pw-agent --instance 0 --yes --debug \
--output-format stream-json --print "..." 2>debug.log
Choosing the model
pw-agent is not read-only about models — the dashboard is the default, not a lock.
| Behaviour | |
|---|---|
--model <tag> |
Overrides everything. Ollama loads that tag on the first request, evicting whatever is resident. |
--model omitted |
Auto-detect: brain's last_known_chat_model (the dashboard pick, trusted over VRAM state) → resident chat model → /api/tags. |
/model in the REPL |
Alias for /models — displays the fleet, switches nothing. There is no in-session switch; restart with --model. |
The brain proxies /api/chat straight through with no model filtering, so
nothing server-side rejects your choice. Two constraints:
- The tag must already be pulled on that box. pw-agent never downloads —
preflight fails with
Model 'X' not found. Available: ...and exit 2. - On a 24 GB card a forced swap costs a cold load (~90 s for a 30B) and evicts the dashboard's model. If the dashboard swaps back you'll fight over the card.
Metadata
Release files for pw-agent 1.59.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pw_agent-1.59.0.tar.gz | 139.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pw_agent-1.59.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 280.9 kB
Release files / pw_agent-1.59.0.tar.gz
| Download URL | pw_agent-1.59.0.tar.gz |
|---|---|
| Size | 139.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7d66a6d038f12d3b486c58913a9536801e6dffb88756fcb069ab88f853ed80e4
|
|
BLAKE2b-256 checksum How to use checksums |
1319bdbba5168eddab6a56641c5fc9b2c017f1813caf95c1740f0be7d80c5719
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.13
|
Release files / pw_agent-1.59.0-py3-none-any.whl
| Download URL | pw_agent-1.59.0-py3-none-any.whl |
|---|---|
| Size | 141.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
99e7aebd4a45066a40e5f51d23f1ea44d485626e1bc093da85e18130968cd6c9
|
|
BLAKE2b-256 checksum How to use checksums |
a9d00649778e2ef6bc9c8d186bef53495537c8b3bb1287758b4c079eab20cb15
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.13
|