whatsapp-agent-cli
A client for the WhatsApp Agent Platform — a library first, with a command-line tool on top.
Send and receive WhatsApp messages from a script, a cron job, a git hook, or a coding agent with shell access. It handles the parts that are tedious to get right — the long-poll and its cursor, per-method rate limits, the 4,096-character send cap, the two-hop media fetch, an error table where one code means "back off" and another means "this token is dead, stop", and voice notes turned into text by Gemini, OpenRouter, or fully offline on your own machine.
wa-agent send "deploy finished, 3 tests failing"
It knows nothing about coding agents, folders, permissions or models. It moves messages.
First product built on it: Hisab — a ledger that texts back. Double-entry bookkeeping for small businesses, run entirely from WhatsApp. More below.
Contents
- What it does differently
- Install
- Get a token
- Use it
- Transcription
- Use it from Python
- Where it keeps things
- When something fails
- Built on it: Hisab
- Coming next: the relay
- Contributing
- License
What it does differently
Sending a WhatsApp message is one HTTP call. What takes a real bot to get right is everything around it: the process that dies halfway through a batch, the platform that says slow down, the voice note that arrives when there is no key to transcribe it. Those behaviours are the point of this package. Each was learned running Hisab against real messages, and each is pinned by the repo's check.
Nothing is lost, nothing is stolen
| Behaviour | What it means for you |
|---|---|
| The cursor moves after the batch | It is saved only once a batch is delivered, and any message id already seen is skipped. A crash can repeat a message; it can never lose one |
| One poller per token | The platform allows a single long-poll and answers 409 when a second takes it. recv exits with another_poller rather than silently competing for your messages |
| A dead token stops the loop | A 401 exits 4 and is never retried. In recv --follow, a 429 or 503 backs off from 1s to 60s and keeps going. The two never look alike |
doctor cannot steal the poll |
Checking a setup with a real poll would take the long-poll from a running recv. doctor uses a read that cannot be a poll, so it is safe to run beside one |
Sending and media
| Behaviour | What it means for you |
|---|---|
| Long messages split cleanly | Cut on paragraph boundaries under WhatsApp's 4,096-character cap, numbered (i/n), with markdown converted to WhatsApp formatting |
| A half-failed send tells you what arrived | Every part that already went out is printed and recorded before a later one can fail, so a retry duplicates nothing |
| Rate limits are built in | Paced per method, with one retry past a 429. You do not write the backoff |
| Size caps are checked before the upload | 5 MB an image, 500 KB a sticker, 16 MB anything else, refused locally without spending a request |
--dry-run shows exactly what would go out |
Needs no token and sends nothing |
Voice notes, where the opinions are
| Behaviour | What it means for you |
|---|---|
| You choose the provider, it never guesses | gemini unless you say otherwise. A key for another provider is never spent because it was lying in your environment |
| Offline is opt-in, never a fallback | A default install has no offline engine. A missing key does not switch one on: recv warns once and delivers the note marked transcribed: false |
| Nothing downloads behind your back | A model comes only from wa-agent model pull, which says its size first. Never during recv |
| A failed transcription is never passed off as speech | The note is delivered marked, with the reason in transcription_error. An empty result or a provider's error message is an error, never words nobody said |
| A weak transcript says so | A locally transcribed note carries transcribed_by, so your code can weigh it. The docs say what the engine is bad at |
| The message keeps its shape | The words are added at text.body; type, audio and everything else are untouched, so code that already reads a message keeps working |
Failures, and your machine
| Behaviour | What it means for you |
|---|---|
| Every failure is a code | One fixed line, an exit status a script can branch on, and a verdict on whether retrying helps. wa-agent errors prints the table |
| The platform's words stay in one place | They appear only on a detail: line, and a token or key is never printed in an error or by doctor |
| Nothing lands in your working directory | State lives under XDG, per profile, and downloaded models beside it. Downloads are swept after a day |
| Secrets come from the environment | Or from ./.env, or a token file. No flag takes one, so none reaches your shell history or the process list. What came from ./.env is named on stderr, never shown |
| One dependency | requests. The heavy things are extras you ask for |
Install
pip install wa-agent
The package, the command and the module are all wa-agent / wa_agent. (This repository is named whatsapp-agent-cli; that name and whatsapp-agent both belong to unrelated projects on PyPI.)
That is the whole install: sending, receiving, media, and transcription through Gemini or OpenRouter, which needs only a key because it is an ordinary HTTP call. Transcribing offline is an extra you ask for. See Transcription.
Get a token
In WhatsApp: Settings → Agents → Create an agent → Chat info → API key. An agent may only message its own creator — you — which is why there is no recipient management here.
export WHATSAPP_AGENT_TOKEN='…' # or: wa-agent --token-file ~/.wa-token …
Or put it in a .env file in the directory you run wa-agent from, as WHATSAPP_AGENT_TOKEN=…. Every command reads ./.env, and only that one: not a parent directory, not your home directory. It only fills gaps, so a variable your shell already exports always wins. For the token, the order is the shell, then --token-file, then ./.env. When ./.env supplied a variable, or set one the shell overrode, the command says so on stderr, by name and never by value:
env: WHATSAPP_AGENT_TOKEN from ./.env; GEMINI_API_KEY from the shell (./.env also sets it, not used)
That line is how you notice a .env you did not mean to use. Run wa-agent inside another project that keeps the same variable in its .env, a Hisab checkout for one, and that project's token is the one used. wa-agent --no-env-file … skips the file entirely. The library never reads it: a program that imports wa_agent sees only its own environment.
Use it
# say something to yourself
wa-agent send "the backup finished"
# read what arrives, one JSON object per line, until you stop it
wa-agent recv --follow --json
# attach a file; the text becomes its caption
wa-agent send "this week's numbers" --file chart.png
# fetch something someone texted you, and open it
open "$(wa-agent media get <media-id>)"
The first recv records who you are, after which send needs no --to.
A voice note keeps its shape and gains the words, so code that reads text.body finds them and nothing about the message is lost:
{"id": "wamid.A", "type": "audio", "audio": {"id": "media-1", "voice": true},
"text": {"body": "call me back at six"}, "transcribed": true}
recv --transcribe takes the same --provider. Without a key for it, it warns once and delivers voice notes marked transcribed: false rather than stopping. More on transcription.
recv --download fetches each photo, document and voice note into the state directory as it arrives and adds a path to the message. It is opt-in because it puts a fetch inside the delivery loop; a download that fails is delivered marked with download_error, never dropped. With --transcribe as well, a voice note is fetched once, kept, and transcribed from that copy.
| Command | Does |
|---|---|
send <text> |
Send a message. Splits a long body on paragraph boundaries, numbers the parts (i/n), converts markdown to WhatsApp formatting |
send --file <path> |
Upload and attach. --media <id> attaches something already uploaded |
send --dry-run |
Print exactly what would be sent, send nothing, need no token |
recv |
Messages since the last run. --json for one object per line, --follow to stream, --typing to show a typing indicator while you work, --transcribe (with --provider) to add words to voice notes, --download to keep photos and files as they arrive |
transcribe <file> |
Audio in, text out. Gemini or OpenRouter, chosen with --provider, or --provider local to run offline (see Transcription) |
model pull [size] |
Download a Whisper model for --provider local: tiny, base (the default) or small. Says the size first. It is the only thing that ever downloads one |
media get <id> |
Download to the state directory, or --out DIR. Prints the path and nothing else |
media put <path> |
Upload, print the media id |
doctor |
Check a setup, a line each: Python, token, each transcription key, the local engine, state directory, creator. Says what to fix, exits 15 if anything fails. It never polls, so it is safe beside a running recv, but it does make real, free metadata requests to the platform and to every provider whose key is set in your environment |
errors |
The exit-code table |
Global options — --token-file, --state-dir, --profile, --no-env-file — go before the subcommand, as in git:
wa-agent --profile work recv --follow # yes
wa-agent recv --follow --profile work # no: unrecognized argument
Transcription
Three providers, and one rule: you choose, it never guesses.
| Gemini | OpenRouter | Local | |
|---|---|---|---|
| Choose with | nothing (the default), or --provider gemini |
--provider openrouter |
--provider local |
| Needs | GEMINI_API_KEY |
OPENROUTER_API_KEY |
no key |
| Install | nothing extra | nothing extra | pip install "wa-agent[local]", then wa-agent model pull |
| Cost and privacy | billed by the provider; the audio goes to them | billed by the provider; the audio goes to them | free; nothing leaves your machine |
Marked in recv --json |
nothing added | nothing added | "transcribed_by": "local:base" |
export GEMINI_API_KEY='…'
wa-agent transcribe voice-note.ogg
export OPENROUTER_API_KEY='…'
wa-agent transcribe voice-note.ogg --provider openrouter # --model takes an OpenRouter model id
Offline, on your own machine
Nothing leaves the machine and nothing is billed. It is an extra because it is heavy (about 150 MB of dependencies) and a model is a separate download, so it takes three deliberate steps:
pip install "wa-agent[local]" # 1. the engine
wa-agent model pull # 2. a model: says the size first; base is about 150 MB
wa-agent transcribe voice-note.ogg --provider local # 3. ask for it: tiny, base or small; --model small
It is never chosen for you, not even when a key is missing: --provider local is a decision. Nothing downloads during recv. Models are kept in $XDG_DATA_HOME/wa-agent/models/, shared by every profile.
When there is no key
Nothing is imposed on you, and nothing is used in its place:
| You run | With no key and no local engine |
|---|---|
wa-agent transcribe note.ogg |
Exits 12 (no_transcription_key), and the detail: line names the variable to set |
wa-agent recv --transcribe |
Warns once, up front, then delivers each voice note marked transcribed: false and keeps going |
wa-agent recv --transcribe --provider local before model pull |
The same: one warning that names the fix, notes delivered marked, and no download |
What the local engine is bad at
- Fine: clear, accented English.
- Poor: Urdu, Roman Urdu, and speech that switches between languages. A small Whisper model tends to write fluent, confident English that was never said, rather than failing, so a wrong transcript looks exactly like a right one.
- Slower than an API call, and
recvwaits while it works. - So it is tagged:
recv --jsonadds"transcribed_by": "local:base"to a locally transcribed voice note, and adds nothing to a Gemini or OpenRouter one, so whatever reads a message can weigh it differently. - Three sizes only:
tiny,baseandsmall. There is no English-only (.en) model, because it would turn Urdu into English even more confidently.
Use it from Python
from wa_agent import WhatsApp, Store, WhatsAppError
client = WhatsApp(token)
for sent in client.send_iter("user:123", "**done** in 40s"):
print(sent.id, sent.text) # one per part, as each leaves
messages, cursor = client.poll(offset=None)
for message in messages:
print(message["from"], message.get("text", {}).get("body"))
send_iter yields each part as it is delivered, so a failure halfway never hides what already arrived. Store is the message log and the cursor, keyed by the platform's own message ids.
Where it keeps things
Nothing is written into your working directory. State lives at $XDG_STATE_HOME/wa-agent/<profile>/ (or ~/.local/state/…), holding the poll cursor, the message log and downloaded media. --state-dir moves it; --profile keeps two agents apart. Downloaded transcription models are kept apart from it, in $XDG_DATA_HOME/wa-agent/models/ (or ~/.local/share/…), because they are large and the same for every profile.
One poller per token. The platform allows a single long-poll per agent and answers 409 when a second one takes the cursor, so recv exits rather than silently competing for your messages.
When something fails
Every failure names its code and exits with a number a script can branch on:
error [platform_rejected]: the platform refused this request; retrying will not help
detail: POST /messages: HTTP 400 error.code 131009 …
Not sure where a setup stands? wa-agent doctor checks it in one go and says what to fix.
wa-agent errors lists them all. docs/errors.md says what to do about each and which are worth retrying — the short version is that exit 7 is, and 4 and 6 never are.
Built on it: Hisab
Hisab is a plain-language ledger you keep by texting WhatsApp — "2500 coffee" posts an entry, "how much do I owe Metro?" gets an answer — in English, Urdu or Roman Urdu, by voice, photo or text, with every entry checked by hledger before it is written.
It is where this package came from. The cursor that only advances after a batch, the dedup, the per-method rate limits and the dead-token exit were all learned running Hisab against real messages, then extracted here so nothing else has to learn them again. Hisab is the first product on this transport, and moves onto the published wa-agent package next.
The two repositories split the work cleanly:
| whatsapp-agent-cli (this) | hisab-whatsapp | |
|---|---|---|
| Is | the transport: messages, media, transcription | a product: a ledger with a model and six tools |
| Knows about | tokens, cursors, rate limits | accounts, entries, hledger |
| You use it | from a script, a cron job, or your own agent | by texting it |
Coming next: the relay
The transport moves messages. The relay is what makes it an agent in your pocket.
wa-agent relay --folder ~/code/my-project # coming soon
Text it from your phone — "why is the deploy failing?", a screenshot of an error, a voice note describing a bug — and it runs a coding agent such as Claude Code over that folder and sends back what it says. The design is settled; the code starts once this release is out:
- Read-only by default. The agent can read the folder and nothing else. Writable paths are declared, never assumed, and a folder created later is denied until you say otherwise.
- The relay owns the session. Starting fresh, switching models and compacting a long conversation happen in the relay, before the agent is called, because none of them survive a non-interactive run otherwise.
- Everything it hears, it can use. Voice notes arrive as words, photos and files arrive by path, and a quoted reply arrives with the message it quoted — all from this package, underneath.
- One command in this package, optional.
pip install wa-agentnever makes you run it. The transport stays usable on its own, and the relay uses it exactly as your own scripts would.
Claude Code comes first, Codex after. Follow along on the issues.
Contributing
cp .env.example .env # fill in your agent token, and a transcription key if you want it
make dev # a virtualenv with this checkout installed
make check # the check: no network, no token, a couple of seconds
make up # a live inbox: text your agent and watch it land, until Ctrl-C
make live # a scripted round trip: send, wait for your reply, read it back
make up and make live keep their state in .live-state/, never your real one, and make clean removes it. develop is the trunk and PRs target it; main is what is published. Conventions live in AGENTS.md.
License
MIT.
Release files for wa-agent 0.4.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| wa_agent-0.4.0.tar.gz | 78.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| wa_agent-0.4.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 126.4 kB
Release files / wa_agent-0.4.0.tar.gz
| Download URL | wa_agent-0.4.0.tar.gz |
|---|---|
| Size | 78.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
896739b14e03719e5898a92f33079088ac6023bb090a12a742b99f4c87526fc3
|
|
BLAKE2b-256 checksum How to use checksums |
fc6520a7c75b73aa95235696aff243b1f835cca1dc87216d75e20e9f43c479b1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 19, 2026.
Transparency logRelease files / wa_agent-0.4.0-py3-none-any.whl
| Download URL | wa_agent-0.4.0-py3-none-any.whl |
|---|---|
| Size | 47.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
39ae6016242646ef6ef653f2d8a9c584b42b7853357ea25e6922028341603e77
|
|
BLAKE2b-256 checksum How to use checksums |
0369f5d50807c3675b3d0ace12272987ce5c48e7f50c89e77c564555aab9421d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 19, 2026.
Transparency log