Skip to main content

lindwyrm

tests pypi python license

A small, dependency-light coding agent for the terminal. It talks to model APIs directly over HTTP — DeepSeek out of the box, plus any other Anthropic- or OpenAI-compatible provider you add as a preset: an aggregator, Kimi, a local vLLM or Ollama server, whatever you use.

Built because the off-the-shelf agents carry telemetry and a long supply chain of packages, some of which have a habit of reaching for API keys. This is one package with two runtime dependencies — httpx for the wire and rich for rendering — and no SDKs. You can read all of it in an afternoon.

$ lwyrm
lindwyrm - coding agent
  preset:   deepseek-flash (anthropic, deepseek-flash)
  thinking: on
  root:     ~/work/shop
  perms:    read=allow write=confirm delete=confirm bash=confirm
  /help for commands, /exit to quit

you › rewrite total() in prices.py as a one-line sum() expression

→ read_file prices.py
  ok 1  def total(items):

→ edit_file prices.py

  WRITE requested: edit prices.py
    prices.py: +1 -4
    @@ -1,5 +1,2 @@
     def total(items):
    -    result = 0
    -    for item in items:
    -        result = result + item["price"] * item["qty"]
    -    return result
    +    return sum(item["price"] * item["qty"] for item in items)
  Allow? [y]es / [n]o / [a]lways (write) / [q]uit: y
  ok Edited prices.py

Rewrote total() as a one-line sum() expression, keeping identical behaviour.

Install

pip install lindwyrm        # installs the `lwyrm` command (`lindwyrm` also works)
export DEEPSEEK_API_KEY=sk-...
lwyrm                       # start in the current directory

pip install 'lindwyrm[socks]' adds SOCKS proxy support. pip install lwyrm works too — an alias package that pulls in the same thing.

Requires Python 3.11+ on Linux or macOS. Windows works under WSL; a native Windows shell won't, because the bash tool relies on POSIX process groups to kill a command's whole process tree.

Use

lwyrm                    # interactive REPL in the current directory
lwyrm -m pro             # use the deepseek-pro preset
lwyrm -m kimi            # use a custom preset from your config
lwyrm -C ~/myproject     # set the project root
lwyrm -p "add type hints to utils.py"   # one-shot, then exit
lwyrm --continue         # pick up the last session in this project
lwyrm --resume <id>      # pick up a specific one (see /sessions)
lwyrm --read-only        # answer only; never write or run
lwyrm --no-thinking      # turn reasoning off
lwyrm --no-save          # leave no trace on disk
lwyrm --init             # write a starter AGENTS.md and exit (no API key needed)

lwyrm --help lists the rest.

In the REPL: /model /presets /thinking /think /markdown /perm /policy /context /compact /sessions /init /proxy /clear /help /exit. /help explains each one.

Ctrl+C stops whatever is running and clears the line you were typing, the way a shell does. Press it twice, or use Ctrl+D or /exit, to leave. Stopping a turn — Ctrl+C, or q at a prompt — keeps the session usable: the tool calls it cut short are recorded as interrupted, and the next message carries on from there.

With -p the exit status says how the turn ended: 0 finished, 1 an API or other error, 3 incomplete (the step ceiling, or a reply cut off at max_tokens), 130 interrupted.

Contents

Tools · Permissions · Project instructions · Context and cost · Sessions · Proxies · Configuration · Limits · Tests · Releasing

Tools

read_file, write_file, edit_file, list_dir, glob, grep, delete_file, bash, read_offloaded, view_image — identical no matter which provider is active.

  • bash streams output line by line as the command runs, with stdout and stderr interleaved in the order they actually happened, and kills the whole process tree on timeout or Ctrl+C. stdin is closed, so a command that prompts gets end-of-file at once instead of waiting out the timeout.
  • edit_file keeps a file's line endings: a CRLF file stays CRLF.
  • edit_file requires a unique match by default; when the text appears several times the error names the lines that matched, so the next attempt has something to go on. replace_all changes every occurrence.
  • grep reads files line by line and skips binaries and vendored directories. list_dir shows dotfiles and marks symlinks with their target.
  • Every write and every overwrite is previewed as a real unified diff.
  • view_image shows the model a screenshot, diagram or chart. It is only offered to models that can see, so the rest don't waste a turn calling it; set vision = true on a preset whose model supports images. PNG, JPEG, GIF and WebP, detected from the file's contents rather than its name, because that is how the provider decides too.

Reasoning is streamed where the provider supports it. DeepSeek's Anthropic endpoint returns 400 unless thinking blocks are echoed back in history, so assistant content is stored and replayed verbatim. OpenAI-format providers don't require that, and some reject unknown fields, so there reasoning is shown live but not sent back.

Permissions

Every read, write and delete resolves to allow (silent), confirm (ask first) or deny (blocked), from a global default plus per-path rules. The most specific rule wins, so you can open a folder broadly and still lock a subfolder inside it. Defaults: read=allow, write=confirm, delete=confirm.

Reads outside the project are held to at least read_outside (default confirm), so ~/.ssh or your key file can't be read and sent to the provider without you seeing it. A path rule still wins — [[policy.rules]] with path = "~/docs" and read = "allow" opens that folder — and read_outside = "allow" restores the old behaviour. Answering [a]lways covers paths inside the project only.

../ and symlinks are resolved before matching, so a link can't dodge a rule — and the confirmation prompt shows the real destination, not the path that was typed.

bash sits outside this system, because a shell command isn't bound to a path: it has one global level (default confirm), an always-on denylist and an optional allowlist for harmless commands. An allowlist entry matches the command itself plus plain arguments — ls allows ls -la src, but anything containing ;, &, |, <, >, $, backticks or parentheses is asked about, since ls && rm -rf ~ starts with ls too. Confirmation prompts show control characters as escapes, so a command can't repaint itself into something harmless-looking. Set rules in config under [[policy.rules]], or during a session with /perm src write=allow delete=confirm.

There is also a --read-only mode and a JSON-lines audit log.

Project instructions

AGENTS.md in the project root is read at the start of every session and prepended to the system prompt, so the agent begins knowing your build commands and conventions instead of rediscovering them. /init or lwyrm --init writes a commented starter template.

Keep it short and factual — it costs tokens on every turn. Write down what you catch yourself repeating in chat; leave out anything the agent can discover in two seconds by looking:

## Commands
```bash
python -m unittest discover -s tests   # tests
ruff check .                           # lint
```

## Conventions
- stdlib `unittest`; never add pytest (zero test dependencies is a hard rule)
- line length 88

## Gotchas
- the version lives only in `lindwyrm/__init__.py`; CI stamps the rest

Commands belong there verbatim so they aren't guessed, and so do decisions the agent can't infer from the code — "we chose X over Y because Z" is what stops it from helpfully reintroducing Y.

LINDWYRM.md takes precedence for notes specific to this agent, and ~/.config/lindwyrm/AGENTS.md holds preferences that follow you between projects. Both are loaded as reference material, explicitly not as instructions that outrank you: a file that arrives with someone else's clone shouldn't be able to give orders to an agent that can run commands.

Context and cost

A model remembers nothing between requests, so the whole conversation is sent again on every turn — your questions, its answers, and the full contents of every file it has read. The history grows, each request grows with it, and that is what costs money and eventually fills the window.

Three mechanisms manage that. They differ sharply in what they cost and in what they destroy, so it is worth knowing which is which.

1. Offloading on arrival — free, always on

The agent reads an 800-line file. Before the next request is even built, the result is written to disk and a stub takes its place in the conversation:

[offloaded: read_file big.py -- 1200 lines, 42.3 KB, ref off_0001]
def parse(source):
    ...
This is a snapshot from when the tool ran; the file may differ now. Use
read_offloaded("off_0001") for the full snapshot, or read_file for current
contents.

Nothing is lost — the full text is on disk and the model can fetch it back. And it costs nothing: the result sits at the very end of the context, so there is nothing after it that would need re-caching.

It applies to the tools that return bulk data — read_file, bash, grep, glob, list_dir. A one-line result like "Wrote 412 chars" is left alone; a stub would be longer than the thing it replaced.

The trigger

Until the conversation crosses a threshold, nothing else happens at all. The threshold is 75% of the window or 200k tokens, whichever comes first:

model window reclaiming starts at
16,000 12,000 (75%)
128,000 96,000 (75%)
1,000,000 200,000 (the ceiling)

The ceiling exists because a share of the window stops being a sensible rule at a million tokens. 75% of 1M is 750k, where a turn whose cache has gone cold costs tens of times more than a warm one, prefill takes real time, and recall degrades. It isn't set lower because reclaiming isn't free either: rewriting history re-charges everything after the edit at cache-miss rates, which only pays for itself after 50–90 turns. Late, but bounded.

2. Offloading retroactively — lossless, but not free

Once over the threshold, older tool results above offload_threshold_tokens are moved to disk, leaving the same recoverable stub. No information is lost.

This one does cost something: editing the middle of the history invalidates the cache from that point on, and the tail is re-charged once at cache-miss rates. That is precisely why it waits for pressure instead of running continuously.

3. Summarizing — lossy, last resort

If space is still short, the model is asked to summarize the older part of the conversation, and all of it is replaced by that summary:

[Summary of earlier conversation, compacted to save context. Treat this as
established background:]
The user was fixing the parser. src/parser.py and tests/test_a.py were
edited. pytest was ruled out. Tests pass. The README still needs updating.

Anything not in the summary is gone for good — unlike an offloaded result, there is nothing to fetch it back from.

What is never touched

A protected zone of recent conversation, measured in tokens rather than messages: four messages might be forty tokens or half the window, depending on whether a test run landed in them.

History is also only ever cut at the boundary of one of your turns. A tool call and its result cannot be separated — the APIs reject a history where they are.

Settings

setting default what it does
offload true master switch for moving results to disk
offload_eager_tokens 0 a single result this big goes to disk on arrival; 0 picks a share of the window
offload_threshold_tokens 1000 old results this big are moved once over the threshold
auto_compact true master switch for automatic reclaiming
compact_threshold 0.75 share of the window that triggers it
compact_max_tokens 200000 absolute ceiling; 0 disables, leaving only the share
compact_keep_tokens 8000 size of the protected zone, capped at a quarter of the window
compact_keep_last 4 messages always kept, however large they are

offload_eager_tokens defaults to a thirty-second of the context window, with a floor of 1,000 — 4,000 tokens on a 128K model, 31,250 on a 1M one. A fixed number cannot be right for both: the same 8,000 tokens are half of a 16K window and a rounding error in a 1M one.

Set it lower than that and you are usually paying rather than saving. The model asked for this content and gets eight lines of it back, so anything it actually needs costs a read_offloaded round trip — and what you reclaimed was, thanks to caching, the cheapest part of the context. Measured on a real session against a 1M window: 111 tool results, 29K tokens between them in a 166K context, 99.5% of input served from cache. Offloading the two largest would have freed 6% of the context for two extra round trips.

Watching it

/context
  ████····················  8% to compaction  (16,240 tokens, measured)
  window 1,000,000 · reclaims at 200,000
  messages: 24   output so far: 3,120 tokens
  served from cache: 142,880 input tokens (94% of input)
  offloaded: 3 result(s), 128 KB on disk

/compact runs it on demand, and /compact keep the API decisions, drop the debugging tells the summarizer what matters.

How hard the model thinks

Providers disagree on how to control reasoning length. Anthropic's own API takes a token budget; DeepSeek accepts thinking_budget and ignores it, choosing instead between effort levels. thinking_effort covers the second kind:

thinking_effort = "low"    # minimal | low | medium | high | xhigh | max

Measured on one question, minimal produced 13k characters of reasoning and max produced 26k. It is worth turning down for routine work: reasoning is billed as output, and on DeepSeek it comes out of the same max_tokens allowance as the answer — a hard question at a high effort can spend the whole budget thinking before it starts to reply.

Prices

Add prices to a preset and a running total appears after each turn:

[[presets]]
name = "flash"
price_input = 0.3         # per million tokens, cache miss
price_cache_read = 0.006  # cache hit — around 50x cheaper
price_output = 1.2

Fresh input, cache reads and cache writes are counted separately, because they are billed separately — on a typical session most of the input comes from cache, and lumping them together overstates the bill several times over. No prices ship built in: they change — DeepSeek's moved twice while this was being written, and now vary by time of day. lindwyrm.example.toml carries the current published figures with the date they were taken; they are the peak ones, since overstating is the safer error. Time-of-day pricing is not modelled, so halve them if you work off-peak.

That gap between a hit and a miss — about 50x on Flash — is why the system prompt and tool schemas are kept byte-stable across turns: anything that shifts the start of the prompt re-charges the whole conversation at miss rates.

Commits

Commits the agent makes carry whatever identity git is already configured with. It is told never to pass --author or set user.name itself, and to stop and ask if the repository has no identity at all — a commit goes out under someone's name, and it should be the name they chose.

To mark that a model helped, set a trailer:

commit_trailer = "Co-Authored-By: {model} via lindwyrm <noreply@lindwyrm.invalid>"

{model} and {preset} are substituted, so the trailer records which model actually did the work:

Make greet() take a name argument

Co-Authored-By: deepseek-v4-flash via lindwyrm <noreply@lindwyrm.invalid>

There is no default: attributing work to an invented identity isn't something to do unasked. GitHub only links a co-author to an account when the address is a real one, so pick accordingly.

Sessions

Every turn is written to ~/.local/share/lindwyrm/sessions/, so closing the terminal doesn't throw the conversation away. --continue resumes the most recent session for the current project — sessions are scoped to the directory they ran in, so one checkout never resumes another's work. /sessions lists them, --resume <id> picks one, and /clear starts a new session without touching what is already saved.

"Most recent" means last worked in, not last started. A session you opened this morning and abandoned does not outrank the one you have been in all week.

What gets saved is the already-compacted history, so resuming costs no more than the session did when you left it. Offloaded results are restored too, so the [offloaded: ...] markers still resolve.

Files are written 0600 in a 0700 directory: a session holds whatever the agent read, which is often source code and sometimes more. --no-save skips writing entirely, and sessions older than 30 days are swept.

Proxies

Off by default, and HTTP_PROXY/ALL_PROXY are ignored unless you ask for them with proxy = "system" — a stray variable in your shell shouldn't silently reroute API traffic. Settings resolve most-specific-first: --proxy, then the preset's own value, then the global one.

proxy = "socks5h://127.0.0.1:1080"   # everything goes through here...

[[presets]]
name = "local-llm"
base_url = "http://localhost:11434/v1"
proxy = false                        # ...except this one, which goes direct

localhost and 127.0.0.0/8 are always reached directly, including under --proxy, which otherwise overrides everything: a proxy resolves localhost on its own side, so proxying a local model server would send the request to a stranger's machine rather than yours. For a model server elsewhere on the LAN, list it — exact host, domain suffix or CIDR:

no_proxy = ["192.168.0.0/16", ".internal", "ollama.box"]

socks5h:// and socks5:// behave identically here: httpx hands the hostname to the proxy, so DNS is resolved on the proxy side either way.

Transient failures (429, 5xx, connect timeouts) are retried with exponential backoff and jitter, honoring Retry-After. A request is only retried when nothing has been streamed yet — never mid-answer.

Configuration

Copy lindwyrm.example.toml to ./.lindwyrm.toml (project) or ~/.config/lindwyrm/config.toml (user). Project settings override user ones — within limits, because a project file arrives with git clone. Unless the project is listed in your user config's trusted_projects, its .lindwyrm.toml cannot set base_url, presets, key_file, proxy, no_proxy, audit_log or session_retention_days, cannot name a context_file outside the project, and its [policy] can only make things stricter. Whatever it tried is listed in a warning at startup.

# ~/.config/lindwyrm/config.toml
trusted_projects = ["~/work"]   # your own checkouts: project files apply in full

The example file documents every option; the essentials:

default_preset = "flash"   # "flash", "pro", or a name from [[presets]]

# Any Anthropic- or OpenAI-compatible endpoint:
[[presets]]
name = "kimi"
format = "openai"
base_url = "https://api.tokenrouter.com/v1"
model = "moonshotai/kimi-k3-free"
api_key_env = ["TOKENROUTER_API_KEY"]
thinking = false

[policy]
read   = "allow"
write  = "confirm"
delete = "confirm"
bash   = "confirm"
bash_allowlist = ["git status", "ls"]
bash_denylist  = ["sudo", "curl"]

[[policy.rules]]        # write freely in generated/, still confirm deletes
path = "generated"
write = "allow"

[[policy.rules]]        # keep the agent out of your keys
path = "~/.ssh"
read = "deny"

API keys are never read from the config file itself — only from the environment or a key_file you point at, so a committed config can't leak one.

How long one turn may run

A single request can run many model→tool cycles. max_tool_steps (default 50) caps them, so a model stuck in a loop can't burn your budget unattended. It is a backstop, not a budget: a genuinely long job — a port, a refactor across thirty files — can reach it with work still to do. When that happens lindwyrm says so, and "continue" picks up exactly where it stopped, with the whole history intact.

max_tool_steps = 200
max_retries = 4     # retries on 429/5xx and connection errors

Security notes and limits

  • bash runs through the shell. The denylist is a backstop, not a jail — it ignores quotes and backslashes (c''url is still curl), but $(printf cu)rl gets past any list. The real protection is that bash defaults to confirm, so you see every command before it runs. For stronger isolation, run lindwyrm in a container.
  • A cloned repo is untrusted input: its .lindwyrm.toml is restricted (see Configuration), and an AGENTS.md that is a symlink out of the project is not read.
  • Path permissions guard against strayed file operations, and resolve symlinks and ../ before matching. They are not a defense against a write you confirm yourself.
  • Summarizing is lossy by nature: it replaces older turns with a model-written summary, cutting only at user-turn boundaries so tool calls are never split from their results. Offloading, which runs first, loses nothing — the full text stays under ~/.local/share/lindwyrm/offload/ (0600 files in 0700 directories) for as long as its session is saved; without one — --no-save, or a crash — it is removed at exit or swept after 7 days. An offloaded result is a snapshot, which is exactly why it is copied rather than re-read later: the file may have changed since.

Tests

python -m unittest discover -s tests

Stdlib unittest, no test dependencies, about a second to run. Covers permission resolution and path escaping, both wire formats, compaction and offload boundaries, proxy resolution, cost accounting, session round-trips, the streaming display, and retry backoff — including an end-to-end suite against a fake provider served over real HTTP. CI runs it on every push across Python 3.11, 3.12 and 3.13, and a release cannot publish on a red suite.

Releasing

Publishing runs on tag push via GitHub Actions using PyPI Trusted Publishing (OIDC) — no API token is stored in this repo. One tag publishes two distributions: lindwyrm and the lwyrm alias.

# bump __version__ in lindwyrm/__init__.py, commit it last, then:
git tag -a v0.5.1 -m "..." && git push origin main && git push origin v0.5.1

lindwyrm/__init__.py is the single source of truth for the version: setuptools reads it as dynamic metadata and the workflow stamps the same value into the alias package. Tag the commit that bumps it, and bump it last — the workflow refuses to build when the tag doesn't match the version in the code. Without that check the mismatch is silent, because skip-existing makes uploading an already-published version succeed, so the run goes green having published nothing.

A version on PyPI can never be re-uploaded; mistakes are fixed by releasing forward, never by moving a tag.

License

MIT — see LICENSE. Copyright (c) 2026 miron404.

Metadata

Release files for lindwyrm 0.9.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for lindwyrm 0.9.0
File Size Uploaded
lindwyrm-0.9.0.tar.gz 129.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for lindwyrm 0.9.0
File Interpreter ABI Platform
lindwyrm-0.9.0-py3-none-any.whl Python 3 none any Details

Total release size: 212.1 kB

Release files / lindwyrm-0.9.0.tar.gz

Download URL lindwyrm-0.9.0.tar.gz
Size 129.1 kB
Tags Source
SHA-256 checksum
How to use checksums
660889da7b3337818902934496a2cef2c66659ad0931e4c2c8e14fa07861849f
BLAKE2b-256 checksum
How to use checksums
de66e2f211449ccf547f41fb9bf83f085601ce19a682ea33d91c671ce4492c42
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.

Transparency log

Release files / lindwyrm-0.9.0-py3-none-any.whl

Download URL lindwyrm-0.9.0-py3-none-any.whl
Size 83.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
bf8bf90784df8d9b39cbd0892d74fa68e9ba42e827642803079f4932f23db026
BLAKE2b-256 checksum
How to use checksums
b8ba06a17eb3e612e38963565fe7001be3323dc91c6803c1161c6b54d92e27d3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.9.0 This release

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page