Skip to main content

litmus — the CLI, and how a repo talks to a workspace

One file, cli/litmus.py, stdlib only. It is the whole client. Everything it does is an HTTP call you could make by hand — the CLI is a convenience, never a requirement.

Getting it (three shapes, one file)

Shape Command When it wins
Installed pip install litmus-ai-sdk → litmus-ai … A customer's agent bootstrapping inside their own repo, and anyone who runs it daily. Names no infrastructure, so it can be pasted into a prompt that goes to someone we do not host.
Served by the workspace curl -sO <workspace>/litmus.py No registry, no install, and the client always matches the workspace that served it. Works from a cloud dev session, a container, an air-gapped-ish laptop, or a machine with no PyPI reachability.
No client at all curl -X POST <workspace>/v0/evidence … CI, another language, a script. The API is the contract.

The first two are the same bytes. The wheel does not contain a rendering of cli/litmus.py; it contains that file, renamed on the way into the wheel so it imports as litmus_ai_sdk (pyproject.toml, force-include). So the digest litmus-ai --version prints is the digest GET /litmus.py.sha256 publishes, a customer can check one against the other, and the probe brief the console lifts out of the served text cannot drift from the client the customer installs.

GET /litmus.py is deliberately unauthenticated: it serves client code, which holds no secrets. The token is what is protected, and it never appears here.

The name

litmus-ai-sdk on PyPI. Sidharth's call, 2026-09-15, after weighing litmus-sdk / litmus-ai-sdk / litmussdk / litmusai — and made because the alternative was worse: the probe prompt was telling a customer's agent to curl a script off a shared-workspace hostname built out of another customer's project name — which reads, to the customer running the command, as an unexplained third party inside their own install instructions. (The reasoning in full, with the names, is internal: DIRECTION.md §5, 2026-09-15/16. This file is the distribution's description on PyPI, so it names no customer.) A pip install line names nobody's infrastructure. It is not up for revisiting.

Two names he did not specify, which are just as permanent, and why they are what they are:

Name Why
distribution litmus-ai-sdk His decision. Also the only one available: PyPI has long held litmus for an unrelated pytest-skeleton generator (1.0.1), so pip install litmus already gets a stranger's package.
import litmus_ai_sdk The PEP 503 normalisation of the distribution name, one-for-one. A reader who sees the import knows exactly what to pip install, and vice versa — the failure mode of a clever short import name is a support conversation that starts by working out which package it came from. It also cannot collide with the internal product's top-level litmus package, which matters below.
console script litmus-ai Not litmus. See the collision.

The collision, and why the console script is litmus-ai

This repo declares two distributions and they share a word:

  • the root pyproject.toml — distribution litmus, package litmus/, console script litmus. The internal V1.1 product: run, grade, analyze, export, compare, world, validate, provision, fixture. Customers never install it and it is published nowhere.
  • cli/pyproject.toml — distribution litmus-ai-sdk, module litmus_ai_sdk, console script litmus-ai. This file. What a customer installs.

If the second one also installed a script called litmus, then in any environment holding both — which is every Litmus engineer who reproduces a customer issue — litmus would be whichever was pip-installed most recently. That is not a theoretical hazard, because the two tools share the verbs export and compare with entirely different meanings, different flags and different outputs. litmus export would quietly do the other thing, succeed, and produce a wrong artifact. A name that can be shadowed is a name that will be.

So: one script, litmus-ai, which can neither shadow nor be shadowed, matches the distribution name the reader just typed, and fails loudly (command not found) for anyone following an older doc — which is the correct failure. Python entry points are per-distribution and per-environment; nothing about pip install litmus-ai-sdk can now reach the internal tool, in either direction.

Treat all three names as permanent: they go into customer prompts, agent instructions and support transcripts, and the cost of changing one later is every one of those.

Publishing (not done, and gated on a human)

The artifacts build locally and are verified locally; nothing has been uploaded. Publishing is outward-facing and irreversible — a PyPI name and version can never be reused — so it is a human's act:

cd cli
uv build --out-dir dist/                       # or: python -m build
python -m twine upload dist/*                  # needs a token

twine reads TWINE_USERNAME (the literal __token__) and TWINE_PASSWORD — an env var name; the value is a PyPI API token that lives in the secret store and never in a file, a log, a commit or a URL. Check the name is still free first (https://pypi.org/pypi/litmus-ai-sdk/json answering 404 means yes); as of 2026-09-16 it was.

Verify the download. A reviewer refused to run this file because it arrived with "no checksum, no integrity guarantee" — correctly. The response carries an X-Litmus-SHA256 header, and GET /litmus.py.sha256 returns the digest on its own so it can be fetched by a different route than the code:

curl -sO <workspace>/litmus.py
curl -s <workspace>/litmus.py.sha256          # {"sha256": "…", "bytes": …}
shasum -a 256 litmus.py                       # compare

Until the workspace has TLS this detects tampering rather than preventing it — anyone able to rewrite the body can rewrite the header. The digest is worth having anyway (it pins exactly what you ran), and it is not a substitute for either TLS or, better, an installed package with a pinned version.

A CLI is not the non-technical surface, and pretending otherwise is how products get built for the wrong person. The non-technical path is the console (/console) and the dashboard: paste a URL, read a report. What makes the CLI usable by a non-expert is that they never type it — their coding agent does, after reading litmus brief probe. The human types English.

The loop

curl -sO http://<workspace>/litmus.py

# the token is a per-org API token from the console (Settings → Tokens, scope
# `push`); it pins the org, so `push` needs no --org. Never the service token.
python3 litmus.py init --url http://<workspace> --token litk_…

# 1. map a real tool surface (nothing leaves the machine yet)
python3 litmus.py probe --mcp "npx -y @acme/mcp-server"
python3 litmus.py probe --mcp "https://acme.example.com/mcp" --header "Authorization: Bearer …"

# 2. ship the evidence
python3 litmus.py push .litmus/sessions/<session>.mcp.jsonl

# 3. get a mock back — compile + mint, one live MCP URL out
python3 litmus.py serve acme .litmus/sessions/*.mcp.jsonl

# 4. the question the product rests on: is the mock like the real thing?
python3 litmus.py compare --real "npx -y @acme/mcp-server" --mock <that URL>

Installed, the same loop is litmus-ai init …, litmus-ai probe … and so on — identical bytes, identical verbs, just on $PATH. litmus-ai --version prints the version, the file's sha256 (compare it with GET /litmus.py.sha256) and the workspace it is pointed at; litmus-ai status is the one that asks the workspace for its build version, because that is a fact about a different machine and --version makes no network call.

Steps 1 and 4 need no Litmus account for the real half — probe writes local files only. A workspace URL and token are needed from step 2 on. A per-org token opens exactly one org: push, and reading that org's evidence, jobs and pipeline back; used against another org it is refused with the mismatch named, and a revoked or expired one is refused with the reason. serve, shell and status still need the workspace's service token, which customers do not get.

The verbs

  • probe --mcp <server cmd | url> — mechanical sweep of every tool: required args filled from the schema, ids harvested from earlier replies and reused, optionals both supplied and omitted, and a deliberate bad id at the end because error envelopes are part of the surface. --read-only skips anything that does not look like a read.
  • call --mcp <…> <tool> --args '{…}' — one recorded call. This is the verb an agent drives, thinking between calls; that is why the protocol in litmus brief probe beats an unattended sweep. Calls append to the same session file.
  • record --wrap gh,gt -- <command> — for surfaces that are CLIs, not MCP: shims wrap the real binaries and log argv/stdout/exit.
  • push <file> — traces, tickets, workflow text, telemetry.
  • serve <name> <sessions…> — compile through the replay gate and mint a world. A mock that cannot reproduce its own evidence refuses to serve, and says why.
  • compare --real <…> --mock <…> — replay identical calls against both.
  • shell --bundle <name-or-dir> --mint — materialize a world locally and PATH-shadow its fake binaries, so an agent runs against the mock unmodified. A name rather than a path is fetched from the workspace and cached in .litmus/bundles/, digest-checked. This is the path for agents whose surface is CLIs rather than MCP: litmus shell --bundle <bundle-id> --mint gives you working linearis, huginn, gh and gt on a machine where none of them are installed, and logs every invocation into the world.
  • brief probe — the two-phase protocol, as instructions for someone else's agent: phase 1 maps the surface (schemas in and out, auth, errors, refusals, which tools write), phase 2 plays out whole jobs so the corpus carries user turns a task can quote. Intelligence flows down; evidence flows up. The fidelity 1.00 on Litmus's battery was measured on brief v1, which was phase 1 alone; v2 does not inherit it.
  • export --org <org> --surface <surface> — assemble the thing a customer receives (a customer PRD's technical requirement 7): the pipeline's suite and mock byte-for-byte, the runner vendored as source with its dependencies pinned by hash, this file with its digest in the README, a config that names secrets by env var only, a manifest of every pin, and a run.sh that runs cells, grades them and renders a scorecard — with no workspace in the loop. It assembles; it authors nothing, and a pin the pipeline did not produce is null in the manifest with the reason. Reproducible: two exports differ only in the timestamp.
  • scorecard <results-dir> — render a scorecard from stored runs alone: per cell n, the five terminal states with infra_failure/invalid shown and excluded, the key's facts one by one, each detector's named findings, and a tier and a path on every number. No bare ratio.

How compare grades

Byte-identity is the wrong bar: a simulator that mints its own ids is still faithful if the shape and the outcome hold. So each replayed call lands in one of five buckets, worst-first in the report:

Verdict Meaning
exact identical payloads
shape_match same fields and types, different values (ids, timestamps)
shape_mismatch the mock returned a different structure
outcome_mismatch one succeeded where the other failed — the serious one
missing_tool the real surface has it, the mock does not

Plus a catalog diff: tools the mock invented, and tools it missed. --push files the whole report in the workspace as evidence.

One thing compare cannot do for you

If the replayed calls mutate, the two sides are not starting from the same place. The mock gets a freshly minted world; the real surface is wherever your last probe left it — cart already full, order already placed, the id you are fetching now existing. Divergence then measures state drift, not fidelity, and it will look like the mock is wrong when it is not.

So for a trace with mutations in it, one of these has to be true:

  • the real surface is reset between the probe and the comparison (a fresh staging tenant, a seeded test account, a teardown script), or
  • you compare on a read-only subset, or
  • you read the result as a state diff and not a fidelity number.

There is no way for us to reset your system from here, so this is a fact about the measurement rather than a bug we can close. compare reports what it saw; knowing which world the real side was in is your half of the contract.

Metadata

Release files for litmus-ai-sdk 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for litmus-ai-sdk 0.1.0
File Size Uploaded
litmus_ai_sdk-0.1.0.tar.gz 92.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for litmus-ai-sdk 0.1.0
File Interpreter ABI Platform
litmus_ai_sdk-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 184.3 kB

Release files / litmus_ai_sdk-0.1.0.tar.gz

Download URL litmus_ai_sdk-0.1.0.tar.gz
Size 92.5 kB
Tags Source
SHA-256 checksum
How to use checksums
115571b666a5d4c0000e764f00653882c58286dab5e57e884faf73e0521800e3
BLAKE2b-256 checksum
How to use checksums
cabd7aeb10a921dd71ae346786e8f254073cd869bd2ed4a64a2530f0fd83f0f1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.10.9 {"installer":{"name":"uv","version":"0.10.9","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / litmus_ai_sdk-0.1.0-py3-none-any.whl

Download URL litmus_ai_sdk-0.1.0-py3-none-any.whl
Size 91.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b422fee27f16208737050ca83706750631501ec0c022bb96a8c47442d2b9d288
BLAKE2b-256 checksum
How to use checksums
d3586913638261b058d6af6a54acd0cce377d16bbd4604d5a15014dfb14ce363
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.10.9 {"installer":{"name":"uv","version":"0.10.9","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

0.9.0

2 release files

0.8.0

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.3.7

2 release files

0.3.6

2 release files

0.3.5

2 release files

0.3.4

2 release files

0.3.3

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page