flanj — Flanj SDK for Python
Your integration didn't break. It started being wrong. Every call succeeded. That's why nothing caught it.
Early — MCP only. See Status.
flanj instruments an MCP client session (mcp.ClientSession). It records the tools/list catalogue
each MCP server hands your agent and the tools/call traffic that follows. It redacts sensitive data at the
source, then exports the redacted records over OTLP to a Flanj collector. The collector checks every call
against the contract the server published and flags drift.
Trust posture. Capture is out of band: the SDK wraps the session your agent already uses, and never proxies, rewrites, delays or blocks a call. Arguments pass through by reference, the tool's own result object is returned, and exceptions propagate unchanged. Redaction runs in your process, before a record is exported. Redacted records go only to a collector you run in your own environment; bodies never leave it. See REDACTION.md for the floor and how it is held identical across languages.
Quick start
Needs Python 3.10+ and a running Flanj collector. The collector README's
Run it on Kubernetes section is the
preferred way to deploy it; Run it with Docker
is the other. Python 3.10 is the floor of the official mcp package this SDK instruments, so there is no
supported MCP client below it; an older runtime is refused with one sentence rather than failing somewhere
inside a capture path.
pip install flanj
Needs Python 3.10+: on an older interpreter
pipdoes not refuse, it silently installs the unrelated0.0.1placeholder that once held the name.
Make this the first line of your program — before anything imports mcp:
import flanj.register # noqa: F401
That is one line, and nothing else changes. It starts the OTLP pipeline, flushes on exit and
auto-instruments every MCP client session your program opens afterwards, placing each server on its own
edge. It prints one line naming the endpoint and the resolved service name (FLANJ_QUIET=1 silences it).
It is the counterpart of the TypeScript SDK's node -r @flanj/sdk/register.
If your collector is running with Docker, there is nothing to point it at: http://localhost:4318/v1/logs
is this SDK's default FLANJ_OTLP_ENDPOINT already, and that is where that collector listens.
On Kubernetes
The chart creates a front Service with a fixed name, so the one address that is right for everyone who
installed with the Run it on Kubernetes
command is http://flanj-collector.flanj:4318/v1/logs. Set it on your own workload, not your shell:
the SDK runs in the agent's pod, where localhost is not the collector.
env:
- name: FLANJ_OTLP_ENDPOINT
value: http://flanj-collector.flanj:4318/v1/logs
The chart also renders a ConfigMap/flanj-endpoint for teams that would rather use envFrom than a
literal value. A pod can only reference a ConfigMap in its own namespace, so the chart has to be told
which namespaces to render it into — see the
chart README.
Verify — after your agent has made at least one tool call. On Docker, docker compose up -d already
starts the bridge that publishes the collector's UI at 127.0.0.1:5335. On Kubernetes, forward it first:
kubectl -n flanj port-forward sts/flanj-flanj-collector-store 5335:5335
curl -s http://127.0.0.1:5335/api/health
then open http://127.0.0.1:5335 and look at the Traffic tab: your tool call should be there, redacted.
If that curl answers Failed to connect: on Docker, confirm the compose stack is still up
(docker compose ps); on Kubernetes, confirm the port-forward above is still running. Ingest on :4318 is
a separate, ordinary port either way.
Load flanj first
A Python ClientSession holds two in-memory streams and no URL, so the SDK learns where each MCP server is
when its transport opens — it wraps streamable_http_client, sse_client and stdio_client. A program that
did from mcp.client.stdio import stdio_client before flanj loaded holds the unwrapped function, and flanj
never sees those transports. The same rule applies to Datadog's import ddtrace.auto, for the same reason.
A session whose transport flanj never saw is not guessed at: its edge is unknown, its calls are recorded
without bodies (it might be an internal server, whose bodies are never read), and flanj says so once on
stderr, naming the fix. unknown is the one behaviour this SDK has that the TypeScript SDK does not; a
JavaScript MCP client keeps its transport, so it can always place a server.
MCP quick start
Nothing is configured per server. With the zero-code entry loaded, an ordinary session is already captured:
import flanj.register # noqa: F401 — first line, before `mcp` is imported
from mcp import ClientSession
from mcp.client.streamable_http import streamable_http_client
async with streamable_http_client("https://mcp.acme.com/mcp") as (read, write):
async with ClientSession(read, write) as session:
await session.initialize()
await session.list_tools() # -> a contract snapshot
await session.call_tool("get_balance", {"account_id": "x"}) # -> a captured, redacted call
For each tools/call: the arguments as the request body; structuredContent — else the content[]
text — as the response body; the outcome (isError); the server's identity from the result's _meta; and
the JSON-RPC request id your client generated, labelled as client-generated. A call the server rejected
(a JSON-RPC error rather than a result with isError) also records the error's code. A Tasks handle
(tasks/get) is recorded as an envelope with no body, so nothing models the envelope as the tool's own
output shape.
For each complete tools/list: the server's own declared schemas, verbatim, never re-inferred. The
catalogue is refetched and re-snapshotted on notifications/tools/list_changed
(refetch_on_list_changed, on by default), so a server that changes its tools mid-session is caught when it
does.
For a server your app launched over stdio: how it was launched (npx @stripe/mcp@0.2.1 …) — the command
and its arguments only, each floor-redacted, never the environment or the working directory.
Your service's name is the only thing you configure, and it defaults to your app's own name
(python -m myapp → myapp), then flanj-sdk. It travels as the OTLP resource service.name, and the
collector shows it as a Service column beside the counterparty and as a Traffic filter. It names your
own internal topology, so it stays on your collector: a service name is never sent to the control plane.
Instrumenting a client yourself
import flanj # still first
handle = flanj.start() # the OTLP pipeline; wraps the transport openers
from mcp import ClientSession
from mcp.client.streamable_http import streamable_http_client
async with streamable_http_client("https://mcp.acme.com/mcp") as (read, write):
async with ClientSession(read, write) as session:
handle.instrument(session) # this handle's logger and body cap
await session.initialize()
# Use `session` exactly as before. Nothing about its behavior changes.
result = await session.call_tool("get_balance", {"account_id": "acct_1"})
To instrument every session without the zero-code entry, call
flanj.register_mcp_auto_instrumentation(logger=handle.logger) after start().
Configuration
| Option / variable | Default | Meaning |
|---|---|---|
OTEL_SERVICE_NAME / service_name= |
the app's own name, else flanj-sdk |
service.name on exported records. Precedence: explicit service_name=, then OTEL_SERVICE_NAME, then the running script or module's own name (python -m myapp → myapp; python path/app.py → app), then flanj-sdk. The collector derives each record's integration itself; this SDK never sends flanj.integration. |
FLANJ_OTLP_ENDPOINT / otlp_endpoint= |
http://localhost:4318/v1/logs |
The collector's logs endpoint. Then OTEL_EXPORTER_OTLP_LOGS_ENDPOINT, then OTEL_EXPORTER_OTLP_ENDPOINT (+ /v1/logs). If an export fails, the first failure prints one line. |
FLANJ_BODY_CAP_BYTES / body_cap_bytes= |
16384 |
Per-body capture cap. |
endpoint= / server_kind= |
detected | Only for a session whose transport flanj could not see: the streamable-HTTP URL (its host[:port] is the edge key), or "stdio". |
refetch_on_list_changed= |
True |
On notifications/tools/list_changed, refetch the catalogue and re-snapshot. Same default as the TypeScript SDK. |
FLANJ_QUIET |
unset | 1 silences the flanj.register startup line. |
FLANJ_SILENCE_CAPTURE_WARNINGS |
unset | Silences the one-time lines for a failed capture or an unknown edge. Same variable, same lines, in the TypeScript SDK. |
Records are flushed on normal exit and on SIGTERM / SIGINT (flanj.flush_on_exit(handle), which
flanj.register installs), bounded to five seconds, after which your own signal handling runs as before.
What is captured
Captured: every tools/call and tools/list on an instrumented session, over any transport the
official client supports (streamable HTTP, stdio). The fields are listed under
MCP quick start.
Not captured:
- HTTP request/response bodies. Python has no
node:httpchoke point to patch the way the TypeScript SDK does. Per-library HTTP capture (requests/httpx/aiohttp) is not started. This is the one intended difference between the two SDKs. tasks/getpayloads. A Tasks handle is recorded as an envelope with no body, so nothing models the envelope as the tool's output shape.
Edges. A server reached over a URL is external or internal by its host, the same rule the collector
uses; internal servers are metadata-only. A server your app launched over stdio is local-process, keyed
by the name it reports, and its bodies are captured: it usually wraps someone else's API. A server flanj
could not place is unknown, metadata-only. Every captured body is redacted in your process before it is
exported; see REDACTION.md for what is redacted and how.
When capture itself fails it stops collecting and says so — once, on stderr, naming what broke and that
your application is unaffected. Silence is the failure mode this SDK exists to remove, and a collector
showing nothing looks exactly like an agent making no calls. FLANJ_SILENCE_CAPTURE_WARNINGS=1 turns the
line off once you have read it.
MCP clients: the contract arrives with the traffic
REST drift detection needs a spec somebody published and kept accurate. MCP servers publish their contract
on every single call — tools/list is the spec. So the collector has the baseline from the first call
your agent makes, for every MCP server it touches, with nothing to configure and nothing to upload: the
observed tools/list is forwarded as a contract snapshot, versioned by content hash, and every later
tools/call is checked against it.
That is not a convenience difference. "Nobody publishes an accurate OpenAPI spec" is the strongest practical objection to drift detection on REST, and it does not apply to MCP at all. It matters most for agents, the most drift-fragile API consumers anyone has built: an agent reads a tool's description to decide what to do, so a description that changes under it changes what it does, and nothing anywhere logs an error. For this SDK that is not one feature among several — it is the whole product.
There is nothing to configure per server: the collector derives each MCP server's integration from its peer host (or, over stdio, the name it reports), so two servers never share a baseline.
It stays out of the way
call_tool is a coroutine, so three properties matter more here than in a synchronous SDK. Each is
locked by a test rather than asserted in prose (tests/mcp/test_async_transparency.py):
- cancellation passes through untouched: a cancelled call is the caller withdrawing, not an outcome, so it is neither captured nor delayed;
- no suspension point is added: the wrapper awaits the wrapped coroutine and nothing else, so scheduling order is unchanged;
- the result is read, never consumed: capture holds no reference past the call and iterates nothing.
Status
Early — MCP only. Not Supported: a language is called supported only once the whole loop runs on it end to end in our own integration harness, with that suite's assertions green. Until then it says early, here and everywhere else.
Languages. Node / TypeScript — supported: @flanj/sdk, HTTP
egress and ingress plus the MCP client. Python — early: this package, MCP client only. Apart from HTTP
capture the two are the same SDK: same defaults, same records, same entry points.
Where this stops, said out loud. No HTTP body capture (see What is captured). A call the collector cannot check against a contract is captured and reported as not validated, never as conforming. Pre-release (v0); see docs/CONCEPTS.md for the engineering model and CLAUDE.md for the repo map.
Flanj turns a detection into something you can act on with the other team. The SDK is open source under Apache-2.0; the collector is source-available under the Elastic License 2.0; the network layer that carries a flagged finding between the two teams is hosted.
Development
uv venv --python 3.10 && uv pip install -e ".[dev]"
pytest # the floor's contract suites, the wire seam, the async contract
ruff check . && mypy # lint and types
bash scripts/smoke-pack.sh # build the wheel, install it into a FRESH venv, drive a real MCP server
Security
Report vulnerabilities, including any redaction gap, privately — see SECURITY.md. Never include a real card number or real personal data.
License
Apache-2.0. Contributions require a DCO sign-off; see CONTRIBUTING.md.
Release files for flanj 0.2.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| flanj-0.2.1.tar.gz | 77.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| flanj-0.2.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 182.2 kB
Release files / flanj-0.2.1.tar.gz
| Download URL | flanj-0.2.1.tar.gz |
|---|---|
| Size | 77.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
1d416c93726ae65f03942b6b52ea3a0c8b245bfac54bfec201c84e823c6450f6
|
|
BLAKE2b-256 checksum How to use checksums |
16253dcada9c5d529f23bdc70768b68856a256e9b85821e0d891a6dd0bb7ed1a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 22, 2026.
Transparency logRelease files / flanj-0.2.1-py3-none-any.whl
| Download URL | flanj-0.2.1-py3-none-any.whl |
|---|---|
| Size | 104.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5bd1db6805ad526923f7da3ce72f55cdb0d79eadbc37f625e31aa43af92545b9
|
|
BLAKE2b-256 checksum How to use checksums |
e07e7afe8ea5ad268a21c04ffb97bb1c9e39366b4d7bb9f9c86f406563f02fd8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 22, 2026.
Transparency log