Caveman Python middleware
Native framework adapters for the Caveman compression runtime. Your framework keeps its inference client, tools, retries, streams, and original conversation. Caveman projects eligible tool-result text into a copied outbound request. Inference stays with your provider.
This quickstart targets middleware 1.0.0 and SDK 1.2.0. Read the release notes and limitations before upgrading.
Run a complete example
Follow the LangChain quickstart for a fresh environment and runtime installation. Start the local runtime separately; the client package does not include it.
python -m pip install 'caveman-sdk==1.2.0' 'caveman-middleware[langchain]==1.0.0' 'langchain==1.4.1' 'langchain-core==1.6.3' 'langgraph==1.2.11'
curl -fsSLo quickstart.py https://docs.caveman.so/examples/middleware/quickstart.py
DEMO_MODE=record python quickstart.py
DEMO_MODE=compress python quickstart.py
DEMO_MODE=off python quickstart.py
The default example makes no provider request. It runs a deterministic native model/tool loop against the real local runtime, checks compression and exact paginated recovery, and asserts that application history retains originals. The optional provider run is separately labeled and can incur charges.
Choose an integration
- Framework guide: public entrypoints, native APIs, recovery ownership, transports, and limitations.
- Compatibility matrix: resolver ranges versus accepted ranges versus exact validation evidence. A range is not an exhaustive test result.
- Deployment: process/container lifecycle, remote TLS/authentication, persistence, session affinity, deadlines, and rollback.
- Recovery and scope: namespace/session/branch/cache epoch, exact originals, excerpts, and expiry.
- Troubleshooting: final reason codes and strict readiness versus normal inference fallback.
- Measurement: quality, latency, retries, recovery calls, cache effects, and provider usage.
Adapters, tiers, and what compresses
Certified adapters (langchain, openai, anthropic, litellm) gate every release. Experimental adapters get the same fail-open guard, logging, and behavioral tests, but their APIs may change in a minor release. caveman_middleware.COMPATIBILITY exposes each family's tier and tested version range.
Compression needs a recovery tool the model can call, so only some entry points compress. The record-only ones never change input: in compress mode each call reports recovery_unbound, logged once.
| Extra | Tier | Compresses | Record-only |
|---|---|---|---|
langchain |
certified | with_caveman_agent, or CavemanMiddleware with its recovery_tool in the agent's tools |
with_caveman_model; CavemanDocumentCompressor without source_expansion |
openai |
certified | with_caveman_openai_tools |
with_caveman_openai |
anthropic |
certified | beta.messages.tool_runner on a with_caveman_anthropic client |
messages.create on that client |
litellm |
certified | CavemanLiteLLM with operator_recovery |
calls without operator_recovery |
google |
experimental | with_caveman_google / with_caveman_google_chat with callable AFC tools, over the Caveman transports |
calls without callable tools |
strands |
experimental | with_caveman_agent |
with_caveman_model |
agno |
experimental | with_caveman_agent |
with_caveman_model |
crewai |
experimental | with_caveman_agent for an agent with tools |
with_caveman_llm alone |
pydantic-ai |
experimental | CavemanCapability on the Agent |
with_caveman_model |
autogen |
experimental | with_caveman_agent, or CavemanWorkbench with its client |
with_caveman_model |
llama-index |
experimental | CavemanFunctionAgent, with_caveman_tools |
with_caveman_model; CavemanNodePostprocessor without source_expansion |
mcp |
experimental | CavemanMCPHost.register plus project_result |
none |
asgi |
experimental | CavemanASGIMiddleware with an ASGIContext.recovery binding |
the same middleware without one |
If your app already has a tool named caveman_retrieve, that tool keeps its name, recovery stays off for that registration, and the adapter logs recovery_name_conflict once.
Providers
- OpenAI: openai 2.x (on
httpx) and 3.x (onhttpx2). The adapter follows whichever HTTP library the installed SDK uses. A Responses API call that the provider would store (storedefaults to true) is sent unchanged with reasonprovider_state_retained. Passstore=False, orallow_stored_responses=Trueto opt in. - Anthropic:
Anthropic,AnthropicBedrock, andAnthropicVertex, sync and async. - LiteLLM: the async and proxy paths project any provider, because they edit LiteLLM's OpenAI-format input. The sync
Routerpath edits the translated request and supports only theopenaiandanthropicproviders; other providers pass through with reasonunsupported_provider. - LlamaIndex: the
OpenAIandAnthropicLLMs, including subclasses such asAzureOpenAI. Bedrock Converse, Vertex, and other LLMs pass through with reasonunsupported_provider. - Pydantic AI:
OpenAIChatModel(including Azure and OpenAI-compatible providers) andAnthropicModel(including Bedrock and Vertex clients).FallbackModel, Google/Gemini, Bedrock Converse, and other models pass through with reasonunsupported_provider. - Google: wrapping returns a clone, so your
ClientorChatis never modified. Wrapping a wrapped object does not stack, andunwrap_google()returns the original. One Caveman transport serves every clone: each call carries the scope of the clone that made it.
Framework versions
Each adapter checks the installed framework version against its tested range:
- Outside the range: the adapter skips (original input, reason
unsupported_version) and logs a warning once. Passaccept_framework_version=Trueafter testing a newer release yourself. - Unreadable version (vendored or bundled builds): the adapter runs and logs
version_unverifiedonce. - Too old or too new to import: importing the adapter raises an
ImportErrorthat names the installed version and the tested range, withcode == "unsupported_version".
Wrapping never raises for a version problem, even in strict mode. To fail fast at startup, call caveman_middleware.preflight(runtime, "langchain"), which reports unavailable for an untested framework. caveman_middleware.ready(...) raises instead.
| Extra | Accepted (gate and extra) | Oldest tested | Newest tested |
|---|---|---|---|
langchain |
langchain 1.1–<2, langchain-core 1.1–<2, langgraph 1.0.2–<2 | 1.1.0 / 1.1.0 / 1.0.2 | 1.4.2 / 1.6.5 / 1.2.12 |
openai |
openai 2.20–<4 | 2.20.0 | 2.54.0, 3.19.2 |
anthropic |
anthropic 1.0–<2 | 1.0.0 | 1.8.0 |
litellm |
litellm 1.95–<2 | 1.95.0 | 1.102.1 |
google |
google-genai 2.18–<3 | 2.18.0 | 2.25.0 |
strands |
strands-agents 1.43–<2 | 1.43.0 | 1.57.0 |
agno |
agno 3.0–<4 | 3.0.0 | 3.0.11 |
crewai |
crewai 1.15.3–<2 | 1.15.3 | 1.15.22 |
pydantic-ai |
pydantic-ai-slim 2.36–<3 | 2.36.0 | 2.49.0 |
autogen |
autogen-agentchat, -core, -ext 0.7–<0.8 | 0.7.0 | 0.7.5 |
llama-index |
llama-index-core 0.14.5–<0.15 | 0.14.5 | 0.14.25 |
mcp |
mcp 2.0–<3 | 2.0.0 | 2.2.0 |
asgi |
none (ASGI 3) | n/a | n/a |
CI tests the oldest versions on Python 3.11 and the newest on Python 3.13, using constraints/floor.txt and constraints/latest.txt. It also runs the certified families on 3.12 and 3.14. Every lane tests the built wheel, installed. A nightly run installs each extra's newest releases with no constraints and opens an issue when one breaks an adapter. The extras stop at the next major version, so a new major is only tested once its range is widened.
[openai] accepts openai 2.20 through 3.x. LiteLLM and CrewAI require openai<3, so [openai,litellm] or [openai,crewai] installs openai 2.x. The adapter supports 2.x, but you can't get openai 3.x in the same environment as those two frameworks. [asgi] has no dependencies; it works with any ASGI 3 server or framework.
Failures, scopes, and logging
- Fail-open: outside strict mode, any adapter or runtime failure sends your original request and reports a reason (
adapter_errorfor a bug in adapter code). In strict mode those reasons raiseMiddlewareError, except version problems (see above). - One warning per problem: every adapter logs one
WARNINGper adapter and reason on thecaveman.middlewarelogger. The line never contains content, scope values, or credentials. Calls that are never LLM calls (embeddings, token counting, other routes) pass through with no report and no warning. - Recovery errors: when the model calls
caveman_retrievewith an unknown or expired handle, or the runtime is unavailable, the tool answers{"error": "<code>"}through the framework's own tool-error result (an errorToolMessagein LangChain,ToolFailedin Pydantic AI,is_errorin Anthropic, MCP and AutoGen, an errorToolResultin Strands) and the run continues, even in strict mode. Cancellation still propagates. - Scopes: thread and session IDs may be free-form. An email or
"user 42 / chat #7"is hashed into a valid scope token. A missing scope, such asscope_from_configwithout athread_id, passes through with reasoninvalid_scope; strict mode raises. - Runtimes: every adapter accepts either
MiddlewareRuntimeorAsyncMiddlewareRuntimeand uses the matching view on sync and async paths. - Long histories: history hashing stops at 2 MiB (
manifest_bytes) and images or bytes are hashed, so a long or multimodal history never skips the whole call. The ASGI adapter's body limit ismax_body_bytes(default 2 MiB); a larger body passes through with reasonpayload_budget.
Contracts to keep
Keep original stored history. Register the actual recovery executor through the native helper; a tool schema alone does not attest recovery. Handles are scope-bound and expire according to runtime retention. Recoverability does not guarantee model quality.
Client modes are off, record, and compress. The client defaults to compression; the standalone runtime defaults to recording. Set both deliberately. Runtime unavailability normally retains original inference input. Strict mode, startup ready(), cancellation, and requested recovery failures have different error contracts.
Final decision reports contain status, reason, transform IDs, replacement/reuse counts, and call IDs. They contain no token counters. Local segment estimates are inferred; provider usage and billed savings are separate evidence. Nothing in this local example verifies billing savings.
Close the runtime client and native framework/provider resources at shutdown. Closing the client does not stop the runtime process. See the deployment guide before sharing a runtime across workers or tenants.
Licence and support
The client, adapters, and Engine runtime are all Apache-2.0. Read LICENSING.md. This package is separate from Caveman Agent SDK. File sanitized reproducible issues in Caveman.
Metadata
Release files for caveman-middleware 1.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| caveman_middleware-1.0.0.tar.gz | 116.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| caveman_middleware-1.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 214.3 kB
Release files / caveman_middleware-1.0.0.tar.gz
| Download URL | caveman_middleware-1.0.0.tar.gz |
|---|---|
| Size | 116.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
0886e93d9684ab76a3627c13e518aa6b93dfdab602a90ed52d017d7fe3913faf
|
|
BLAKE2b-256 checksum How to use checksums |
a33d98a64f78e597a6eab32b9a2868c903791692ca2d6f03bf8e1b40f0eccc25
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.13
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency logRelease files / caveman_middleware-1.0.0-py3-none-any.whl
| Download URL | caveman_middleware-1.0.0-py3-none-any.whl |
|---|---|
| Size | 97.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
3fe070982fadce8a8d382d7ac6aafd16d8caa82a204bf5cd48178b8bc2ffb24c
|
|
BLAKE2b-256 checksum How to use checksums |
ed5fc62eeec50a0ce8d1022a4848dde732f1903f5acf16a7826aaae76d02ed9c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.13
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency log