Skip to main content

stapel-agent

CI coverage pypi downloads python license llms.txt

LLM facade: JSON completion, translation, transcription and summarization in front of swappable model / speech-to-text / image-generation providers (Anthropic, OpenAI-compatible, Claude Code CLI, whisper-http, ElevenLabs, AssemblyAI, openai-images), with a prompt cache and a per-call token ledger. Python port of a prior NestJS service — same HTTP paths/contracts, plus an in-process comm surface for monolith deployments.

Part of the Stapel framework — composable Django apps that deploy as a monolith or as microservices without changing module code.

Install

pip install stapel-agent

At a glance

Fact Value
Version 0.11.0
Python >=3.11 (3.11, 3.12, 3.13, 3.14)
HTTP operations 5
Config axes 4
Usage surface 5
Extension points 4
Fleet dependencies stapel-core

Documentation

capabilities.json · llms.txt (for agents)

What this is

Python port of a prior NestJS service. Same HTTP paths and contracts (stapel-translate's AgentProvider keeps working unchanged), plus a comm surface so monolith deployments call it in-process without HTTP.

Quick start

pip install stapel-agent[anthropic]  # + the Anthropic SDK for the default provider
# settings.py
INSTALLED_APPS = [
    ...
    'stapel_agent',
]

STAPEL_AGENT = {
    "ANTHROPIC_API_KEY": "sk-ant-...",
}

# urls.py — mount the app under agent/
urlpatterns = [
    ...
    path("agent/", include("stapel_agent.urls")),
]

Two surfaces, same contracts:

Surface HTTP comm Function Does
Complete POST /agent/api/llm/complete llm.complete JSON completion: {"prompt", "model": "small|medium|large", "system_prompt"?, "provider"?, "images"?} → parsed JSON in result, prose in comment. images (vision) entries are {url} or {data_b64, mime?} — OCR, screenshots, photo moderation
Translate POST /agent/api/llm/translate llm.translate {"from"/"from_lang", "to", "entries": {key: text}}{key: translated} (cached by prompt)
Transcribe POST /agent/api/llm/transcribe llm.transcribe {"audio_url", "language"?, "diarization"?, "provider"?, "timeout_seconds"?} → a normalized transcript (words, utterances, speakers, timings) via the STT router
Summarize POST /agent/api/llm/summarize llm.summarize exactly one of text / transcript (+ language?, model?, provider?) → Markdown summary + aggregated usage; long inputs are map-reduced
Generate image POST /agent/api/llm/generate-image llm.generate_image {"prompt", "size"?, "n"? (1-10), "provider"?} → raw provider results [{url? | data_b64?, mime}] — storing them in a CDN/asset library is the caller's job
# HTTP (service-to-service: X-API-KEY, or a staff session)
POST /agent/api/llm/complete   {"prompt": "...", "model": "small|medium|large",
                                "provider"?: "...", "system_prompt"?: "...",
                                "images"?: [{"url": "https://..."} | {"data_b64": "...", "mime"?: "image/webp"}]}
POST /agent/api/llm/translate  {"from": "auto", "to": "de", "entries": {"key": "text"}}
POST /agent/api/llm/transcribe {"audio_url": "https://...presigned...", "language"?: "en",
                                "diarization"?: true, "provider"?: "elevenlabs"}
POST /agent/api/llm/summarize  {"text": "..."} | {"transcript": {...llm.transcribe output...}}
POST /agent/api/llm/generate-image {"prompt": "a cat", "size"?: "1024x1024", "n"?: 2}
# comm (in-process in a monolith, transport chosen by STAPEL_COMM)
from stapel_core.comm import call

call("llm.complete", {"prompt": "...", "model": "small"})
call("llm.complete", {"prompt": "what is on this?", "model": "small",
                      "images": [{"url": "https://..."}]})  # url or data_b64 — never raw bytes
call("llm.translate", {"from_lang": "auto", "to": "de", "entries": {...}})
call("llm.transcribe", {"audio_url": "https://..."})  # URLs only — never raw bytes
call("llm.summarize", {"transcript": result["transcript"]})
call("llm.generate_image", {"prompt": "a cat", "n": 2})

LLM failures are HTTP 200 with {"status": "failure", "reason": ...} — 4xx/5xx are reserved for request validation and auth. Successful completions return the parsed JSON in result, prose around it in comment, and snake_case usage (input_tokens / output_tokens).

Every provider call writes a PromptLog row: model, size, source, status, duration and the full token ledger (input / output / thinking / cache-read / cache-write) — per-user and per-source cost accounting needs no other table. Transcriptions land there too (source=transcribe, model = STT provider name, token columns NULL, the fallback walk in metadata.attempts); each summarize pass is a normal LLM row (source=summarize); image generations log as source=generate_image with {count, mimes, bytes_total} in metadata — image bytes never touch the ledger, and multimodal completions never collide with the text-keyed prompt cache.

Transcription routes through an STT provider chain — explicit provider in the request beats STT_LANGUAGE_ROUTES[lang], which beats DEFAULT_STT_PROVIDER + STT_FALLBACK_CHAIN — and falls back only on transient failures (429/5xx/timeouts); bad audio or auth errors never retry on another provider.

Settings — STAPEL_AGENT

Key Default Meaning
MODELS {"small": "claude-haiku-4-5-20251001", "medium": "claude-sonnet-5", "large": "claude-opus-4-8"} Size → model-name map
PROVIDERS {} Overlay merged over the built-in registry (anthropic / openai-compat / claude-code) — add/override entries, None removes one; resolved lazily per request
DEFAULT_PROVIDER "anthropic" Provider used when a request names none
ANTHROPIC_API_KEY "" Key for the Anthropic SDK provider (read lazily)
OPENAI_COMPAT_BASE_URL "" Base URL of any OpenAI-compatible endpoint
OPENAI_COMPAT_API_KEY "" Bearer token for that endpoint
OPENAI_COMPAT_MODELS {} Optional size → model map for openai-compat (missing sizes fall back to MODELS)
CLI_BINARY "claude" Claude Code CLI binary (opt-in provider only)
CLI_TIMEOUT 120 Provider timeout, seconds
MAX_TOKENS 4096 Completion token cap
STT_PROVIDERS {} Overlay merged over the built-in STT registry (whisper-http / elevenlabs / assemblyai) — same semantics as PROVIDERS
DEFAULT_STT_PROVIDER "whisper-http" STT provider used when a request pins none and no language route matches
STT_FALLBACK_CHAIN [] STT providers tried in order after the default — on transient failure only
STT_LANGUAGE_ROUTES {} {"ru": ["gigaam", "whisper-http"], ...} language matrix (beats the default chain)
STT_TIMEOUT 1800 Hard cap (seconds) on one STT provider's submit+poll cycle
STT_DOWNLOAD_MAX_BYTES 134217728 Byte cap on an audio-URL download (128 MiB), enforced mid-stream
STT_DOWNLOAD_TIMEOUT 30.0 Per-socket connect/read timeout for one hop of that download
STT_DOWNLOAD_TOTAL_DEADLINE 300.0 Ceiling on the whole download incl. redirects; a per-call timeout= may only lower it
STT_DOWNLOAD_ALLOWED_HOSTS [] Exact-host allowlist for audio URLs, applied to every hop. Empty = the download is REFUSED (empty is not a wildcard) and W015 warns at boot. Settings-only (NO_ENV) — an env var cannot fill it
STT_DOWNLOAD_ALLOW_ANY_HOST False Opt-out: accept audio from any public host when the allowlist is empty. The other fetch guards still apply
WHISPER_BASE_URL / WHISPER_API_KEY / WHISPER_MODEL "" / "" / "whisper-1" OpenAI-compatible Whisper endpoint — the OpenAI API or self-hosted faster-whisper (key optional)
ELEVENLABS_API_KEY / ELEVENLABS_STT_URL / ELEVENLABS_STT_MODEL "" / Scribe URL / "scribe_v2" ElevenLabs Scribe
ASSEMBLYAI_API_KEY / ASSEMBLYAI_BASE_URL / ASSEMBLYAI_MODEL "" / "https://api.assemblyai.com" / "universal" AssemblyAI (async submit+poll)
IMAGE_PROVIDERS {} Overlay merged over the built-in image registry (openai-images) — same semantics as PROVIDERS
DEFAULT_IMAGE_PROVIDER "openai-images" Image provider used when a request pins none
IMAGES_BASE_URL / IMAGES_API_KEY "" / "" OpenAI-compatible /images/generations endpoint + key; both fall back to the OPENAI_COMPAT_* pair
IMAGES_MODEL "" Optional model name (gpt-image-1, ...); empty = omitted from the request
CACHE_LOOKUP {"llm_facade": False, "translate": True, "summarize": False} Per-source cache-by-prompt toggle (used by the default cache policy)
CACHE_TTL 604800 Cache window in seconds (7 days); older rows are ignored (default policy)
CACHE_ALLOW_UNSCOPED [] Sources whose content the host declares non-personal; a call without user_id otherwise skips the cache
PROMPT_LOG_RETENTION_DAYS 90 Age after which purge_prompt_logs scrubs a ledger row's text, keeping its token counters
PROMPT_LOG_RETENTION_SCHEDULED False Declares that the host's scheduler runs purge_prompt_logs; otherwise a system check (W014) says the window is unenforced
CACHE_POLICY "stapel_agent.cache.PromptLogCachePolicy" Dotted path to a CachePolicy subclass — swap the prompt cache (Redis, no-op, ...) without forking

Provider matrix

Name Class Backend Needs
anthropic (default) providers.anthropic.AnthropicProvider Anthropic SDK anthropic extra + ANTHROPIC_API_KEY
openai-compat providers.openai_compat.OpenAICompatProvider Any /chat/completions dialect: OpenAI, DeepSeek, MiMo, GLM, Kimi OPENAI_COMPAT_BASE_URL (+ key)
claude-code providers.claude_cli.ClaudeCodeCLIProvider Spawns claude -p ... --output-format json The CLI in the host image

No CLI in any default path. claude-code is strictly opt-in: it exists for hosts that ship the Claude Code CLI in their image and want the CLI to handle its own authentication (provider: "claude-code" per request, or DEFAULT_PROVIDER override). There is no OAuth credential reading and no background token-refresh — that plumbing was deliberately dropped.

Adding, overriding and removing providers (merge semantics)

STAPEL_AGENT["PROVIDERS"] is an overlay merged over the built-ins, not a replacement dict — adding your provider never requires restating the three shipped ones, and setting a name to None removes it:

# settings.py — one line per change, built-ins stay available
STAPEL_AGENT = {
    "PROVIDERS": {
        "acme": "myproject.llm.AcmeProvider",  # add a custom backend
        "claude-code": None,                    # remove a built-in
    },
    "DEFAULT_PROVIDER": "acme",
}

Or register at runtime from your app's AppConfig.ready() (highest precedence):

from stapel_agent import register_provider

register_provider("acme", AcmeProvider)  # class or dotted path

A custom backend is just a stapel_agent.LlmProvider subclass returning a ProviderResult. Django system checks (stapel_agent.E001/W001/W002) flag a DEFAULT_PROVIDER that is not in the effective registry, unimportable dotted paths and non-LlmProvider entries at startup. See MODULE.md for the full extension-point map.

STT provider matrix

Name Class Backend Needs
whisper-http (default) stt.providers.whisper_http.WhisperHttpProvider OpenAI Whisper API or any self-hosted server speaking the same dialect (faster-whisper, whisper.cpp); accepts url/path/bytes refs WHISPER_BASE_URL (+ key for OpenAI)
elevenlabs stt.providers.elevenlabs.ElevenLabsProvider ElevenLabs Scribe (synchronous multipart), diarization ELEVENLABS_API_KEY; URL refs only
assemblyai stt.providers.assemblyai.AssemblyAIProvider AssemblyAI async submit+poll, diarization, 99-language universal model ASSEMBLYAI_API_KEY; URL refs only

The STT registry has the same merge semantics as the LLM one (STAPEL_AGENT["STT_PROVIDERS"] overlay, None removes, or register_stt_provider() at runtime), and its own startup checks (stapel_agent.W003/W004). A custom engine is a stapel_agent.SttProvider subclass returning a NormalizedTranscript — MODULE.md walks through a self-hosted GigaAM adapter as the worked example.

Vision & image generation

The two vision-capable LLM providers map images into their dialects automatically — anthropic to image content blocks (URL or base64 source), openai-compat to image_url parts (URL or data URI). claude-code has no vision: an image request through it fails fast with status: "failure".

Image generation ships one built-in backend, openai-images — anything speaking the OpenAI POST {base}/images/generations dialect (OpenAI, Together, self-hosted). It is the third instance of the merge-registry pattern (STAPEL_AGENT["IMAGE_PROVIDERS"] overlay, None removes, register_image_provider() at runtime; startup checks stapel_agent.W005/W006). Vendors with their own protocols (Stability, Ideogram, ...) are an app-layer stapel_agent.ImageGenProvider subclass — recipe in MODULE.md. The agent returns raw results and writes the ledger; storage/placement belongs to the calling tier.

Privacy & retention

PromptLog is both an accounting ledger and a copy of the conversation, so it is treated as customer content:

  • The cache is per-tenant. A cached answer belongs to the user_id that paid for it; a call with no scope skips the cache entirely unless its source is listed in CACHE_ALLOW_UNSCOPED.
  • Subject requests reach it. AgentGDPRProvider (section agent) is registered into stapel_core.gdpr.gdpr_registry at app startup — export returns the user's rows, erasure scrubs their text and unlinks the subject while the token counters survive for accounting.
  • Text expires. Schedule manage.py purge_prompt_logs (--days N, --dry-run): it scrubs prompt, system prompt, response and error text on rows older than PROMPT_LOG_RETENTION_DAYS and keeps the counters.

License

MIT — see LICENSE.


This page is assembled by stapel-readme from docs/readme.md plus the contract artifacts in docs/. Edit the prose in docs/readme.md; the badges, facts and links above and below it are generated — do not hand-edit README.md.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

stapel_agent-0.11.0.tar.gz (279.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

stapel_agent-0.11.0-py3-none-any.whl (228.4 kB view details)

Uploaded Python 3

File details

Details for the file stapel_agent-0.11.0.tar.gz.

File metadata

  • Download URL: stapel_agent-0.11.0.tar.gz
  • Upload date:
  • Size: 279.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for stapel_agent-0.11.0.tar.gz
Algorithm Hash digest
SHA256 d48cba56241a261ae5cef053108ce2d7bf94a6ee46d9533d14b7bdfed86ebe65
MD5 a6d8560ddc6d50ca3bc38f46be7c8e2d
BLAKE2b-256 4a05525fa3ed1981614c70334ad01c1049d6790c315ef36285fb8fddd7b6cfee

See more details on using hashes here.

Provenance

The following attestation bundles were made for stapel_agent-0.11.0.tar.gz:

Publisher: publish.yml on usestapel/stapel-agent

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file stapel_agent-0.11.0-py3-none-any.whl.

File metadata

  • Download URL: stapel_agent-0.11.0-py3-none-any.whl
  • Upload date:
  • Size: 228.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for stapel_agent-0.11.0-py3-none-any.whl
Algorithm Hash digest
SHA256 318992981e03dcfc64902cfdf3cbe29d87f857fde17d219ac558a06f5d097946
MD5 b93c7dc513b95ceb532dd6189c7b5bc1
BLAKE2b-256 951a051e9faea868c02e15a4384f237c73e4fcd41959ba4db1d5b4bb96efbeb6

See more details on using hashes here.

Provenance

The following attestation bundles were made for stapel_agent-0.11.0-py3-none-any.whl:

Publisher: publish.yml on usestapel/stapel-agent

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.13.1

2 files

0.13.0

2 files

0.12.0

2 files

This release

0.11.0 This release

2 files

0.10.0

2 files

0.9.0

2 files

0.8.1

2 files

0.8.0

2 files

0.7.1

2 files

0.7.0

2 files

0.6.7

2 files

0.6.6

2 files

0.6.5

2 files

0.6.4

2 files

0.6.3

2 files

0.6.2

2 files

0.6.1

2 files

0.6.0

2 files

0.5.0

2 files

0.4.1

2 files

0.4.0

2 files

0.3.0

2 files

0.2.11

2 files

0.2.10

2 files

0.2.9

2 files

0.2.8

2 files

0.2.7

2 files

0.2.6

2 files

0.2.4

2 files

0.2.3

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page