Skip to main content

Composable LLM reasoning patterns with budget-aware execution

Project description

ExecutionKit

Composable LLM reasoning patterns. Consensus voting · Iterative refinement · ReAct tool loops · Structured JSON · Zero SDK lock-in.

Python 3.11-3.13 License: MIT PyPI Releases CI coverage Linting: ruff Type checked: mypy Docs


ExecutionKit fills the gap between raw chat calls and full orchestration stacks — more power than one-off prompts, less weight than a framework. Provider-agnostic, zero runtime dependencies (stdlib only), mypy --strict clean, with lightweight eval, tracing, routing, workflow, planning, and approval primitives.

📚 Full documentation: tafreeman.github.io/executionkit

Architecture

ExecutionKit is the execution-primitive layer of a two-tier stack. The companion repo, agentic-runtime-platform, handles orchestration above it — ExecutionKit patterns run inside each agent step there.

File Contents
docs/architecture.md Module map, dependency graph, error hierarchy, security notes
CONTRIBUTING.md — Anti-Scope What the library does not do, and why
examples/ OPENAI_API_KEY=<your-key> python examples/quickstart_openai.py

For implementation details, start with docs/architecture.md and the public docs site. The diagram below shows the intended layering.

flowchart TB
    subgraph ARP ["agentic-runtime-platform  —  orchestration layer"]
        W["Persistent DAGs · YAML · multi-agent scheduling"]
    end

    subgraph EK ["ExecutionKit  —  pattern library  (this repo)"]
        direction LR
        C["consensus()"] ~~~ R["refine_loop()"] ~~~ RA["react_loop()"] ~~~ S["structured()"] ~~~ PI["pipe()"]
        PI ~~~ O["Router · Workflow · Plan · ApprovalGate · TraceEvent · evals"]
    end

    subgraph P ["LLM Provider  —  any OpenAI-compatible endpoint"]
        API["OpenAI · Ollama · vLLM · Groq · Together · Azure"]
    end

    ARP -->|"calls patterns inside each agent step"| EK
    EK -->|"HTTP POST /chat/completions"| P

Platform role (ADR-023). ExecutionKit is the OpenAI-message-format execution kernel of the stack: the runtime aligns its provider seam onto ExecutionKit's LLMProvider / LLMResponse contract rather than maintaining a parallel one. The decision, migration plan, and functionality-preservation matrix live in the runtime repo at docs/adr/ADR-023-*. The shared value types (LLMResponse, ToolCall, TokenUsage) and the error hierarchy live directly in executionkit/ today. ADR-023 reserves a future extraction path if agentic-runtime-platform ever needs a separate contracts package, but there is no standalone executionkit-contracts distribution in v0.2.0.

Development note: Built with AI-assisted development under human review; architecture, tests, release gates, and public documentation remain maintainer-owned and verified through the repo's lint, type, test, and security checks.

Quick Start

pip install executionkit
import asyncio
import os
from executionkit import Provider, consensus

async def main() -> None:
    async with Provider(
        "https://api.openai.com/v1",
        api_key=os.environ["OPENAI_API_KEY"],
        model="gpt-4o-mini",
    ) as provider:
        result = await consensus(provider, "What is the capital of France?", num_samples=5)
        print(result.value, result.metadata["agreement_ratio"], result.cost)

asyncio.run(main())

What you see when you run it:

$ pip install executionkit
$ export OPENAI_API_KEY=<your-key>
$ python examples/quickstart_openai.py
Answer: Paris
Agreement: 100%
Cost: TokenUsage(input_tokens=57, output_tokens=6, llm_calls=3)

See the Quick Start guide for a complete walkthrough.

What shipped in v0.1.0

Five security fixes in the initial release: prompt-injection sandboxing in the default evaluator via XML delimiters and 32 KB input truncation; API key masking in Provider.__repr__; credential redaction in HTTP error messages; information hiding in tool error returns; and supply-chain hardening with Bandit SAST + pip-audit in CI. Six net-new features including the structured() pattern, optional httpx pooling backend, max_history_messages capping, JSON Schema tool-arg validation, async context manager lifecycle for Provider, and MkDocs Material docs with ADRs. Full notes: CHANGELOG.md · docs site.

Patterns

Pattern What it does
Consensus Run N parallel calls, vote on the result, return the majority answer with confidence.
Iterative Refinement Generate, score, refine. Bounded loop with a quality gate.
ReAct Tool Loop Think-act-observe loop with JSON-Schema-validated tool calls.
Structured Output Parse JSON responses with custom validators and automatic repair retries.
Pipe Chain patterns end-to-end with a shared budget.

Lightweight primitives

ExecutionKit also exposes small stdlib-only primitives for the glue code around pattern calls:

  • Evals. EvalCase and run_eval_suite() run deterministic golden checks in CI; live_provider_from_env() enables opt-in live checks via EXECUTIONKIT_LIVE_EVAL=1, EXECUTIONKIT_BASE_URL, and EXECUTIONKIT_MODEL.
  • Observability. TraceEvent callbacks can receive structured events for LLM calls, retries, tool calls, workflow steps, plan steps, approvals, cost, and latency.
  • Routing. Router and RouteRule select a provider before a pattern call without changing the pattern implementation.
  • Workflow and planning. Workflow/Step execute simple dependency-ordered fan-out DAGs; Plan/PlanStep execute ordered plan-then-act flows.
  • Approval gates. ApprovalGate can require human or policy approval before tool execution, workflow steps, or plan steps.

Why ExecutionKit

  • Provider-agnostic. OpenAI, Ollama, vLLM, GitHub Models, Together, Groq, llama.cpp, and Azure via an OpenAI-compatible gateway.
  • Zero SDK lock-in. Structural LLMProvider protocol — any conforming object works without inheritance.
  • Composable. Patterns are async functions. Wrap them, chain them with pipe(), or drop them inside a larger orchestrator like agentic-runtime-platform.
  • Budget-aware. TOCTOU-safe max_cost enforcement across parallel calls; llm_calls counts every dispatched wire attempt, including failed retries.
  • Secure-by-default. API key masking, broad credential redaction in errors, top-level JSON-Schema tool validation, prompt-injection-hardened default evaluator, and optional approval gates.
  • Eval-aware. A deterministic golden suite and a model-failure corpus assert output correctness (not just coverage) in normal CI, with EvalReport.accuracy/summary() metrics; judge-calibration and live-provider regression tiers stay explicitly env-gated.

Built for Platform Teams

ExecutionKit targets three groups who need LLM reliability without runtime coupling:

  • Platform / infra engineers dropping a reasoning primitive into an existing service — no SDK to pin, no dependency conflict. pip install executionkit adds one package with zero transitive dependencies; provider swap is one constructor call.
  • Solutions architects evaluating multi-vendor strategies — the structural LLMProvider protocol means vendor A and vendor B are runtime-swappable with no code changes outside the constructor.
  • AI-native teams building beyond chat — consensus voting, iterative refinement, and ReAct tool loops are the building blocks for production-grade LLM behaviour without pulling in a full framework.

If you need persistent, declarative, multi-agent orchestration on top, agentic-runtime-platform layers over ExecutionKit and handles scheduling, runtime state, and fleet-level evaluation gating.

Relationship to agentic-runtime-platform

ExecutionKit and agentic-runtime-platform occupy different layers of the same stack:

ExecutionKit agentic-runtime-platform
Role Pattern library Orchestration runtime
Scope Reasoning patterns plus lightweight Python routing/workflow/planning primitives Multi-agent DAG workflows with tiered model routing
Workflow authoring Python functions and named async steps Declarative YAML
Dependencies Zero (stdlib only; httpx optional) FastAPI, LangGraph, Pydantic, provider SDKs
Use when You need a reasoning primitive — vote, refine, tool loop, trace, route, simple DAG You need to orchestrate many agents with scheduling, persistence, retries, and evaluation

agentic-runtime-platform uses ExecutionKit patterns internally as the execution primitive for each agent step. Build atop agentic-runtime-platform for free; install ExecutionKit alone if you want the patterns without the orchestration overhead.

Documentation

The canonical reference is the docs site:

Development

pip install -e ".[dev]"
ruff check . && ruff format . --check
mypy --strict executionkit/
pytest --cov=executionkit --cov-fail-under=80

See CONTRIBUTING.md for the full dev workflow.

License

MIT — see LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

executionkit-0.2.0.tar.gz (357.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

executionkit-0.2.0-py3-none-any.whl (53.7 kB view details)

Uploaded Python 3

File details

Details for the file executionkit-0.2.0.tar.gz.

File metadata

  • Download URL: executionkit-0.2.0.tar.gz
  • Upload date:
  • Size: 357.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for executionkit-0.2.0.tar.gz
Algorithm Hash digest
SHA256 7115ee1b08ba6ed67f53b88a9ae5872f1b96c6a2d401aea2fea4430675772686
MD5 083d088577cbdae8f9a5ff55451c8d18
BLAKE2b-256 72f1a89f1cc6f5d7506e222322d245e91a64b18cfd254d6ca7901511523a8689

See more details on using hashes here.

Provenance

The following attestation bundles were made for executionkit-0.2.0.tar.gz:

Publisher: publish.yml on tafreeman/executionkit

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file executionkit-0.2.0-py3-none-any.whl.

File metadata

  • Download URL: executionkit-0.2.0-py3-none-any.whl
  • Upload date:
  • Size: 53.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for executionkit-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 db4acadb7eab6275bfaf4a2c1552b0ca841f81bb1eb7dfd2db5eac4db19629c8
MD5 7e381e4aff17e52e3a1809b4321496a4
BLAKE2b-256 8dfeefd362026654fb9e2da03acf655d15c88c31b0c58b298d2e4ac3e3d4f8cb

See more details on using hashes here.

Provenance

The following attestation bundles were made for executionkit-0.2.0-py3-none-any.whl:

Publisher: publish.yml on tafreeman/executionkit

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page