Skip to main content

Provider-agnostic AI Runtime SDK

Project description

AI Runtime

AI Runtime is a provider-agnostic Python runtime for building AI applications and agent workflows.

Features

  • Provider abstraction with a unified LLM contract
  • Built-in default provider registry and plugin discovery
  • Custom provider registration with test doubles
  • Conversation management and session-based execution
  • Streaming responses with event processing
  • Streaming timeout/cancellation boundary handling
  • Automatic tool-calling loop (model requests tools → runtime executes → re-invokes)
  • Capability-gated request mapping (tools, structured output, vision, metadata)
  • Expanded provider contract: chat, stream, embeddings, image, transcription
  • Retries via ProviderConfig.max_retries (forwarded to the backend)
  • Uniform events: chat() and stream() both emit StreamEvents
  • Rich streaming events: text, usage, tool call, tool result, thinking, permission
  • Execution engine and pipeline stages
  • Provider metadata, capabilities, and request mapping
  • Fully tested runtime and provider integration coverage

Agentic capabilities (v0.6.0)

ai_runtime now closes the gap with agentic coding tools (Claude Code, OpenAI Codex, Cursor):

  • Plan modeAgentRunner.plan() / ExecutionEngine.plan() produce a reviewable Plan (read-only, no tools run) before execution.
  • Sub-agents — declare SubAgentSpecs on an Agent; the SupervisorStage fans them out in parallel with isolated contexts and aggregates results.
  • PermissionsPermissionPolicy + GuardedToolExecutor enforce allow/deny/ask rules (glob-matched) over tool calls, mirroring tiered permission modes.
  • HooksHookRegistry with PreToolUse / PostToolUse / PreLLM / PostLLM / OnPlan / OnCompact / OnError lifecycle hooks.
  • Auto compactionCompactionStage summarizes or drops old turns when the context window exceeds its token budget.
  • Memory consolidationMemoryConsolidationStage writes durable learnings (LEARNING: ...) back to the agent's MemoryStore.
  • MCP clientMCPClient + StdioTransport speak JSON-RPC over stdio; register_mcp_tools() wraps server tools as runtime Tools.
  • Background tasksBackgroundTaskRegistry for resumable async tasks (à la codex resume / Claude /tasks).
  • Skill scopingSkill gains paths / globs / disable_model_invocation.
  • Reasoning controlsProviderConfig.reasoning_effort / thinking_enabled / thinking_budget_tokens, forwarded when the provider supports reasoning.

Agentic workflow types (v0.8.0)

Beyond the single-agent loop, ai_runtime now ships the higher-level agent primitives that modern agentic frameworks (LangGraph, AutoGen, CrewAI, Reflexion) are built from. They compose with the existing Agent, Skill, Command, and pipeline machinery.

New agent types (ai_runtime.agents):

  • WorkflowAgent — a declarative DAG of WorkflowSteps. Steps run in dependency order with concurrent layers (bounded by max_concurrency); each step's output is exposed to downstream steps via {step_name} placeholders. Mirrors the graph/orchestration layer of LangGraph/AutoGen.
  • RouterAgent — intent-based dispatcher. Routes a message to a specialist Agent via explicit predicates, keyword matching, or an LLM classifier, falling back to a default_agent. Mirrors multi-agent orchestrators (CrewAI routers, AutoGen groups).
  • CriticAgent — Reflexion-style actor/critic loop. Produces a candidate with an actor, evaluates it with a critic (or a validator callable), and feeds critiques back for up to max_iterations until approved. Mirrors self-reflection / CriticGPT patterns.

New skill types (ai_runtime.skills):

  • RetrievalSkill — RAG-backed skill that injects retrieved context (via a ai_runtime.rag.Retriever) into the prompt before the LLM call.
  • GuardrailSkill — validation hook that gates model output via a guardrail callable, with reject / warn / rewrite on-fail policies. ComposedSkills aggregates retrieval context and applies guardrails automatically.

New command types (ai_runtime.commands):

  • Categorized slash commands: /review, /explain, /test, /workflow (plus the existing /compact, /context, /clear). Commands carry a category and args and support with_args() binding.

New streaming event:

  • WorkflowEvent — surfaces DAG/router/critic progress (queuedrunningcompleted/failed) so clients can render step status, and is recorded on ExecutionContext.metadata["workflow_steps"].

Built-in agents, skills & commands (v0.8.1)

The framework ships ready-to-use presets so you don't have to hand-roll common agents/skills/commands, and it uses them on itself to become more agentic.

Built-in agents (ai_runtime.agents.builtin): reviewer_agent, explainer_agent, tester_agent, summarizer_agent, plus critic_agent (Reflexion loop) and router_agent (intent router with default review/explain/test routes).

Built-in skills (ai_runtime.skills.builtin): self_review_skill, explain_code_skill, generate_tests_skill, summarize_skill, no_secrets_guardrail (secret-blocking GuardrailSkill), retrieval_skill (RAG-backed), and default_builtin_skills().

Built-in commands (ai_runtime.commands.builtin): reusable Command factories (review_command, explain_command, test_command, workflow_command, compact_command, context_command, clear_command) consumed by default_commands().

Self-agentic wiring — the runtime turns its own primitives on itself:

  • AgentRunner(self_review=True) runs a CriticAgent self-review (reflexion) pass over its own output before returning, in both run() and stream().
  • CompactionStage uses the framework's own summarizer_agent as an LLM-backed compaction summarizer (via make_agentic_compaction_summarizer) when a provider is present, instead of naively dropping old turns.

Integration surfaces (v0.7.0)

To embed ai_runtime in a web app, desktop app, VS Code, or CLI like Claude Code / Codex / Cursor / Copilot, the framework now ships:

  • Built-in toolsReadFileTool, WriteFileTool, EditFileTool, GlobTool, GrepTool, BashTool (scoped to a sandbox root via register_builtin_tools). These mirror the file/shell/grep primitives of the four tools.
  • Checkpoints / undoCheckpointManager snapshots files before edits so the UI can roll back (à la Cursor/Claude/Copilot checkpoints).
  • Agent config filesload_project_instructions() discovers .github/copilot-instructions.md, AGENTS.md, CLAUDE.md, .cursor/rules from the project root and folds them into the system prompt.
  • Transport-agnostic serverAgentServer exposes an Agent over a JSON-line protocol (AgentRequest / AgentResponse / StreamEvent) via serve_stdio() (VS Code / CLI) and serve_http() (web / desktop).
  • CLIai-runtime console script: ai-runtime "prompt" --model ..., --mode plan|stream, --serve (stdio), --http (HTTP server), --yolo.
  • Workspace / projectProject scopes memory, tools, permissions, and checkpoints to a project root (the unit you mount in a client).
  • Slash commandsCommandRegistry / default_commands() (/compact, /context, /clear) mirroring Copilot's / menu.
  • BYO providerProviderConfig.from_env() reads COPILOT_PROVIDER_* vars to point at Ollama / vLLM / any OpenAI-compatible endpoint.

Web application (web/)

A ready-to-run web UI + REST/WebSocket API that exercises the full framework:

export AI_RUNTIME_API_KEY=sk-...      # or AI_RUNTIME_PROVIDER/AI_RUNTIME_BASE_URL
./venv/bin/python web/run.py          # serves http://127.0.0.1:8787

Features exposed in the UI:

  • Streaming chat (WebSocket streams StreamEvents: text, tool calls, tool results, thinking, usage, completion)
  • Plan mode (/api/chat with mode: plan returns a reviewable Plan)
  • ProjectsProject scoping with auto-loaded instruction files + sandboxed built-in tools (Read/Write/Edit/Glob/Grep/Bash)
  • Checkpoints / undo — snapshot + restore files before agent edits
  • PermissionsPermissionPolicy allow/deny/ask rules per tool + params
  • Sub-agents — configure SubAgentSpecs; supervisor fans them out
  • Background tasks — submit/resume/cancel via BackgroundTaskRegistry
  • Slash commands/compact, /context, /clear, plus categorized agentic commands /review, /explain, /test, /workflow
  • MCP — connect a stdio MCP server and register its tools
  • Reasoning controlsreasoning_effort / thinking_enabled per request
  • Built-in agents panel — run reviewer / explainer / tester / summarizer / router / critic presets from the UI (GET /api/builtin/agents, POST /api/builtin/agents/run)
  • Built-in skills panel — compose self_review, explain_code, generate_tests, summarize, no_secrets_guardrail, retrieval skills into a session (GET /api/builtin/skills, POST /api/builtin/skills/apply)
  • Built-in command palette — invoke categorized commands from the UI (GET /api/builtin/commands, POST /api/builtin/commands/run). Both the composer slash-autocomplete and the command palette show a worked example for each command (e.g. /review def add(a, b): …).
  • Self-review toggle — enable the reflexion self-review pass per session (POST /api/builtin/self-review)

See web/README.md for the full API reference.

Built-in commands

The web UI (and CLI) ship a categorized / command menu. Type / in the composer to see the full list with live examples, or open the command palette (Cmd/Ctrl+K). Every command accepts free-form text after its name:

Command Category Example
/compact general /compact — summarize the conversation to free context
/context general /context — show a token/context breakdown
/clear general /clear — clear the conversation history
/review review /review def add(a, b):\n return a + b
/explain explain /explain the quicksort implementation in sort.py
/test test /test def divide(a, b):\n return a / b
/workflow workflow /workflow refactor Extract the parser into its own module
  • General commands (/compact, /context, /clear) run as side effects against the active session.
  • Agentic commands (/review, /explain, /test, /workflow) render their prompt template with the supplied text and run it through the session's own agent runner, streaming the result back into the chat.

Installation

From PyPI

pip install ai-runtime

From source

git clone https://github.com/KiritVaghela/ai-runtime.git
cd ai-runtime
pip install -e .

Quick Start

import os

from ai_runtime import AgentRuntime
from ai_runtime.conversation import ChatMessage
from ai_runtime.providers.enums import ProviderType

runtime = AgentRuntime.from_provider(
    provider=ProviderType.GROQ,
    model="groq/llama-3.3-70b-versatile",
    api_key=os.getenv("GROQ_API_KEY"),
)

session = runtime.create_session()

response = await session.chat(
    ChatMessage.user("Hello!")
)

print(response.message.content)

Streaming

from ai_runtime.conversation import ChatMessage
from ai_runtime.streaming import TextDeltaEvent

async for event in session.stream(
    ChatMessage.user("Write a haiku.")
):
    if isinstance(event, TextDeltaEvent):
        print(event.delta, end="")

Streaming with timeout

The runtime supports stream timeout propagation through ChatRequest.timeout. If a provider stream stalls, an ErrorEvent is emitted and the stream stops.

from ai_runtime.conversation import ChatMessage, ChatRequest
from ai_runtime.streaming import ErrorEvent, TextDeltaEvent

request = ChatRequest(
    messages=[ChatMessage.user("Generate a poem.")],
    timeout=10.0,
)

async for event in session.stream(request):
    if isinstance(event, TextDeltaEvent):
        print(event.delta, end="")
    elif isinstance(event, ErrorEvent):
        print("\nStream timed out or failed:", event.message)

Provider registry

The runtime uses a provider registry so you can register custom providers or replace the default provider implementation during tests.

from ai_runtime import AgentRuntime
from ai_runtime.providers import ProviderRegistry
from ai_runtime.providers.enums import ProviderType

registry = ProviderRegistry()
registry.register(ProviderType.OPENAI, MyCustomProvider)

runtime = AgentRuntime.from_provider(
    provider=ProviderType.OPENAI,
    model="gpt-4.1",
    api_key=os.getenv("OPENAI_API_KEY"),
    registry=registry,
)

Architecture

AgentRuntime
    └── Session
          └── ExecutionEngine
                └── ExecutionPipeline
                      ├── RequestBuilderStage
                      ├── LLMStage
                      └── ToolLoopStage
                            └── EventProcessor

Tools

Register tools with a ToolRegistry, wrap it in a ToolExecutor, and attach it to the session context. When the model requests a tool, the runtime executes it and feeds the result back automatically.

from ai_runtime.tools import ToolRegistry, ToolExecutor, FunctionTool
from ai_runtime.conversation import ChatMessage

registry = ToolRegistry()
registry.register(
    FunctionTool("get_weather", lambda ctx, inp: f"Weather in {inp['city']}: sunny")
)

session.context.tool_executor = ToolExecutor(registry)

response = await session.chat(
    ChatMessage.user("What is the weather in Paris?")
)
print(response.message.content)

Agents

An Agent bundles a provider, system prompt, tools, memory, and skills. AgentRunner drives execution (including the automatic tool-call loop) and persists conversation memory across turns.

from ai_runtime import AgentRuntime, Agent, AgentRunner
from ai_runtime.tools import ToolRegistry, ToolExecutor, FunctionTool

runtime = AgentRuntime.from_provider(
    provider=ProviderType.GROQ,
    model="llama-3.3-70b-versatile",
    api_key=os.getenv("GROQ_API_KEY"),
)

registry = ToolRegistry()
registry.register(FunctionTool("ping", lambda ctx, inp: "pong"))

agent = runtime.create_agent(
    name="helper",
    system_prompt="You are a helpful assistant.",
    tool_registry=registry,
)

runner = AgentRunner(agent)
response = await runner.run("Ping the tool for me.")
print(response.message.content)

Memory, RAG & Skills

  • ai_runtime.memoryMemoryStore, ConversationMemory, SemanticMemory.
  • ai_runtime.ragDocument, VectorStore, Retriever for retrieval.
  • ai_runtime.skillsSkill + SkillRegistry for composable behaviors.
  • ai_runtime.contextContextWindow for token budgeting/truncation.

Documentation

See CHANGELOG.md for version history.

Testing

pytest

Current status:

  • Provider registry with custom provider registration
  • Streaming response support with timeouts and completion events
  • Session-based execution and conversation accumulation
  • Provider integration coverage with test doubles

Roadmap

v0.3.x

  • Tool registry
  • Tool execution
  • Permission manager
  • Filesystem tools
  • Bash tools

Future

  • MCP support
  • Multi-agent workflows
  • Desktop and terminal integrations
  • AI Factory integration

Publishing to PyPI

  1. Update version in pyproject.toml.
  2. Build:
python -m build
  1. Check:
twine check dist/*
  1. Upload to TestPyPI:
twine upload --repository testpypi dist/*
  1. Upload to PyPI:
twine upload dist/*

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

forge_ai_runtime-0.8.2.tar.gz (68.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

forge_ai_runtime-0.8.2-py3-none-any.whl (94.3 kB view details)

Uploaded Python 3

File details

Details for the file forge_ai_runtime-0.8.2.tar.gz.

File metadata

  • Download URL: forge_ai_runtime-0.8.2.tar.gz
  • Upload date:
  • Size: 68.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.3

File hashes

Hashes for forge_ai_runtime-0.8.2.tar.gz
Algorithm Hash digest
SHA256 a8f962806d840f90374a845ff486363ac75409853b6c5b11f5faa2265be7769a
MD5 f771986cba402795630889408559b07c
BLAKE2b-256 4c13503e2b9c801bdd9dc81f8f54d4cddc026a827ef1192f1cc0f8366040d632

See more details on using hashes here.

File details

Details for the file forge_ai_runtime-0.8.2-py3-none-any.whl.

File metadata

File hashes

Hashes for forge_ai_runtime-0.8.2-py3-none-any.whl
Algorithm Hash digest
SHA256 b93dbd6cbb74acb5e452aad96d79f7a8c9b4198036713e86b0b6dc0141609f27
MD5 4471b12ece4722f57ab8b7efd496e7cd
BLAKE2b-256 e4656e8818d4e9222bfe11d152e12c164e915c67a2f0a8180fc1e28e546170c0

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page