Tools for developing and optimizing side effect free background agents
Project description
Weak Incentives (WINK)
WINK is an open source toolkit for developing and optimizing side-effect-free background agents (e.g., research, review, and coding agents) that translate end-user instructions into deterministic actions. The library is designed for human developers, but this README is written so automated coding agents (OpenAI Codex web/CLI, Claude Code, Cursor background agents, etc.) can be productive consumers of the API surface without depending on internals.
The public API centers on typed prompts, declarative tool contracts, inspectable sessions, and provider-agnostic adapters so you can keep determinism, observability, and safety front and center while iterating on agent behaviors.
This document is the package README published to PyPI. It focuses on the supported, public API surface you can depend on when wiring WINK into your own agent or orchestration system.
Architecture Overview
An agent harness in WINK wires together five core components:
-
PromptTemplate (
weakincentives.prompt.PromptTemplate): Immutable blueprint defining sections, tools, and structured output schema. Import fromweakincentives.prompt. -
Prompt (
weakincentives.Prompt): Wraps a template with parameter bindings and optional overrides. Pass params to the constructor or call.bind()to attach additional params before evaluation. -
Session (
weakincentives.runtime.Session): Event-driven state container that records all prompt renders, tool invocations, and custom state. State changes flow through pure reducers. Creates its ownDispatcherinternally (access viasession.dispatcher). -
ProviderAdapter (
OpenAIAdapter,LiteLLMAdapter): Bridges prompts to LLM providers. Calladapter.evaluate(prompt, session=session)to execute. -
Tool handlers: Functions with signature
(params: ParamsT, *, context: ToolContext) -> ToolResult[ResultT]that implement side effects when the model requests tool calls.
Data flow: PromptTemplate → Prompt (with bound params) →
adapter.evaluate() → LLM response → Tool calls dispatched → ToolResult
returned → Session updated → Events published
Installation
WINK targets Python 3.12+. Install the core library:
pip install weakincentives
Optional extras enable specific providers or tooling:
pip install "weakincentives[openai]"for the OpenAI adapter.pip install "weakincentives[litellm]"for the LiteLLM adapter.pip install "weakincentives[claude-agent-sdk]"for the Claude Agent SDK adapter.pip install "weakincentives[asteval]"to enable the sandboxed Python eval tool.pip install "weakincentives[podman]"for Podman-based sandboxes.pip install "weakincentives[wink]"for the demo CLI (wink).
The Claude Agent SDK adapter also requires the Claude Code CLI:
npm install -g @anthropic-ai/claude-code
Key Concepts
- Prompts (
weakincentives.prompt.Prompt): Composable, dataclass-driven blueprints that render deterministic model inputs while exposing tool contracts. - Tools (
weakincentives.prompt.Tool): Declarative descriptions of capabilities the model can invoke. Tools surface type-checked handlers and renderable schemas. - Sessions (
weakincentives.runtime.Session): Event-driven state containers that record every prompt render and tool invocation as immutable events. State changes flow through pure functions called "reducers", keeping state deterministic and inspectable. - Adapters (
weakincentives.adapters.ProviderAdapter): Bridges to model providers that negotiate tool calls and structured outputs without locking you into a single vendor. - Structured Output (
parse_structured_output): JSON-schema-backed parsing that turns model responses into typed dataclass instances. - Overrides (
PromptOverride,LocalPromptOverridesStore): Hash-based prompt overrides that let you refine prompt text safely in version control.
Public API
weakincentives: Curated entrypoints for building prompts, tools, and sessions.- Classes and functions:
Budget: Resource envelope combining time and token limits.BudgetExceededError: Exception raised when a budget limit is breached.BudgetTracker: Thread-safe tracker for cumulative token usage against a Budget.Deadline: Immutable value object describing a wall-clock expiration.DeadlineExceededError: Exception raised when a deadline is exceeded.FrozenDataclass: Decorator providing immutable dataclass utilities (copy, asdict, normalization).JSONValue: Type alias for JSON-compatible primitives, objects, and arrays.MarkdownSection: Render markdown content usingstring.Template.Prompt: Coordinate prompt sections and their parameter bindings.PromptResponse: Structured result emitted by an adapter evaluation.StructuredLogger: Logger adapter enforcing a minimal structured event schema.SupportsDataclass: Protocol satisfied by dataclass types and instances.Tool: Describe a callable tool exposed by prompt sections.ToolContext: Immutable container exposing prompt execution state to handlers.ToolHandler: Callable protocol implemented by tool handlers.ToolResult: Structured response emitted by a tool handler.ToolValidationError: Raised when tool parameters fail validation checks.WinkError: Base class for all weakincentives exceptions.TransactionError: Base for transaction errors (renamed fromExecutionStateError).RestoreFailedError: Failed to restore from snapshot.configure_logging: Configure the root logger with sensible defaults.get_logger: Return aStructuredLoggerscoped to a name.parse_structured_output: Parse a model response into the structured output type declared by the prompt.
- Modules:
adapters,cli,contrib,deadlines,debug,evals,formal,optimizers,prompt,runtime,serde,types.
- Classes and functions:
weakincentives.adapters: Provider integrations, configuration, and throttling primitives.- Constants:
CLAUDE_AGENT_SDK_ADAPTER_NAME,LITELLM_ADAPTER_NAME,OPENAI_ADAPTER_NAME. - Types:
AdapterName: Type alias for adapter names.PromptEvaluationError: Raised when evaluation against a provider fails.PromptResponse: Structured result emitted by an adapter evaluation.ProviderAdapter: Abstract base class describing the synchronous adapter contract.SessionProtocol: Protocol describing the session interface required by adapters.ThrottleError: Raised when a provider throttles a request.ThrottlePolicy: Configuration for automatic retry/backoff on throttling.
- Configuration:
LLMConfig: Base configuration for common LLM parameters (temperature, max_tokens, top_p, etc.).OpenAIClientConfig: Configuration for OpenAI client instantiation (api_key, base_url, timeout).OpenAIModelConfig: OpenAI-specific model configuration extending LLMConfig.LiteLLMClientConfig: Configuration for LiteLLM client instantiation.LiteLLMModelConfig: LiteLLM-specific model configuration extending LLMConfig.ClaudeAgentSDKClientConfig: Configuration for Claude Agent SDK (permission_mode, cwd, max_turns, isolation).ClaudeAgentSDKModelConfig: Claude Agent SDK model configuration extending LLMConfig.
- Factory:
new_throttle_policy: Factory for creating throttle policies. - Claude Agent SDK Isolation (
weakincentives.adapters.claude_agent_sdk):IsolationConfig: Hermetic isolation configuration (network_policy, sandbox, env, api_key, skills).NetworkPolicy: Network access constraints (allowed_domains). UseNetworkPolicy.no_network()for API-only.SandboxConfig: OS-level sandboxing (enabled, writable_paths, readable_paths, bash_auto_allow).EphemeralHome: Temporary HOME directory for isolation (auto-created when IsolationConfig is set).PermissionMode: Literal type for SDK permission levels ("default", "acceptEdits", "plan", "bypassPermissions").
- Claude Agent SDK Workspace (
weakincentives.adapters.claude_agent_sdk):ClaudeAgentWorkspaceSection: Section that materializes host files into a temp directory for SDK access.HostMount: Configuration for mounting host paths (host_path, mount_path, include_glob, exclude_glob, max_bytes).HostMountPreview: Preview of mount contents before materialization.WorkspaceBudgetExceededError: Raised when mount exceeds max_bytes.WorkspaceSecurityError: Raised when accessing paths outside allowed_host_roots.
- Constants:
weakincentives.prompt: Prompt authoring, rendering, and override helpers.- Authoring:
PromptTemplate: Immutable prompt blueprint with sections (import from here).MarkdownSection: Render markdown content usingstring.Template.Prompt: Coordinate prompt sections and their parameter bindings.RenderedPrompt: Result of rendering a prompt.ResourceRegistry: Typed container for runtime resources available to tool handlers. Supportsbuild(),merge(), and typedget().Section: Base class for prompt sections.SectionNode: Node in section tree.SectionPath: Path to a section.SectionVisibility: Enum controlling how a section is rendered (FULL,SUMMARY).Tool: Describe a callable tool exposed by prompt sections.ToolContext: Immutable container exposing prompt execution state to handlers.ToolExample: Representative invocation for a tool documenting inputs and outputs.ToolHandler: Callable protocol implemented by tool handlers.ToolRenderableResult: Protocol for tool results that can be rendered.ToolResult: Structured response emitted by a tool handler.SupportsDataclass: Protocol satisfied by dataclass types and instances.SupportsDataclassOrNone: Protocol for dataclass types or None.SupportsToolResult: Protocol for tool results.PromptProtocol: Protocol for prompts.PromptTemplateProtocol: Protocol for prompt templates.RenderedPromptProtocol: Protocol for rendered prompts.ProviderAdapterProtocol: Protocol for provider adapters.
- Composition:
OpenSectionsParams: Parameters for progressive disclosure of sections.
- Overrides:
LocalPromptOverridesStore: Store for local prompt overrides.PromptDescriptor: Descriptor for a prompt.PromptLike: Protocol for objects that look like prompts.PromptOverride: Override for a prompt.PromptOverridesError: Raised when prompt overrides fail.PromptOverridesStore: Protocol for prompt override stores.SectionDescriptor: Descriptor for a section.SectionOverride: Override for a section.ToolDescriptor: Descriptor for a tool.ToolOverride: Override for a tool.hash_json: Hash a JSON value.hash_text: Hash a text value.
- Structured output and validation:
OutputParseError: Raised when structured output parsing fails.StructuredOutputConfig: Configuration for structured output.parse_structured_output: Parse a model response into the structured output type declared by the prompt.PromptError: Base class for prompt errors.PromptRenderError: Raised when prompt rendering fails.PromptValidationError: Raised when prompt validation fails.VisibilityExpansionRequired: Raised when model requests expansion of summarized sections.
- Tool Policies (declarative constraints for safe tool invocation):
ToolPolicy: Protocol for tool invocation constraints.SequentialDependencyPolicy: Enforce unconditional tool ordering (e.g.,deployrequirestestandbuildto have run first).ReadBeforeWritePolicy: Prevent file overwrites without reading first; new files can be created freely.
- Enhanced override types:
TaskExampleOverride: Override entire task examples (objective, outcome, steps).TaskStepOverride: Modify individual steps within task examples.ToolExampleOverride: Override tool example descriptions, inputs, and outputs.
- Authoring:
weakincentives.runtime: Session, event, and orchestration primitives.- Logging:
StructuredLogger: Logger adapter enforcing a minimal structured event schema.configure_logging: Configure the root logger with sensible defaults.get_logger: Return aStructuredLoggerscoped to a name.
- Events:
Dispatcher: Interface for publishing events.HandlerFailure: Event emitted when a handler fails.InProcessDispatcher: Simple in-process event bus.PromptExecuted: Event emitted when a prompt is executed.PromptRendered: Event emitted when a prompt is rendered.DispatchResult: Result of publishing an event.TokenUsage: Token usage data from provider responses.ToolInvoked: Event emitted when a tool is invoked.
- Main loop orchestration:
MainLoop: Abstract base class for standardized agent workflow orchestration.MainLoopConfig: Configuration for default deadline/budget/resources.MainLoopRequest: Event requesting execution with optional constraints (budget, deadline, resources).MainLoopCompleted: Success event published via bus.MainLoopFailed: Failure event published via bus.
- Lifecycle management:
Runnable: Protocol for loops supporting graceful shutdown (run(),shutdown(),running,heartbeatproperties).ShutdownCoordinator: Singleton for SIGTERM/SIGINT handling and coordinated callback invocation.LoopGroup: Runs multiple loops in dedicated threads with coordinated shutdown, optional health endpoints, and watchdog monitoring.Heartbeat: Thread-safe timestamp tracker for worker liveness.Watchdog: Daemon thread that monitors heartbeats and terminates the process via SIGKILL when workers stall.HealthServer: Minimal HTTP server for Kubernetes liveness (/health/live) and readiness (/health/ready) probes.wait_until: Helper for polling predicates with timeout.
- Transactions:
CompositeSnapshot: Combines session + resource snapshots with JSON serialization.SnapshotMetadata: Context for when/why a snapshot was taken.PendingToolTracker: Thread-safe tracker for hook-based tool execution.PendingToolExecution: Metadata for an in-flight native tool execution.create_snapshot: Capture session and resource state.restore_snapshot: Restore session and resource state.tool_transaction: Context manager for automatic rollback on exception.
- Mailbox:
Mailbox: Protocol for point-to-point message delivery with visibility timeout and acknowledgment.Message: Message wrapper withacknowledge(),nack(),reply(), andreply_mailbox()methods.InMemoryMailbox: Single-process queues for testing and development.MailboxResolver: Protocol for backend-specific mailbox resolution.ReplyNotAvailableError: Raised when reply mailbox cannot be resolved.MailboxError,MailboxConnectionError,MailboxFullError,ReceiptHandleExpiredError,SerializationError: Error types.
- Session ledger:
DataEvent: Event carrying data.ReducerContext: Context for reducers.ReducerContextProtocol: Protocol for reducer context.ReducerEvent: Type alias for reducer events (dataclasses).Session: Immutable event ledger with Redux-like reducers.SessionProtocol: Protocol for sessions.SessionView: Read-only wrapper for Session (used in reducer contexts).Snapshot: Session snapshot.SnapshotProtocol: Protocol for session snapshots.SnapshotRestoreError: Raised when snapshot restoration fails.SnapshotSerializationError: Raised when snapshot serialization fails.TypedReducer: Reducer for typed state.append: Append an event to a session.build_reducer_context: Build a reducer context.iter_sessions_bottom_up: Iterate over sessions bottom-up.QueryBuilder: Fluent query builder for session slices.replace_latest: Replace the latest value in a session.replace_latest_by: Replace the latest value in a session by key.upsert_by: Upsert a value in a session by key.
- Slice storage:
SliceView[T]: Read-only protocol for accessing slice values.Slice[T]: Mutable protocol for slice storage operations.SliceFactory: Protocol for creating slices by type.SliceOp: Algebraic type for slice mutations (Append | Extend | Replace | Clear).Append[T]: Append a single value to a slice.Extend[T]: Extend a slice with multiple values.Replace[T]: Replace all values in a slice.Clear: Clear all values from a slice.InitializeSlice[T]: System event for initializing a slice.ClearSlice[T]: System event for clearing a slice.MemorySlice/MemorySliceView: In-memory tuple-backed storage.JsonlSlice/JsonlSliceView: JSONL file-backed persistent storage.
- Logging:
weakincentives.optimizers: Prompt optimization algorithms and utilities.- Protocol and base classes:
PromptOptimizer: Protocol for prompt optimization algorithms.BasePromptOptimizer: Abstract base class for prompt optimizers.OptimizerConfig: Base configuration dataclass withaccepts_overridesfield.
- Context and results:
OptimizationContext: Immutable context bundle with adapter, event bus, deadline, and overrides.OptimizationResult: Generic result container with response, artifact, and metadata.WorkspaceDigestResult: Result of workspace digest optimization.PersistenceScope: Enum for artifact storage location (SESSION,GLOBAL).
- Concrete implementations live in
weakincentives.contrib.optimizers. - Events:
OptimizationStarted: Event emitted when optimizer begins work.OptimizationCompleted: Event emitted on successful completion.OptimizationFailed: Event emitted when optimization raises exception.
- Protocol and base classes:
weakincentives.contrib: Optional, domain-specific tools and optimizers.weakincentives.contrib.tools: Planning (PlanningToolsSection,Plan), VFS (VfsToolsSection,VirtualFileSystem,HostMount), workspace digest (WorkspaceDigestSection), and (with extras)AstevalSection/PodmanSandboxSection.weakincentives.contrib.optimizers: Concrete optimizers (currentlyWorkspaceDigestOptimizer).weakincentives.contrib.mailbox: Distributed mailbox implementations.RedisMailbox: Distributed queues using Redis lists and sorted sets for visibility timeout management.RedisMailboxFactory: Factory for creating Redis-backed mailboxes.
weakincentives.evals: Evaluation framework for WINK agents.- Core types:
Sample[InputT, ExpectedT]: Single evaluation case with input and expected output.Dataset[InputT, ExpectedT]: Immutable collection of samples with JSONL loading.Score: Result from evaluating a single sample (0.0–1.0 with optional metadata).EvalResult: Pairs a sample with its output and score.EvalReport: Aggregated metrics across all samples.Evaluator: Type alias for evaluation functions.SessionEvaluator: Evaluator that receivesSessionViewfor inspection.
- Built-in evaluators:
exact_match: Strict equality comparison.contains: Substring matching withall_of/any_ofcombinators.llm_judge: LLM-as-Judge with categorical ratings.
- Session-aware evaluators (behavioral assertions):
tool_called(name): Assert a specific tool was invoked.tool_not_called(name): Assert a tool was never invoked.tool_call_count(name, min_count, max_count): Assert tool call count within bounds.all_tools_succeeded(): Assert no tool failures occurred.token_usage_under(max_tokens): Assert token budget was respected.slice_contains(T, predicate): Assert session slice contains matching value.
- Combinators:
all_of: All evaluators must pass (mean score).any_of: At least one must pass (max score).adapt: Convert standard evaluator to session-aware.
- Orchestration:
EvalLoop: Mailbox-driven evaluation orchestration.EvalRequest: Request wrapper for eval samples.submit_dataset: Helper to submit dataset samples to a mailbox.collect_results: Helper to collect eval results from a mailbox.
- LLM-as-Judge:
JudgeOutput: Structured output from judge prompt.JudgeParams: Parameters for judge prompt.Rating: Rating scale enum.RATING_VALUES,PASSING_RATINGS: Rating scale constants.JUDGE_TEMPLATE: Default judge prompt template.
- Core types:
weakincentives.formal: TLA+ formal specification support.@formal_spec: Decorator for embedding TLA+ metadata in Python classes.StateVar: Declares a TLA+ state variable with name and type.Action: Declares a TLA+ action with guard and state updates.ActionParameter: TLA+ action parameter with domain.Invariant: Declares a TLA+ invariant with ID, name, and predicate.FormalSpec: Container for all specification metadata withto_tla()andto_tla_config()methods.- Testing utilities (
weakincentives.formal.testing):extract_spec(cls): Extract TLA+ specification from decorated class.write_spec(spec, output_dir): Write TLA+ and config files to disk.model_check(spec): Run TLC model checker with 3-minute timeout.extract_and_verify(cls, output_dir): Combined extraction and verification.
weakincentives.resources: Resource injection with scoped lifecycles.Binding[T]: Associates protocol type with provider function and scope.Scope: Enum for instance lifetime (SINGLETON,TOOL_CALL,PROTOTYPE).ScopedResourceContext: Resolution context with dependency graph walking.ResourceRegistry: Container for bindings withof(),build(), andopen().ResourceResolver: Protocol for dependency resolution in providers.Provider[T]: Type alias for factory functions accepting a resolver.Closeable: Protocol for resources withclose()method (auto-cleanup).PostConstruct: Protocol for resources withpost_construct()hook.- Errors:
CircularDependencyError,DuplicateBindingError,ProviderError,UnboundResourceError,ResourceError.
weakincentives.skills: Agent Skills specification support.Skill: Skill metadata container with name, description, and content.SkillMount: Configuration for mounting a skill (source path, name override, enabled flag).SkillConfig: Configuration for skill mounting (mounts tuple, validate_on_mount flag).- Validation functions:
validate_skill(path): Comprehensive skill validation (structure, frontmatter, size limits).validate_skill_name(name): Validate skill name format.resolve_skill_name(path, override): Derive skill name from path or override.
- Constants:
MAX_SKILL_FILE_BYTES(1 MiB),MAX_SKILL_TOTAL_BYTES(10 MiB). - Errors:
SkillError(base),SkillValidationError,SkillNotFoundError,SkillMountError.
weakincentives.filesystem: Filesystem protocol and implementations.Filesystem: Protocol for file operations (read, write, list, glob, grep).SnapshotableFilesystem: Extended protocol with snapshot/restore support.HostFilesystem: Host filesystem implementation with git-based snapshots.- Note:
InMemoryFilesystemis inweakincentives.contrib.tools. - Binary operations:
read_bytes(path, *, offset=0, limit=None): Read file as bytes.write_bytes(path, content, *, mode="overwrite", create_parents=True): Write bytes to file.
- Result types:
ReadResult,ReadBytesResult,WriteResult,ListResult,GlobResult,GrepResult,GrepMatch. SnapshotError: Raised when snapshot/restore operations fail.
weakincentives.serde: Dataclass serialization helpers.clone: Clone a dataclass.dump: Dump a dataclass to JSON-compatible types.parse: Parse JSON-compatible types into a dataclass.schema: Generate a JSON schema for a dataclass.
weakincentives.types: JSON typing helpers.ContractResult: Result of a contract check.JSONArray: Type alias for JSON arrays.JSONArrayT: Type variable for JSON arrays.JSONObject: Type alias for JSON objects.JSONObjectT: Type variable for JSON objects.JSONValue: Type alias for JSON-compatible primitives, objects, and arrays.ParseableDataclassT: Type variable for parseable dataclasses.
weakincentives.dbc: Design-by-contract utilities.dbc_active: ReturnTruewhen DbC checks should run.dbc_enabled: Context manager to temporarily enable DbC.disable_dbc: Force DbC enforcement off.enable_dbc: Force DbC enforcement on.ensure: Validate postconditions once the callable returns or raises.invariant: Enforce invariants before and after public method calls.pure: Validate that the wrapped callable behaves like a pure function.require: Validate preconditions before invoking the wrapped callable.skip_invariant: Mark a method so invariants are not evaluated around it.
weakincentives.cli: CLI entrypoints, notably thewinkmodule.wink docs: Print bundled documentation (--referencefor API reference,--guidefor user guide,--changelogfor release history,--specsfor design specs).
Agent-facing operational notes
- WINK does not run unattended background agents by itself. It provides deterministic primitives that research/review/coding agents (or humans) drive explicitly via prompts, tool handlers, and adapters.
- Rendering is side-effect-free:
Prompt.render()produces a typedRenderedPromptcontaining message content, declared tools, and any structured-output schema, but does not contact providers until you pass it to an adapter. - Tool handlers are synchronous callables; use them to gate filesystem or
network access and to enforce policy before applying patches. Handlers
accept the typed params plus a keyword-only
contextand returnToolResultinstances. Use convenience constructorsToolResult.ok(value)for success orToolResult.error(message)for failures. PromptResponsecarries the prompt name, rendered text, and parsed output (when structured output is requested) so you can safely resume after partial failures or retries.- Sessions are immutable ledgers: reducers consume
PromptRendered,PromptExecuted, andToolInvokedevents that includeevent_id,session_id, timestamps, and provider metadata so you can join prompt and tool flows deterministically.
Quickstart Snippets
Minimal harness setup
from dataclasses import dataclass
from weakincentives import MarkdownSection, Prompt
from weakincentives.prompt import PromptTemplate
from weakincentives.adapters.openai import OpenAIAdapter
from weakincentives.runtime import Session
@dataclass(slots=True, frozen=True)
class TaskResponse:
summary: str
next_steps: list[str]
template = PromptTemplate[TaskResponse](
ns="myapp/tasks", key="task-agent", name="task-agent",
sections=[MarkdownSection(title="Instructions", template="...", key="instructions")],
)
session = Session() # Creates event bus internally (access via session.dispatcher)
adapter = OpenAIAdapter(model="gpt-4o-mini")
response = adapter.evaluate(Prompt(template), session=session)
result: TaskResponse = response.output
Defining a tool handler
from weakincentives import Tool, ToolContext, ToolResult
@dataclass(slots=True, frozen=True)
class PatchArgs:
path: str
diff: str
@dataclass(slots=True, frozen=True)
class PatchResult:
applied: bool
def apply_patch(params: PatchArgs, *, context: ToolContext) -> ToolResult[PatchResult]:
# context.session, context.deadline, context.dispatcher available
return ToolResult.ok(PatchResult(applied=True), message="Applied")
patch_tool = Tool[PatchArgs, PatchResult](
name="apply_patch", description="Apply a unified diff.", handler=apply_patch
)
Attaching tools to sections
section = MarkdownSection(
title="Instructions", template="Use apply_patch to edit files.",
key="instructions", tools=(patch_tool,),
)
Multi-turn with session state
response = adapter.evaluate(prompt, session=session)
session[TaskResponse].append(response.output) # Store result
later = session[TaskResponse].latest() # Retrieve later
Error handling
from weakincentives.adapters import PromptEvaluationError
try:
response = adapter.evaluate(prompt, session=session)
except PromptEvaluationError as exc:
print(exc.phase, exc.prompt_name) # "request"/"response"/"tool"/"budget"
Prompt Authoring (weakincentives.prompt)
PromptTemplate and Prompt
from dataclasses import dataclass
from typing import Any
from weakincentives.prompt import Prompt, PromptTemplate, MarkdownSection
@dataclass(frozen=True)
class MyParams:
value: str
template: PromptTemplate[Any] = PromptTemplate(
ns="myapp/agents",
key="my-agent",
name="my-agent",
sections=[MarkdownSection(title="Task", key="task", template="${value}")],
)
prompt = Prompt(template).bind(MyParams(value="...")) # Bind returns self
MarkdownSection with parameters
Use ${param} syntax for dynamic content:
section = MarkdownSection[TaskParams](
title="Task", template="Objective: ${objective}", key="task",
default_params=TaskParams(objective=""),
)
Tool.wrap helper
Creates a Tool using the function's __name__ and docstring:
def search(params: SearchParams, *, context: ToolContext) -> ToolResult[SearchResult]:
"""Search for content.""" # Becomes tool description
return ToolResult.ok(SearchResult(...), message="Done")
search_tool = Tool.wrap(search) # name="search", description="Search for content."
ToolContext fields
Available in tool handlers via context:
context.session- Current Sessioncontext.deadline- Optional Deadline (check withdeadline.remaining())context.resources-ResourceRegistryfor runtime servicescontext.filesystem- Sugar forcontext.resources.get(Filesystem)context.budget_tracker- Sugar forcontext.resources.get(BudgetTracker)context.prompt/context.rendered_prompt/context.adapter
ToolResult convenience constructors
# Success with typed value
ToolResult.ok(MyResult(...), message="Done") # success=True
# Failure with no value
ToolResult.error("File not found") # success=False, value=None
# Full form (when exclude_value_from_context is needed)
ToolResult(message="...", value=MyResult(...), success=True, exclude_value_from_context=False)
Additional components
parse_structured_output: Parse model response into typed dataclass- Overrides:
LocalPromptOverridesStorefor hash-scoped prompt refinements
Adapter Layer (weakincentives.adapters)
OpenAI and LiteLLM adapters
from weakincentives.adapters.openai import OpenAIAdapter
from weakincentives.adapters.litellm import LiteLLMAdapter
from weakincentives.adapters import OpenAIModelConfig, OpenAIClientConfig
# Basic usage
adapter = OpenAIAdapter(model="gpt-4o-mini") # Native JSON schema by default
# With typed configuration
adapter = OpenAIAdapter(
model="gpt-4o",
model_config=OpenAIModelConfig(temperature=0.7, max_tokens=4096),
client_config=OpenAIClientConfig(timeout=30.0),
)
# LiteLLM for multi-provider support
adapter = LiteLLMAdapter(model="claude-3-sonnet-20240229") # Any LiteLLM model
Claude Agent SDK adapter
The Claude Agent SDK adapter provides Claude's full agentic capabilities
through the official claude-agent-sdk package. Unlike OpenAI/LiteLLM
adapters, this runs Claude Code as a subprocess with native tools (Read,
Write, Bash, Glob, Grep).
from weakincentives.adapters.claude_agent_sdk import (
ClaudeAgentSDKAdapter,
ClaudeAgentSDKClientConfig,
ClaudeAgentWorkspaceSection,
HostMount,
IsolationConfig,
NetworkPolicy,
SandboxConfig,
)
# Create workspace section that materializes host files
workspace = ClaudeAgentWorkspaceSection(
session=session,
mounts=(
HostMount(
host_path="/path/to/project",
mount_path="project",
include_glob=("*.py", "*.md"),
exclude_glob=("*.pyc", "__pycache__/*"),
max_bytes=5_000_000,
),
),
allowed_host_roots=("/path/to",),
)
# Configure with hermetic isolation
adapter = ClaudeAgentSDKAdapter(
model="claude-sonnet-4-5-20250929",
client_config=ClaudeAgentSDKClientConfig(
permission_mode="bypassPermissions", # Auto-approve all tools
cwd=str(workspace.temp_dir), # Working directory
isolation=IsolationConfig(
network_policy=NetworkPolicy.no_network(), # API-only access
sandbox=SandboxConfig(
enabled=True,
readable_paths=(str(workspace.temp_dir),),
),
),
),
)
# Evaluate prompt (add workspace section to prompt template)
response = adapter.evaluate(prompt, session=session)
# Clean up temp directory when done
workspace.cleanup()
Isolation modes
# Minimal isolation (development)
adapter = ClaudeAgentSDKAdapter(model="claude-sonnet-4-5-20250929")
# Hermetic with specific domains (documentation access)
adapter = ClaudeAgentSDKAdapter(
client_config=ClaudeAgentSDKClientConfig(
isolation=IsolationConfig(
network_policy=NetworkPolicy(
allowed_domains=("docs.python.org", "peps.python.org"),
),
sandbox=SandboxConfig(enabled=True),
),
),
)
# Full lockdown (sensitive data)
adapter = ClaudeAgentSDKAdapter(
client_config=ClaudeAgentSDKClientConfig(
isolation=IsolationConfig(
network_policy=NetworkPolicy.no_network(),
sandbox=SandboxConfig(enabled=True),
include_host_env=False, # Don't inherit environment
),
),
)
MCP tool bridging
Custom weakincentives tools with handlers are automatically bridged to the SDK via MCP servers. The adapter creates an MCP server for tools from prompt sections:
from weakincentives.contrib.tools import PlanningToolsSection
# Planning tools are bridged as MCP tools
template = PromptTemplate[Result](
ns="app", key="agent",
sections=(
MarkdownSection(title="Task", template="...", key="task"),
PlanningToolsSection(session=session), # planning_* tools
workspace, # No tools - just provides workspace info
),
)
ProviderAdapter.evaluate() signature
response = adapter.evaluate(
prompt,
session=session,
deadline=deadline, # Optional timeout
budget=budget, # Token/time limits
budget_tracker=budget_tracker, # Shared tracker across evaluations
resources=resources, # Custom runtime resources
)
When resources is provided, it is merged with workspace resources (like
filesystem from prompt) to create the final resource registry. User-provided
resources take precedence over workspace defaults.
Progressive disclosure is managed via session state:
from weakincentives.prompt import SectionVisibility
from weakincentives.runtime.session import SetVisibilityOverride, VisibilityOverrides
session.dispatch(
SetVisibilityOverride(path=("details",), visibility=SectionVisibility.FULL)
)
PromptResponse fields
response = adapter.evaluate(prompt, session=session)
response.output # Parsed dataclass
response.text # Raw text
response.prompt_name # Prompt identifier
Throttling
from weakincentives.adapters import ThrottleError
try:
response = adapter.evaluate(prompt, session=session)
except ThrottleError as exc:
# exc.kind, exc.retry_after, exc.attempts
raise
Runtime & Events (weakincentives.runtime)
Session: Immutable event ledger with Redux-like reducers. Feed events in withappend(session, event)or convenience selectors likereplace_latestandupsert_by.Snapshot/SnapshotProtocolprovide persistence helpers.- Slice accessor API: Use
session[T]for reading and writing state slices. Methods includelatest(),all(),where(),seed(),clear(). - Reducers: Use
TypedReducerwithReducerContextto manage typed state slices through event-driven mutations. - Events:
PromptExecutedandToolInvokedevents capture every model exchange.Dispatcher/InProcessDispatcherpublish events to reducers.HandlerFailureandDispatchResultoffer backpressure and error reporting controls. - MainLoop: Abstract orchestrator for agent workflows with automatic visibility expansion handling and budget tracking.
- Logging:
configure_logging()wires a structured logger;get_loggerretrieves a module-level logger.StructuredLoggeris a protocol you can implement for custom sinks.
Contributed Tool Sections (weakincentives.contrib.tools)
VfsToolsSection - Sandboxed file operations
from weakincentives.contrib.tools import VfsToolsSection, HostMount, VfsPath
vfs = VfsToolsSection(
session=session,
mounts=(HostMount(host_path="./repo", mount_path=VfsPath(("workspace",)),
include_glob=("*.py",), exclude_glob=("*.pyc",), max_bytes=600_000),),
allowed_host_roots=(Path("."),),
)
Tools: ls, read_file, write_file, edit_file, glob, grep, rm
PlanningToolsSection - Multi-step planning
from weakincentives.contrib.tools import PlanningToolsSection, PlanningStrategy
planning = PlanningToolsSection(session=session, strategy=PlanningStrategy.PLAN_ACT_REFLECT)
Tools: planning_setup_plan, planning_read_plan, planning_add_step,
planning_update_step
WorkspaceDigestSection
digest = WorkspaceDigestSection(session=session) # Renders workspace summary
Session State Management
Query API
latest = session[MyType].latest()
all_items = session[MyType].all()
filtered = session[MyType].where(lambda x: x.status == "done")
exists = session[MyType].exists()
Dispatch API
All session mutations flow through a single dispatch() method:
# Dispatch event - routes to registered reducers
session.dispatch(AddStep(step="x"))
# Convenience methods dispatch events internally
session[Plan].seed(initial_plan) # → dispatches InitializeSlice
session[Plan].clear() # → dispatches ClearSlice
Mutation API
# Initialize or replace slice values (bypasses reducers)
session[Plan].seed(initial_plan)
# Append value using default reducer
session[Plan].append(new_step)
# Register reducer for custom event types
session[Plan].register(AddStep, my_reducer)
# Remove items from a slice
session[Plan].clear() # Clear all
session[Plan].clear(lambda p: p.done) # Clear matching
# Global operations
session.reset() # Clear all slices
session.restore(snapshot) # Restore from snapshot
Reducer helpers
from weakincentives.runtime import append_all, replace_latest
# These are reducer functions, not session mutators
# Use session[T].append() for direct mutations instead
In tool handlers
def handler(params, *, context: ToolContext) -> ToolResult:
plan = context.session[Plan].latest()
# Tool handlers can read session; adapters record ToolInvoked events
MainLoop Orchestration
MainLoop standardizes agent workflow orchestration: receive request, build
prompt, evaluate, handle visibility expansion, publish result. Implementations
define only the domain-specific factories.
Implementing a MainLoop
from weakincentives.runtime import MainLoop, MainLoopConfig, Session
from weakincentives.prompt import Prompt, PromptTemplate
class CodeReviewLoop(MainLoop[ReviewRequest, ReviewResult]):
def __init__(
self, *, adapter: ProviderAdapter[ReviewResult], bus: Dispatcher
) -> None:
super().__init__(
adapter=adapter,
bus=bus,
config=MainLoopConfig(budget=Budget(max_total_tokens=50000)),
)
self._template = PromptTemplate[ReviewResult](
ns="reviews", key="code-review", sections=[...],
)
def prepare(self, request: ReviewRequest) -> tuple[Prompt[ReviewResult], Session]:
prompt = Prompt(self._template).bind(ReviewParams.from_request(request))
session = Session(bus=self._bus, tags={"loop": "code-review"})
return prompt, session
Direct execution
loop = CodeReviewLoop(adapter=adapter, bus=bus)
response, session = loop.execute(ReviewRequest(...))
Mailbox-driven execution
from weakincentives.runtime import InMemoryMailbox, MainLoopRequest, MainLoopResult
# Create request/response mailboxes
requests: InMemoryMailbox[MainLoopRequest, MainLoopResult] = InMemoryMailbox(
name="requests"
)
responses: InMemoryMailbox[MainLoopResult, None] = InMemoryMailbox(name="responses")
# Send request with reply routing
requests.send(
MainLoopRequest(
request=ReviewRequest(...),
budget=Budget(max_total_tokens=10000), # Overrides config default
deadline=Deadline(expires_at=datetime.now(UTC) + timedelta(minutes=5)),
),
reply_to="responses",
)
# MainLoop processes from requests mailbox and replies via msg.reply()
# Results arrive in responses mailbox
for msg in responses.receive():
result = msg.body
if result.error:
print(f"Failed: {result.error}")
else:
print(f"Done: {result.output}")
msg.acknowledge()
Visibility expansion handling
MainLoop automatically handles VisibilityExpansionRequired exceptions by
accumulating visibility overrides and retrying evaluation. A shared
BudgetTracker enforces limits cumulatively across retries.
Event Subscription
from weakincentives.runtime import (
PromptRendered,
PromptExecuted,
ToolInvoked,
Session,
)
session = Session()
# Define typed handlers
def on_tool_invoked(event: object) -> None:
if isinstance(event, ToolInvoked):
print(event.name)
def on_prompt_executed(event: object) -> None:
if isinstance(event, PromptExecuted):
print(event.usage)
def on_prompt_rendered(event: object) -> None:
print(event)
# Subscribe to events
session.dispatcher.subscribe(ToolInvoked, on_tool_invoked)
session.dispatcher.subscribe(PromptExecuted, on_prompt_executed)
# Unsubscribe handler (returns True if found and removed)
session.dispatcher.subscribe(PromptRendered, on_prompt_rendered)
session.dispatcher.unsubscribe(PromptRendered, on_prompt_rendered)
Session Snapshots
# Capture session state
snapshot = session.snapshot()
# Restore from snapshot
session.restore(snapshot)
# Serialize for persistence
snapshot_json = snapshot.to_json()
restored = Snapshot.from_json(snapshot_json)
Deadlines
from datetime import datetime, timedelta, UTC
from weakincentives import Deadline
deadline = Deadline(expires_at=datetime.now(UTC) + timedelta(minutes=5))
# In handlers: if deadline.remaining() <= timedelta(0): ...
Budgets
Budgets combine time and token limits into a single resource envelope:
from datetime import datetime, timedelta, UTC
from weakincentives import Budget, BudgetTracker, BudgetExceededError, Deadline
# Create a budget with deadline and token limits
budget = Budget(
deadline=Deadline(expires_at=datetime.now(UTC) + timedelta(minutes=10)),
max_total_tokens=100_000,
max_input_tokens=80_000,
max_output_tokens=20_000,
)
# Track usage across evaluations
tracker = BudgetTracker(budget=budget)
tracker.record_cumulative("eval-1", usage) # Record TokenUsage from response
tracker.check() # Raises BudgetExceededError if any limit breached
Mailbox (weakincentives.runtime.mailbox)
Point-to-point message delivery with visibility timeout and acknowledgment:
from weakincentives.runtime import InMemoryMailbox, Message
mailbox: Mailbox[WorkRequest] = InMemoryMailbox()
message_id = mailbox.send(WorkRequest(task="analyze"))
# Receive with visibility timeout (message hidden from other consumers)
messages = mailbox.receive(visibility_timeout=30, wait_time_seconds=5)
for msg in messages:
process(msg.body)
msg.acknowledge() # Remove from queue
Reply-to routing
Workers can send results to dynamic destinations derived from incoming messages:
from weakincentives.runtime.mailbox import InMemoryMailbox, RegistryResolver
# Setup resolver mapping identifiers to mailboxes
responses = InMemoryMailbox(name="client-responses")
resolver = RegistryResolver({"client-123": responses})
requests = InMemoryMailbox(name="requests", reply_resolver=resolver)
# Client sends with reply destination
requests.send(body=Request(...), reply_to="client-123")
# Worker replies - resolver routes to correct mailbox
for msg in requests.receive():
msg.reply(process(msg.body)) # Resolves "client-123" → responses mailbox
msg.acknowledge()
For dynamic mailbox creation (e.g., per-request reply queues), use
CompositeResolver with a MailboxFactory. See specs/MAILBOX_RESOLVER.md.
Serialization (weakincentives.serde)
from weakincentives.serde import dump, parse, schema, clone
data = dump(my_dataclass) # To JSON-compatible dict
obj = parse(MyDataclass, data) # From dict
json_schema = schema(MyDataclass) # JSON schema
copy = clone(my_dataclass) # Deep clone
Additional Patterns
Hierarchical sections
root = MarkdownSection(title="Root", template="...", key="root", children=[
MarkdownSection(title="Child", template="...", key="child"),
])
Tool examples
tool = Tool[P, R](name="search", description="...", handler=h, examples=(
ToolExample(description="Find X", input=P(...), output=R(...)),
))
Design-by-contract
from weakincentives.dbc import require, ensure
@require(lambda params: params.query) # Predicate validates params.query is truthy
@ensure(lambda result: result is not None)
def handler(params, *, context): ...
Section visibility (progressive disclosure)
from weakincentives.prompt import MarkdownSection, SectionVisibility
# Section with summary for progressive disclosure
section = MarkdownSection(
title="Details",
template="Full detailed content...",
key="details",
summary="Brief summary of the section",
visibility=SectionVisibility.SUMMARY, # Show summary by default
)
Prompt optimizers
from weakincentives.contrib.optimizers import WorkspaceDigestOptimizer
from weakincentives.optimizers import OptimizationContext, PersistenceScope
context = OptimizationContext(
adapter=adapter,
dispatcher=session.dispatcher,
overrides_store=overrides_store,
)
optimizer = WorkspaceDigestOptimizer(context, store_scope=PersistenceScope.SESSION)
result = optimizer.optimize(prompt, session=session)
# result.digest contains the workspace summary
Resource injection
Pass custom runtime resources to prompts and MainLoop for cleaner, more testable tool handlers:
from weakincentives.resources import Binding, Scope
from myapp.http import HTTPClient, Config
# Simple case: pre-constructed instances (pass a mapping)
http_client = HTTPClient(base_url="https://api.example.com")
prompt = Prompt(template).bind(params, resources={HTTPClient: http_client})
# Advanced: lazy construction with dependencies (Binding objects in mapping)
prompt = Prompt(template).bind(params, resources={
Config: Binding(Config, lambda r: Config.from_env()),
HTTPClient: Binding(HTTPClient, lambda r: HTTPClient(r.get(Config).url)),
Tracer: Binding(Tracer, lambda r: Tracer(), scope=Scope.TOOL_CALL),
})
# Use prompt.resources context manager for lifecycle
with prompt.resources:
response = adapter.evaluate(prompt, session=session)
# Or configure at MainLoop level (also a mapping)
config = MainLoopConfig(resources={
Config: Binding(Config, lambda r: Config.from_env()),
HTTPClient: Binding(HTTPClient, lambda r: HTTPClient(r.get(Config).url)),
})
loop = MyLoop(adapter=adapter, requests=requests, config=config)
In tool handlers, access resources via the typed registry:
def my_handler(params: Params, *, context: ToolContext) -> ToolResult[Result]:
# Access via typed registry
client = context.resources.get(HTTPClient)
# Common resources have sugar properties
fs = context.filesystem # context.resources.get(Filesystem)
budget = context.budget_tracker # context.resources.get(BudgetTracker)
...
ResourceRegistry.merge() combines registries with the second taking
precedence on conflicts, enabling layered resource injection where
caller-provided resources override workspace defaults.
CLI
pip install "weakincentives[wink]"
wink --help
Example
See code_reviewer_example.py in the repository for a complete production
harness demonstrating all patterns: structured types, tool handlers, built-in
sections (VFS, Planning), event subscription, and prompt overrides.
Versioning & Stability
- Public APIs are the objects exported from
weakincentivesand the submodules documented above. - Adapters are optional; include only the extras you need.
- Keep
StructuredOutputConfig, tool schemas, and overrides in version control so your agents remain deterministic and auditable.
License
Apache License 2.0. See LICENSE for details.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file weakincentives-0.19.0.tar.gz.
File metadata
- Download URL: weakincentives-0.19.0.tar.gz
- Upload date:
- Size: 8.6 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
944f5e2992a99ae71826086892596168863d826e3be077b6e0ab34f126ab418d
|
|
| MD5 |
b3d4fcacb8b2d7154b6d04c1d0b25163
|
|
| BLAKE2b-256 |
390b5fa5516a96aad9810928e16d5d8e653026b6dcf61e11551692d346b688f8
|
Provenance
The following attestation bundles were made for weakincentives-0.19.0.tar.gz:
Publisher:
release.yml on weakincentives/weakincentives
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
weakincentives-0.19.0.tar.gz -
Subject digest:
944f5e2992a99ae71826086892596168863d826e3be077b6e0ab34f126ab418d - Sigstore transparency entry: 803756116
- Sigstore integration time:
-
Permalink:
weakincentives/weakincentives@312bfde1e6773f2ed738ad3831235e511eba3c0e -
Branch / Tag:
refs/tags/v0.19.0 - Owner: https://github.com/weakincentives
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@312bfde1e6773f2ed738ad3831235e511eba3c0e -
Trigger Event:
release
-
Statement type:
File details
Details for the file weakincentives-0.19.0-py3-none-any.whl.
File metadata
- Download URL: weakincentives-0.19.0-py3-none-any.whl
- Upload date:
- Size: 660.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d081fe511b90b91c02e124f8b6c3aabc735dde0eb2e8fcea7d94fed359412d6d
|
|
| MD5 |
bce4450734d430d44378214ff7e218ed
|
|
| BLAKE2b-256 |
c892edf94fc5056e648e38aca05df5e7d5b8743d98ba5ddd2520a28486c07f63
|
Provenance
The following attestation bundles were made for weakincentives-0.19.0-py3-none-any.whl:
Publisher:
release.yml on weakincentives/weakincentives
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
weakincentives-0.19.0-py3-none-any.whl -
Subject digest:
d081fe511b90b91c02e124f8b6c3aabc735dde0eb2e8fcea7d94fed359412d6d - Sigstore transparency entry: 803756179
- Sigstore integration time:
-
Permalink:
weakincentives/weakincentives@312bfde1e6773f2ed738ad3831235e511eba3c0e -
Branch / Tag:
refs/tags/v0.19.0 - Owner: https://github.com/weakincentives
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@312bfde1e6773f2ed738ad3831235e511eba3c0e -
Trigger Event:
release
-
Statement type: