Skip to main content

Python SDK for llm4agents.com — gasless AI agent infrastructure

Project description

llm4agents-sdk

PyPI version Python License: MIT

Python SDK for llm4agents.com — gasless AI agent infrastructure. Chat completions, wallet management, gasless stablecoin transfers, and MCP-powered tools through a single async client.

Install

pip install llm4agents-sdk

eth-account is bundled as a dependency for gasless transfers (no extra install needed).

Get an API Key

  1. Go to api.llm4agents.com/docs
  2. Register your agent to receive a key in the format sk-proxy-...
  3. Pass it to the client constructor

Quick Start

import asyncio
from llm4agents import LLM4AgentsClient

async def main():
    client = LLM4AgentsClient(api_key="sk-proxy-...")

    # Chat completion
    response = await client.chat.completions.create({
        "model": "anthropic/claude-sonnet-4",
        "messages": [{"role": "user", "content": "Hello"}],
    })
    print(response["choices"][0]["message"]["content"])

    # Conversation with MCP tools
    conv = client.chat.conversation({
        "model": "anthropic/claude-sonnet-4",
        "system": "You are a research assistant",
        "tools": client.tools,
    })
    answer = await conv.say("Search for Bitcoin news and summarize the top 3")
    print(answer.content)

    # Gasless stablecoin transfer
    result = await client.transfer.send({
        "chain": "polygon", "token": "USDC",
        "to": "0xRecipient...", "amount": "10.50",
        "private_key": "0x...",
    })
    print(result["tx_hash"], result["explorer_url"])

asyncio.run(main())

Chat

Completions

# Non-streaming
response = await client.chat.completions.create({
    "model": "anthropic/claude-sonnet-4",
    "messages": [{"role": "user", "content": "Hello"}],
})
print(response["choices"][0]["message"]["content"])

# Streaming
from llm4agents import FinalUsage

def on_final_usage(usage: FinalUsage) -> None:
    print(f"\n[usage] prompt={usage.prompt_tokens} completion={usage.completion_tokens} total={usage.total_tokens}")

stream = await client.chat.completions.create(
    {
        "model": "anthropic/claude-sonnet-4",
        "messages": [{"role": "user", "content": "Count to 10"}],
        "stream": True,
    },
    on_final_usage=on_final_usage,   # fires once after the stream ends
)
async for chunk in stream:
    print(chunk.choices[0]["delta"].get("content", ""), end="", flush=True)

# With extended thinking
response = await client.chat.completions.create({
    "model": "anthropic/claude-sonnet-4",
    "messages": [{"role": "user", "content": "Solve step by step: 47 * 83"}],
    "reasoning": True,
    "include_reasoning": True,
})

# Vision (multimodal) input
analysis = await client.chat.completions.create({
    "model": "openai/gpt-4o",
    "messages": [{
        "role": "user",
        "content": [
            {"type": "text", "text": "What is in this image?"},
            {"type": "image_url", "image_url": {"url": "https://example.com/image.png"}},
        ],
    }],
})

# Model fallbacks — try primary, fall back to secondary if it fails
response = await client.chat.completions.create({
    "models": ["anthropic/claude-sonnet-4", "openai/gpt-4o"],
    "messages": [{"role": "user", "content": "Hello"}],
})
# X-Model-Used response header indicates which model actually responded

# Force/restrict tool use
response = await client.chat.completions.create({
    "model": "openai/gpt-4o",
    "messages": [...],
    "tools": await client.tools.fetch_definitions(),
    "tool_choice": "required",  # "none" | "auto" | "required" | {"type": "function", "function": {"name": "..."}}
})

Conversation with Tools

from llm4agents import LLM4AgentsError, ResponseMeta

def on_round_meta(meta: ResponseMeta) -> None:
    cents = meta.cost_usd_cents or 0
    bal = meta.balance_remaining_cents or 0
    print(f"Round cost: ${cents / 100:.4f}, balance: ${bal / 100:.2f}")

conv = client.chat.conversation({
    "model": "anthropic/claude-sonnet-4",
    "system": "You are a research assistant",
    "tools": client.tools,
    "history": [],                # optional: rehydrate from persisted messages
    "on_tool_call": lambda name, args: print(f"Calling {name}...") or True,
    "on_tool_result": lambda name, result: print(f"{name}: {result.text[:80]}"),
    "on_round_meta": on_round_meta,
    "on_tools_ignored": lambda model: print(f"WARN: {model} ignored tools"),
    "enable_prompt_tool_fallback": True,   # retry round 0 with tools in prompt
    "max_tool_rounds": 5,         # default 10
    "tool_choice": "required",    # "auto" (default) | "required" | "none" | {"type": "function", "function": {"name": "..."}}
})

# Single turn
answer = await conv.say("Search for Bitcoin news and summarize the top 3")
print(answer.content)
for call in answer.tool_calls:
    print(f"  {call['name']}{call['result'].text[:60]}")

# Streaming conversation
async for event in conv.stream("Now find the current price"):
    if event["type"] == "text":
        print(event["content"], end="", flush=True)
    elif event["type"] == "reasoning":
        print(f"<think>{event['content']}</think>", end="")
    elif event["type"] == "meta":
        print(f"\n[meta] request_id={event['meta'].request_id}")
    elif event["type"] == "tool_call":
        print(f"\n[tool] {event['name']}({event['args']})")
    elif event["type"] == "tool_result":
        print(f"[done] {event['result'].text[:60]}")
    elif event["type"] == "done":
        print(f"\nUsage: {event['response']['usage']}")

# History management
messages = conv.messages    # list[ChatMessage] — JSON-serializable
conv.clear()                # reset to empty, keeps system prompt and options
branch = conv.fork()        # copy history + all callbacks into a new Conversation

on_tool_call receives (name: str, args: dict) and should return True to proceed or False to cancel the tool call.

on_round_meta fires after each round with a ResponseMeta containing cost_usd_cents, balance_remaining_cents, tokens_input, tokens_output, tokens_reasoning, model_used, and request_id parsed from response headers.

on_tools_ignored(model) fires once if you pass tools but the model returns no tool calls on the first round — useful for detecting models without native function calling.

tool_choice controls tool selection on round 1 only, then reverts to "auto" for every subsequent round. This is the agent-routing pattern: force the model to use a tool on the first turn (so it can't quietly emit JSON-as-text instead), then let it summarize the tool result naturally on the wrap-up turn. Setting "required" on every round forces a tool call forever and the conversation hits max_tool_rounds. Accepts "auto" (default), "required", "none", or {"type": "function", "function": {"name": "..."}} to force a specific named tool.

enable_prompt_tool_fallback=True adds an automatic recovery path for those models. When round 0 ignores tools, the SDK retries the round once with the tool definitions injected into the system prompt and asks the model to emit <tool_call>{"name":"...","arguments":{...}}</tool_call> blocks. Parsed blocks are executed exactly like native tool calls, so the rest of the loop is unchanged. In stream() you'll see a {"type": "fallback", "reason": "tools_ignored", "model": ...} event before the fallback round runs (the fallback round itself is non-streamed). If the fallback also returns no blocks, its plain-text reply is returned as the answer.

ConversationResponse.usage includes prompt_tokens, completion_tokens, total_tokens, and (when the upstream model exposes them) reasoning_tokens accumulated across every round, including any fallback round.

To restore a conversation from a previous session:

conv = client.chat.conversation({
    "model": "anthropic/claude-sonnet-4",
    "history": saved_messages,   # rehydrate from your store
})

McpToolResult

All tool calls return an McpToolResult instead of a plain string:

result = await client.tools.scraper.fetch_html("https://example.com")

result.text          # str — joined text from all text parts (convenience)
str(result)          # same as result.text — backward-compatible
result.content       # tuple[McpContent, ...] — typed content parts

for part in result.content:
    if part.type == "text":
        print(part.text)
    elif part.type == "image":
        print(part.mimeType, len(part.data))    # base64-encoded image
    elif part.type == "resource":
        print(part.uri, part.text)

The transport auto-normalizes raw MCP responses: snake_case mime_type is aliased to mimeType, imageBase64/pngBase64 keys are mapped to data, MIME types are sniffed from base64 magic bytes when missing, and JSON-wrapped image/PDF payloads embedded inside text blocks (e.g. {"imageBase64": "...", "mimeType": "image/png"}) are promoted to typed McpImageContent / McpResourceContent automatically.

Agents

# Register a new agent — call before you have an API key
client = LLM4AgentsClient(api_key="")   # empty key is fine for registration
reg = await client.agents.register("My Agent")
# The returned api_key is shown only once — save it immediately
print(reg.api_key)        # sk-proxy-...
print(reg.uuid)

Wallets

# Generate a deposit wallet
wallet = await client.wallets.generate({"chain": "polygon", "token": "USDC"})
print(wallet["address"])

# Check balance
balance = await client.wallets.balance()
print(balance["available_usd"])
for w in balance["wallets"]:
    print(f'{w["chain"]}/{w["token"]}: ${w["available_usd"]}')

# Transaction history
txs = await client.wallets.transactions({"limit": 20, "type": "deposit"})
for tx in txs["transactions"]:
    print(f'{tx["type"]}: ${tx["amount_usd_cents"] / 100}{tx["description"]}')

type filter accepts "deposit", "usage", "refund", or "gas_sponsored".

Gasless Transfers

# One-call convenience
result = await client.transfer.send({
    "chain": "polygon", "token": "USDC",
    "to": "0xRecipient...", "amount": "10.50",
    "private_key": "0x...",
})
print(result["tx_hash"], result["explorer_url"])

# Two-step — inspect the fee before committing
quote = await client.transfer.quote({
    "chain": "polygon", "token": "USDC",
    "from_": "0xSender...", "to": "0xRecipient...", "amount": "10.50",
})
print(f"Fee: {quote.fee_formatted}")
print(f"Forwarder: {quote.forwarder_address}")

result = await client.transfer.submit(quote, "0xPrivateKey...")
print(result.tx_hash)

x402 Walk-up Payment

The proxy supports the x402 protocol for per-request stablecoin payments on POST /v1/chat/completions. Instead of pre-funding an agent account, the client signs an EIP-3009 TransferWithAuthorization for USDC on Base / Base-Sepolia, attaches it as an X-PAYMENT header, and the proxy settles on-chain after the response is delivered.

Scope: this SDK signs x402 payments going OUT. The SDK builds and signs X-PAYMENT headers so your code can pay any x402-compatible server (the llm4agents API, an x402engine endpoint, or any third party). It does not include server-side helpers: there is no verify_payment, no settle_payment, no require_payment middleware. If your agent needs to receive x402 payments (run its own paywall), use a server library directly — x402-fastapi for FastAPI, x402-flask for Flask, coinbase-x402 for any framework, or the reference servers for other stacks. See Roadmap below for the server-side direction.

Two modes are mutually exclusive — pick one at construction time:

Mode Set via Required Use when
Bearer (default) omit payment or PaymentConfig(mode="bearer") api_key You have an agent and a pre-funded balance
x402 walk-up payment=PaymentConfig(mode="x402", signer=...) signer (eth_account or custom) You want one-shot calls billed per-request from a wallet, no agent registration

Bearer vs x402 — at a glance

from llm4agents import LLM4AgentsClient, PaymentConfig, eth_account_to_signer
from eth_account import Account

# Bearer (existing) — pre-funded agent
bearer = LLM4AgentsClient(api_key="sk-proxy-...")

# x402 walk-up — pay per call from a wallet
account = Account.from_key("0xYOUR_PRIVATE_KEY")
x402 = LLM4AgentsClient(
    api_key="",  # ignored in x402 mode
    payment=PaymentConfig(
        mode="x402",
        signer=eth_account_to_signer(account),
        network="base-sepolia",        # or "base" for mainnet
    ),
)

# Same API surface — the SDK probes the proxy for a 402, signs an
# EIP-3009 authorization, and retries with X-PAYMENT automatically.
res = await x402.chat.completions.create({
    "model": "openai/gpt-4o-mini",
    "messages": [{"role": "user", "content": "Hello"}],
})

Custom signers (no eth_account dependency)

eth_account is already a hard dependency of the SDK (used by gasless transfers), so eth_account_to_signer is the default path. If you want to plug in a hardware wallet, KMS-backed key, or WalletConnect, implement the Signer Protocol directly — the SDK depends on the Protocol, not on eth_account:

from llm4agents import Signer

class HsmSigner:
    address = "0xYourAddress..."

    async def sign_typed_data(self, *, domain, types, primary_type, message):
        # Defer to your HSM / KMS / hardware wallet. Return a 0x-prefixed
        # 65-byte signature.
        return "0x..."

client = LLM4AgentsClient(
    api_key="",
    payment=PaymentConfig(mode="x402", signer=HsmSigner(), network="base"),
)

sign_typed_data may be sync or async — the SDK awaits the result if it's a coroutine.

Streaming receipts

x402-mode streaming responses end with a trailing SSE event after [DONE] containing the on-chain settlement receipt. The Conversation.stream() helper surfaces this as a typed event yielded BEFORE the matching done event:

conv = x402.chat.conversation({"model": "openai/gpt-4o-mini"})
async for ev in conv.stream("Tell me a joke"):
    if ev["type"] == "text":
        print(ev["content"], end="", flush=True)
    elif ev["type"] == "x402_receipt":
        print(f"\nsettled: {ev['transaction']} on {ev['network']}")
    elif ev["type"] == "done":
        print(f"\n{ev['response']['usage']}")

For the lower-level chat.completions.create() API, pass on_x402_receipt:

def handle_receipt(receipt):
    print(f"settled {receipt.amount} on {receipt.network}: {receipt.transaction}")

stream = await x402.chat.completions.create(
    {"model": "openai/gpt-4o-mini", "messages": [...], "stream": True},
    on_x402_receipt=handle_receipt,
)
async for chunk in stream:
    ...

Lower-level helpers — client.x402

For advanced use cases (signing without sending, inspecting the 402 response shape, batch signing), the client.x402 namespace exposes the building blocks:

# Probe the proxy and get the typed PaymentRequirements
requirements = await x402.x402.probe()
print(requirements.max_amount_required, requirements.network)

# Probe + sign in one call — returns SignedPayment(payment_payload, encoded_header, requirements)
signed = await x402.x402.sign()
#   signed.encoded_header     → base64-encoded X-PAYMENT value
#   signed.payment_payload    → the parsed PaymentPayload (typed)
#   signed.requirements       → the proxy-advertised requirements the signature is bound to

# Sign against caller-supplied requirements (no HTTP) — useful for testing
signed2 = await x402.x402.sign_from_requirements(requirements)

Error handling

When the proxy rejects payment (signature invalid, nonce reused, etc.) the SDK raises X402PaymentRequiredError carrying the typed requirements so the caller can re-sign with a different amount or network:

from llm4agents import X402PaymentRequiredError

try:
    await x402.chat.completions.create({...})
except X402PaymentRequiredError as err:
    print("Payment rejected. accepted offers:", err.payment_requirements)
    print("x402 version:", err.x402_version)

Networks: "base" (mainnet, USDC 0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913) and "base-sepolia" (testnet, USDC 0x036CbD53842c5426634e7929541eC2318f3dCF7e) are currently supported. The USDC EIP-712 domain name differs between them (USD Coin vs USDC); eth_account_to_signer handles this automatically.

Endpoints accepting x402 (signed per-call USDC):

  • POST /v1/chat/completions — chat with any model (per-token signed upper bound)
  • POST /v1/scrape/{markdown,fetch_html,links,screenshot,pdf,extract} — one-shot scraping
  • POST /v1/search/{google,news,maps,batch} — Google search (Serper)
  • POST /v1/image/{generate,edit,analyze} — image generation / edit / vision

Per-call x402 prices are seeded ~10% below x402engine.app reference rates (e.g. scrape markdown ~$0.0045, screenshot ~$0.009, image gen ~$0.0135-$0.045). Prices are admin-editable from the operator panel without redeploy.

Browser sessions (session_*) and other endpoints (/v1/embeddings, /api/v1/wallets/*, etc.) stay Bearer-only — sessions are pre-deposit by design. Using x402 mode on a non-allowed path raises LLM4AgentsError with code x402_payment_required and a clear message.

REST scrape / search / image with x402

The same payment=PaymentConfig(mode="x402", signer=..., network=...) client config that works for chat completions also works for the MCP REST surface:

import httpx
from eth_account import Account
from llm4agents import LLM4AgentsClient, PaymentConfig, eth_account_to_signer

account = Account.from_key("0xYOUR_KEY")
client = LLM4AgentsClient(
    api_key="",
    payment=PaymentConfig(
        mode="x402",
        signer=eth_account_to_signer(account),
        network="base-sepolia",
    ),
)

# Probe + sign once, then hit the REST endpoint directly with the X-PAYMENT
signed = await client.x402.sign()
async with httpx.AsyncClient() as http:
    res = await http.post(
        "https://api.llm4agents.com/v1/scrape/markdown",
        headers={"x-payment": signed.encoded_header},
        json={"url": "https://example.com"},
    )
    print(res.json())

The MCP tools accessor (client.tools.scraper.markdown(...)) currently uses Bearer auth via the MCP transport; the REST surface above is the path for walk-up.

Roadmap — server-side x402

Today the SDK is client-only: it signs X-PAYMENT headers so an agent can pay for outbound services. The mirror direction (an agent serving its own paywalled endpoint and receiving x402 payments from third parties) is not implemented and not on the v2.x roadmap.

If you want to monetize your agent with x402, use a server library directly:

Stack Library
FastAPI x402-fastapi
Flask x402-flask
Any framework (low-level, JWT-auth'd CDP facilitator) coinbase-x402
Other stacks see the x402 reference servers

If a first-class llm4agents-sdk-server (verify/settle/middleware helpers around coinbase-x402 with llm4agents-flavoured defaults) would help your use case, open an issue at llmforagents/sdk-python.

MCP Tools

All tool methods return McpToolResult. Access .text for the plain-text representation.

Scraper

html   = await client.tools.scraper.fetch_html("https://example.com")
md     = await client.tools.scraper.markdown("https://example.com")
links  = await client.tools.scraper.links("https://example.com")
shot   = await client.tools.scraper.screenshot("https://example.com", full_page=True)
pdf    = await client.tools.scraper.pdf("https://example.com")
data   = await client.tools.scraper.extract("https://example.com", schema={
    "type": "object",
    "properties": {"title": {"type": "string"}},
})

All scraper methods accept an optional proxy keyword: "none", "datacenter", or "residential".

Browser sessions

session = await client.tools.scraper.session_create(proxy="residential")
result  = await client.tools.scraper.session_run(
    session_id=session.text,
    actions=[{"type": "navigate", "url": "https://example.com"}],
)
status  = await client.tools.scraper.session_status(session_id=session.text)
await client.tools.scraper.session_close(session_id=session.text)

Search

results = await client.tools.search.google("TypeScript SDK design")
news    = await client.tools.search.google_news("Bitcoin", tbs="qdr:d")
places  = await client.tools.search.google_maps("coffee near me")
batch   = await client.tools.search.google_batch(["python", "golang"])

Image

img      = await client.tools.image.generate("A robot writing code")
edited   = await client.tools.image.edit("Make it blue", image_url="https://...")
analysis = await client.tools.image.analyze("What is this?", image_url="https://...")

image.generate and image.edit may return McpImageContent parts (base64 PNG) alongside text.

Tool Definitions

client.tools.definitions returns the cached list[ToolDefinition] (populated after first tool call). Use fetch_definitions() to eagerly fetch and cache the list:

defs = await client.tools.fetch_definitions()  # list[ToolDefinition] in OpenAI function format

Pass these to any LLM that supports function calling, or let conversation() manage them automatically when "tools": client.tools is set.

More tools via client.tools.call(...)

Beyond the typed scraper.*, search.*, image.*, and workspace.* namespaces, ~30 additional MCP tools are callable today through the generic client.tools.call(name, args) method. Pass the tool name and a dict of snake_case arguments; you get back an McpToolResult (.text for plain text, .raw for the structured payload). Typed namespaces for these may land in a later release — the generic call works now.

result = await client.tools.call("ai_summarize", {"text": long_article, "max_words": 80})
print(result.text)        # plain-text summary
print(result.raw)         # structured tool payload

Prices below are in US cents (¢). x402 is the per-call walk-up price; tools marked balance-only are not payable via x402 (Bearer / pre-funded balance only).

AI (Workers AI)

Tool Params Price (balance / x402)
ai_summarize text, max_words? 0.5¢ / 0.45¢
ai_translate text, target_lang, source_lang? 0.5¢ / 0.45¢
ai_embed input (str or str[] ≤100) → 768-dim 0.1¢ / 0.09¢
ai_classify text, labels (2-20) 1¢ / 0.9¢
ai_moderate text 1¢ / 0.9¢
ai_rerank query, documents (≤100), top_k? 1¢ / 0.9¢
image_to_text image_base64 2¢ / 1.8¢
speech_to_text audio_base64 1.5¢/MB / 1.35¢/MB (metered per MB)
res = await client.tools.call("ai_translate", {"text": "Hello", "target_lang": "es"})
labels = await client.tools.call("ai_classify", {"text": "I love this!", "labels": ["positive", "negative"]})

Notify

Tool Params Price (balance / x402)
send_telegram bot_token, chat_id, text, parse_mode? 1¢ / 0.9¢
send_discord webhook_url, content 1¢ / 0.9¢
send_slack webhook_url, text 1¢ / 0.9¢
webhook_post url, body, headers? 1¢ / 0.9¢
send_email to, subject, html? or text?, from? 2¢ / 1.8¢
send_sms to (E.164), body 3¢ / 2.7¢
await client.tools.call("send_slack", {"webhook_url": "https://hooks.slack.com/...", "text": "Scrape done"})
await client.tools.call("send_email", {"to": "ops@example.com", "subject": "Report", "text": "All green"})

Data

Tool Params Price (balance / x402)
dns_lookup name, type? free
ip_geolocate ip free
url_unfurl url 1¢ / 0.9¢
rss_parse url, limit? 1¢ / 0.9¢
youtube_transcript video, lang? 1¢ / 0.9¢
whois domain 1¢ / 0.9¢
crypto_price ids (≤50), vs_currencies? 1¢ / 0.9¢
fx_convert amount, from, to 1¢ / 0.9¢
qr_generate data, size?, ecc? 1¢ / 0.9¢
captcha_solve_create type, website_url, website_key, … 2¢ / 1.8¢
captcha_solve_result task_id free
prices = await client.tools.call("crypto_price", {"ids": ["bitcoin", "ethereum"], "vs_currencies": ["usd"]})
geo = await client.tools.call("ip_geolocate", {"ip": "8.8.8.8"})

Vector

Tool Params Price (balance / x402)
vector_upsert items (≤100, each {id, text? or vector?, metadata?}) 0.5¢ per 100 items / 0.45¢
vector_query query (str or float[768]), top_k?, filter? 1¢ / 0.9¢
vector_delete ids (≤100) free
await client.tools.call("vector_upsert", {"items": [{"id": "doc1", "text": "hello world", "metadata": {"src": "blog"}}]})
hits = await client.tools.call("vector_query", {"query": "greetings", "top_k": 5})

Web Crawl

Tool Params Price
web_crawl start_url, max_pages?, max_depth?, allow_subdomains?, render?, include?, exclude?, output? ('markdown' or 'links'), save_to_workspace? 0.5¢/page (min 2¢) — balance-only
crawl = await client.tools.call("web_crawl", {
    "start_url": "https://example.com",
    "max_pages": 20,
    "output": "markdown",
    "save_to_workspace": True,
})

Memory

Tool Params Price (balance / x402)
memory_set key, value (JSON ≤64KB), ttl_days? 1¢ / 0.9¢
memory_get key free
memory_list prefix?, limit?, cursor? free
memory_delete key free
await client.tools.call("memory_set", {"key": "run:42:status", "value": {"done": True}, "ttl_days": 7})
state = await client.tools.call("memory_get", {"key": "run:42:status"})

Web3 (chains: ethereum, polygon, base, solana)

Tool Params Price (balance / x402)
token_balance chain, address, token? 1¢ / 0.9¢
tx_status chain, tx_hash 1¢ / 0.9¢
nft_metadata chain, contract, token_id 1¢ / 0.9¢
ens_resolve query, direction ('forward' or 'reverse') 1¢ / 0.9¢
bal = await client.tools.call("token_balance", {"chain": "base", "address": "0xabc...", "token": "USDC"})
name = await client.tools.call("ens_resolve", {"query": "vitalik.eth", "direction": "forward"})

Document

Tool Params Price
pdf_parse workspace_file? or url?, pages? 0.5¢/page (min 1¢) — balance-only
doc_extract workspace_file, format ('docx', 'xlsx', or 'csv') 0.5¢/unit (min 1¢) — balance-only
article_extract url 1¢ / 0.9¢
parsed = await client.tools.call("pdf_parse", {"workspace_file": "docs/report.pdf", "pages": "1-5"})
article = await client.tools.call("article_extract", {"url": "https://example.com/post"})

Workspace tools

Every authenticated agent gets a private workspace backed by Cloudflare R2. Files are billed per-MB on upload (covering 1 day of storage) and per-day afterwards; downloads are billed per-MB. The same tools work with Bearer auth and with x402 walk-up.

Tools

Tool Purpose
workspace.create() Idempotent — confirm the workspace exists.
workspace.list(prefix=None, limit=None) List files. Free, rate-limited (60/min).
workspace.stat(filename) Get one file's metadata. Free, rate-limited.
workspace.delete(filename) Delete a file (no storage refund). Free, rate-limited.
workspace.upload(filename, content_base64, days_to_store, content_type=None) Inline upload, ≤10 MB. Billed per-MB + storage days.
workspace.upload_init(filename, size_bytes, days_to_store, content_type=None) Start a large upload. Returns {upload_id, put_url, expires_at, max_bytes}. Reserves cost.
workspace.upload_finalize(upload_id) Confirm the PUT and settle billing. Must be called within 15 min of init.
`workspace.download(filename, format='inline' 'url', url_ttl_minutes=None)`
workspace.extend(filename, additional_days) Extend storage on an existing file.
workspace.copy(source_filename, dest_filename, days_to_store) Server-side copy. Billed for destination storage only.

Pricing

Operation Price $/GB equivalent
Upload base 0.01¢/MB (min 1¢) $0.10/GB
Storage 0.0001¢/MB/day ~$0.03/GB-month
Download 0.004¢/MB (min 1¢) $0.04/GB
Storage extension 0.0001¢/MB/day ~$0.03/GB-month
List / stat / delete / create Free, 60 req/min

x402 walk-up rates are ~10% lower per-MB.

Note: Downloads never expose direct R2 URLs — both inline and url modes route bytes through our worker so per-download billing is enforced.

Quick example (Python SDK)

import base64
import json
import os
from llm4agents import LLM4AgentsClient

client = LLM4AgentsClient(api_key=os.environ["LLM4AGENTS_API_KEY"])

# Upload a small file
await client.tools.workspace.upload(
    filename="scrapes/page-1.md",
    content_base64=base64.b64encode(b"# Hello\n").decode("ascii"),
    days_to_store=7,
    content_type="text/markdown",
)

# List files — result.text is the JSON-stringified response from the MCP tool
list_result = await client.tools.workspace.list(prefix="scrapes/")
data = json.loads(list_result.text)
files = data["files"]

# Get a one-time proxied URL — useful when forwarding bytes to a third party
# (email attachment, frontend hand-off, etc.). The URL is single-use: the
# second hit returns 410. The agent is billed at issuance.
dl_result = await client.tools.workspace.download(
    filename="scrapes/page-1.md",
    format="url",
    url_ttl_minutes=5,
)
download_url = json.loads(dl_result.text)["download_url"]
# Hand `download_url` to whoever needs the bytes — they GET it once.

Models

result = await client.models.list()
for m in result.models:
    print(f"{m['slug']} — ${m['inputPricePer1M']}/1M in, ${m['outputPricePer1M']}/1M out")
    if m.get("feePct") is not None:
        print(f"  platform fee: {m['feePct']}%")

# Filter by name
result = await client.models.list(search="claude")

models.list() returns a ModelListResult with .models (list of dicts), .fee_pct (int | None — the agent's platform fee percentage, applied to every billed call), and .request_id (str | None). Each model dict contains slug, displayName, provider, inputPricePer1M, outputPricePer1M, contextWindow, lastSyncedAt, and an optional feePct (platform fee percentage).

Embeddings

res = await client.embeddings.create(
    model="openai/text-embedding-3-large",
    input="How many vectors fit in a haystack?",
)
print(len(res.data[0].embedding))  # → e.g. 3072
print(res.usage.prompt_tokens, res.model)

# Batch input
batch = await client.embeddings.create(
    model="openai/text-embedding-3-small",
    input=["first", "second", "third"],
)
for item in batch.data:
    print(item.index, item.embedding)

embeddings.create() accepts model (slug), input (string or list of strings, max 2048 entries), and the optional encoding_format, dimensions, and user keyword args. Returns an EmbeddingsResponse with .data: list[EmbeddingItem], .model, and .usage (prompt_tokens, total_tokens). Embeddings have no completion tokens, so billing is input-only.

Catalog: Embedding models do not appear in OpenRouter's public catalog endpoint, so the proxy maintains them by hand. New embedding models can be added through the admin panel — see model_type='embedding' rows.

Audio (TTS)

result = await client.audio.speech.create(
    model="x-ai/grok-voice-tts-1.0",
    input="Estás en el interior de una pirámide.",
    voice="sal",             # eve | ara | rex | sal | leo
    response_format="mp3",
)
open("speech.mp3", "wb").write(result.data)

audio.speech.create() posts to POST /v1/audio/speech and returns a SpeechResult: .data (bytes — the raw audio bytes), .content_type (from the response's content-type header, defaults to audio/mpeg), and optional .request_id, .charged_usd_cents, .model_used surfaced from the x-request-id, x-charged-usd-cents, and x-model-used response headers. input is capped at 15,000 characters server-side; a 402 raises the usual insufficient_balance LLM4AgentsError. Unlike other endpoints, /v1/audio/speech is Bearer-only — it is not on the x402 walk-up allowlist, so an x402-mode client raises immediately with a clear error rather than attempting the request.

MCP tool equivalent, useful inside client.tools.call() dispatch or a tool-calling conversation loop:

result = await client.tools.text_to_speech("Hola, ¿cómo estás?", voice="eve", format="mp3")

tools.text_to_speech() delegates to the text_to_speech MCP tool (metered per 1k input characters). Audio payloads over 256KB are not inlined in the tool result — they land in your workspace storage instead; check result.text / result.raw for the workspace reference.

Video (async)

Video generation is asynchronous: create() returns immediately with a job id and a charge estimate, and you poll get() until the job reaches a terminal status before fetching bytes with content().

import asyncio

job = await client.videos.create(
    prompt="A cat riding a skateboard through a neon city",
    image="https://example.com/first-frame.png",  # x-ai/grok-imagine-video-1.5 is image-to-video
    model="x-ai/grok-imagine-video-1.5",
    duration=5,
    resolution="720p",
    aspect_ratio="16:9",
    generate_audio=True,
)
print(job.id, job.status, job.charged_usd_cents)  # -> 'job_abc', 'pending', 250

# Poll until the job is done
status = await client.videos.get(job.id)
while status.status in ("pending", "in_progress"):
    await asyncio.sleep(5)
    status = await client.videos.get(job.id)

if status.status == "completed":
    result = await client.videos.content(job.id)
    open("output.mp4", "wb").write(result.data)
elif status.status == "failed":
    print(status.error, "refunded:", status.refunded)

videos.create() posts to POST /v1/videos and returns a VideoJobAccepted (202): id, status, polling_url, charged_usd_cents — the estimated charge reserved upfront. videos.get(job_id) polls GET /v1/videos/{job_id} and returns a VideoJobStatus whose status is one of pending | in_progress | completed | failed | cancelled | expired; once completed it carries video_url, and on failed/cancelled/expired it carries error and a refunded flag (the reserved estimate is automatically refunded when a job doesn't complete). videos.content(job_id) fetches GET /v1/videos/{job_id}/content and returns a VideoContentResult: .data (bytes — raw mp4 bytes), .content_type (defaults to video/mp4), and an optional .request_id from the x-request-id response header. Like audio, the entire /v1/videos surface is Bearer-only — it is not on the x402 walk-up allowlist.

MCP tool equivalents, useful inside client.tools.call() dispatch or a tool-calling conversation loop:

started = await client.tools.generate_video(
    "A cat riding a skateboard through a neon city",
    image_url="https://example.com/first-frame.png",  # https only
    model="x-ai/grok-imagine-video-1.5",
    duration=5,
)
polled = await client.tools.video_status("job_abc")

tools.generate_video() delegates to the generate_video MCP tool and tools.video_status(job_id) delegates to video_status (with the job id passed as job_id). Prefer client.videos.* for the typed REST flow above; these wrappers exist for tool-calling loops that dispatch by tool name.

Images

Synchronous image generation — unlike video, generate() returns the finished image(s) inline in the response body, no polling required.

res = await client.images.generate(
    prompt="A robot writing code, studio lighting",
    model="x-ai/grok-imagine-image-quality",
    n=1,
    resolution="1K",
    output_format="png",
)

print(res.created, res.cost_usd)  # -> 1700000000, 0.04 (real cost, USD)

image = res.data[0]
if image:
    import base64
    open("output.png", "wb").write(base64.b64decode(image.b64_json))

images.generate() posts to POST /v1/images/generations and returns an ImagesGenerateResponse: created (unix timestamp, optional), data (list of GeneratedImage.b64_json base64-encoded image bytes, decode with base64.b64decode(...), and .media_type), and cost_usd — the real cost in USD for the images actually generated (not an upfront estimate like videos.create()'s charged_usd_cents). prompt and model are named parameters; any other field accepted by the endpoint (n, resolution, aspect_ratio, quality, output_format, background, output_compression, seed, input_references, ...) can be passed as a keyword argument and is forwarded as-is (omitted when None). The x-charged-usd-cents response header (alongside x-request-id and x-model-used) carries the actual amount debited from your balancecost_usd plus the platform fee, rounded up to the nearest cent — which is not the same number as cost_usd itself. A 402 raises the usual insufficient_balance LLM4AgentsError. Like audio and video, /v1/images/generations is Bearer-only — it is not on the x402 walk-up allowlist.

client.images.generate vs client.tools.image.*: these are two different surfaces and are not interchangeable. client.images.generate() talks to the main proxy's /v1/images/generations REST endpoint directly — synchronous, typed, and billed against your agent balance like chat/embeddings/audio/video. client.tools.image.* (see MCP Tools → Image) instead dispatches to the scraper-worker's PiAPI-backed generate_image / edit_image / analyze_image MCP tools, which support image editing and vision analysis that the /v1/images/generations endpoint does not, and are priced per the MCP tool pricing table rather than the images-service model catalog. Use client.images.generate() for straightforward text-to-image generation against a specific model slug; use client.tools.image.* when you need edit/analyze or are already driving a tool-calling conversation loop.

Error Handling

All errors are instances of LLM4AgentsError:

from llm4agents import LLM4AgentsClient
from llm4agents.errors import LLM4AgentsError

try:
    await client.chat.completions.create(...)
except LLM4AgentsError as err:
    print(err.code, err.status_code, err.request_id, err.message)
code HTTP status Description
auth_error 401, 403 Invalid or missing API key
insufficient_balance 402 Not enough balance to cover the request
rate_limited 429 Too many requests
model_not_found 404 Requested model does not exist in the catalog
model_disabled 422 Model exists but is currently disabled
context_overflow Prompt + max_tokens exceeds the model's context window
gas_spike 409 Network gas price spiked above safe threshold during transfer
signature_mismatch 422 EIP-712 permit signature could not be verified
invalid_token 422 Unsupported token or chain for gasless transfer
operator_unavailable 503 Gasless relayer is temporarily unavailable
deadline_expired 400 EIP-712 permit deadline passed before submission
tool_not_found MCP tool name not found in the server's tool list
tool_execution_error MCP tool returned an error result
tool_loop_limit Conversation exceeded max_tool_rounds without a final answer
network_error Connection failed (DNS failure, TCP reset, etc.)
timeout Request exceeded the configured timeout
api_error 4xx, 5xx Any other non-success response

Constructor Options

client = LLM4AgentsClient(
    api_key="sk-proxy-...",                         # required in Bearer mode; "" in x402 mode
    base_url="https://api.llm4agents.com",          # optional
    mcp_url="https://mcp.llm4agents.com/mcp",       # optional
    timeout=30.0,                                   # optional, seconds, default 30
    payment=PaymentConfig(mode="bearer"),           # optional, default; or PaymentConfig(mode="x402", signer=..., network=...)
)

What's New in v2.8

  • Text-to-speechclient.audio.speech.create(model=..., input=..., voice=..., response_format=None, speed=None) posts to the new POST /v1/audio/speech endpoint and returns a SpeechResult (.data: bytes, .content_type, .request_id, .charged_usd_cents, .model_used). This endpoint is Bearer-only (not on the x402 walk-up allowlist). An MCP-tool equivalent is also available as client.tools.text_to_speech(text, voice=None, model=None, format=None) for tool-calling conversation loops. See Audio (TTS).
  • New types exported: Audio, Speech, SpeechResult.
  • Async video generationclient.videos.create(prompt=..., model=None, image=None, duration=None, resolution=None, aspect_ratio=None, generate_audio=None, seed=None) posts to POST /v1/videos and returns a VideoJobAccepted (202) immediately; poll with client.videos.get(job_id) until status is terminal, then fetch bytes with client.videos.content(job_id) (VideoContentResult.data is bytes, .content_type defaults to video/mp4). Failed/cancelled/expired jobs are automatically refunded (VideoJobStatus.refunded). This endpoint is Bearer-only (not on the x402 walk-up allowlist). MCP-tool equivalents are also available as client.tools.generate_video(prompt, image_url=None, model=None, duration=None, resolution=None, aspect_ratio=None, generate_audio=None) and client.tools.video_status(job_id). See Video (async).
  • New types exported: Videos, VideoJobAccepted, VideoJobStatus, VideoContentResult.
  • Synchronous image generationclient.images.generate(prompt=..., model=None, n=None, **kwargs) posts to POST /v1/images/generations and returns the finished image(s) inline (ImagesGenerateResponse.data[].b64_json, base64) — no polling, unlike video. Extra fields (resolution, aspect_ratio, quality, output_format, background, output_compression, seed, input_references, ...) pass through as keyword arguments and are omitted when None. cost_usd reports the real USD cost of the images actually generated; the x-charged-usd-cents response header instead carries the amount actually charged to your balance (cost_usd plus the platform fee, rounded up to the cent). This endpoint is Bearer-only (not on the x402 walk-up allowlist) and is distinct from the PiAPI-backed client.tools.image.* MCP tools (edit/analyze support, separate pricing). See Images.
  • New types exported: Images, GeneratedImage, ImagesGenerateResponse.
  • The REST API's GET /api/v1/transactions endpoint now accepts an optional ?service= query filter (alongside the existing ?type=) to narrow results to a specific billed service: llm (chat completions + embeddings, default), tts (audio speech), video (video generation), tools (MCP registry tools), scraper, search, image, workspace. Typed SDK support for this filter will land in a follow-up release — for now, pass it via a raw HTTP call against the REST endpoint if needed.

What's New

  • ~30 new MCP tools available via client.tools.call(...) — No SDK upgrade required: these tools are live on the server and callable today through the generic client.tools.call(name, args) method (returns McpToolResult.text plain text, .raw structured). Eight categories: ai (Workers AI summarize / translate / embed / classify / moderate / rerank / image-to-text / speech-to-text), notify (Telegram / Discord / Slack / webhook / email / SMS), data (DNS / IP geolocation / unfurl / RSS / YouTube transcript / WHOIS / crypto price / FX / QR / captcha), vector (upsert / query / delete), web_crawl, memory (set / get / list / delete), web3 (token balance / tx status / NFT metadata / ENS), and document (PDF parse / doc extract / article extract). Most accept x402 walk-up; web_crawl, pdf_parse, and doc_extract are balance-only. Typed namespaces (client.tools.ai.*, etc.) may follow in a later release; the generic call() works now. See More tools via client.tools.call(...).

What's New in v2.6

2.6.1 — streaming meta cost (bugfix)

  • Fix: ResponseMeta.cost_usd_cents was None on streaming rounds. The proxy doesn't emit x-cost-usd-cents in response headers for streams — it attaches the final cost to the terminating SSE chunk's usage.cost field (USD). The SDK now promotes that into cost_usd_cents (cents) so streaming and non-streaming consumers see the same field. Fractional cents are preserved (no truncation), so micro-spend rounds report their real cost.
  • Type widening: cost_usd_cents is now float | None (was int | None). Streaming completions can produce fractional cents (e.g. 0.0035 for a haiku round). Integer values from x-cost-usd-cents headers still satisfy the new type without code changes.

2.6.0

  • Conversation accepts tool_choice — Force tool selection on round 1 only; reverts to "auto" for subsequent rounds so the model can summarize tool results naturally without looping. Critical for agent-routing patterns where the model would otherwise emit JSON-as-text instead of using the tool_calls API. See Chat.

What's New in v2.5

  • Workspace tools (NEW) — Private R2-backed file storage per agent. 10 new MCP tools for upload (inline and pre-signed), download (inline or signed URL), list, stat, extend, copy, and delete. Works with Bearer and x402. See Workspace tools.
  • x402 walk-up payment mode — pay per-request from a wallet on /v1/chat/completions without registering an agent. Pass payment=PaymentConfig(mode="x402", signer=..., network=...) to the client constructor. Supports both eth_account.LocalAccount (via eth_account_to_signer) and any custom Signer Protocol implementation (HSM, KMS, hardware wallets, WalletConnect) thanks to the Ports & Adapters design — sign_typed_data may be sync or async.
  • Streaming responses emit a typed x402_receipt event after [DONE], surfacing the on-chain settlement receipt (transaction, network, amount, payer) to conv.stream() consumers and the on_x402_receipt callback on chat.completions.create().
  • New client.x402 namespace — probe(), sign(recipient=...), and sign_from_requirements(req, recipient=...) helpers for low-level integrations.
  • New top-level exports: PaymentConfig, PaymentPayload, PaymentRequirements, Signer, SignedPayment, X402Namespace, X402Network, X402PaymentRequiredError, X402Receipt, eth_account_to_signer, build_transfer_with_authorization_typed_data, generate_nonce, network_to_caip2, sign_from_requirements, encode_payment_header, decode_payment_required_header, pick_supported_requirements, USDC_ADDRESS_BY_NETWORK, USDC_DOMAIN_NAME_BY_NETWORK, X402_CAIP2_BY_NETWORK, CHAIN_ID_BY_NETWORK, TRANSFER_WITH_AUTHORIZATION_TYPES, DEFAULT_VALID_FOR_SECONDS.
  • x402 allowlist extended to the MCP REST surface — clients in x402 mode can now hit /v1/scrape/*, /v1/search/*, and /v1/image/* in addition to chat. Prices are admin-editable in cents from the operator panel (parallel value for balance / x402_value for walk-up per tool).

What's New in v2.4

  • client.embeddings.create(model=..., input=...) — OpenAI-compatible embeddings against POST /v1/embeddings. Pass a string or a list of up to 2048 strings; receive an EmbeddingsResponse with data: list[EmbeddingItem], model, and usage (prompt_tokens, total_tokens). Embeddings are billed input-only — there are no completion tokens.
  • New types exported from llm4agents: Embeddings, EmbeddingItem, EmbeddingsResponse, EmbeddingsUsage. The embedding-model catalog is curated by hand on the server because OpenRouter omits embedding models from its public catalog endpoint; admins maintain the list via the proxy admin panel.

What's New in v2.3

This release brings the Python SDK to feature parity with the TypeScript SDK at v2.3.1. Seven fixes that an external playground audit flagged as real-world breakages:

  • Prompt-mode tool fallback (enable_prompt_tool_fallback=True) — when a model ignores native tools on round 0, the SDK retries the round with the tool definitions injected into the system prompt and parses <tool_call>{"name":"...","arguments":{...}}</tool_call> blocks from the response. Recovers tool use on models without native function calling. stream() emits a {"type": "fallback", "reason": "tools_ignored", "model": ...} event before the fallback round.
  • MCP Accept header — MCP rpc requests now send Accept: application/json, text/event-stream. Streamable HTTP MCP servers were rejecting every tool call with HTTP 406 without it.
  • Models endpoint trailing slashclient.models.list() now hits /api/v1/models (no trailing slash), matching the deployed API.
  • reasoning_tokens propagationConversationResponse.usage["reasoning_tokens"] now accumulates across all rounds, and ResponseMeta.tokens_reasoning exposes the per-round value parsed from the x-tokens-reasoning header.
  • fee_pct in ModelListResult — the platform fee percentage returned by the API is now surfaced as result.fee_pct: int | None.
  • on_final_usage callback in streaming completionschat.completions.create(params, on_final_usage=cb) fires cb(FinalUsage(prompt_tokens, completion_tokens, total_tokens, reasoning_tokens)) once after the SSE stream ends, using the last usage chunk emitted by providers that send stream_options: {include_usage: true}.
  • Assistant content: null normalization — when a model returns content: null for a tool-only message (OpenAI / Gemini convention), the SDK now normalizes it to "" before pushing to history. Strict backends were rejecting follow-up requests because messages[*].content was null.

What's New in v2.1

  • Conversation.stream() — async generator yielding typed events: text, reasoning, tool_call, tool_result, meta, done. Each round emits a meta event with ResponseMeta parsed from response headers.
  • on_round_meta — callback fired per round with ResponseMeta (cost, balance, tokens, model used, request id).
  • on_tools_ignored(model) — callback fired when a model returns no tool calls on the first round despite being given tools — useful for detecting models without native function-calling support.
  • agents.register(name) — register a new agent and receive its api_key (only shown once).
  • forwarder_address in QuoteResult — the EIP-2771 forwarder contract used for the gasless transfer.
  • Auto-normalization in MCP transport — JSON-wrapped image/PDF payloads embedded inside text blocks ({"imageBase64": "..."}, {"pngBase64": "..."}, {"pdfBase64": "..."}) are auto-promoted to typed McpImageContent / McpResourceContent. Snake_case mime_type is aliased to mimeType and MIME types are sniffed from base64 magic bytes when missing.
  • Top-level exportsfrom llm4agents import AgentRegistration, ResponseMeta, Conversation, McpToolResult, ... (all public types now re-exported from the package root).
  • Transaction.type filter accepts "gas_sponsored" in addition to "deposit", "usage", "refund".
  • ModelInfo.feePct — platform fee percentage now exposed in the model catalog.

Migration from v1.x

Before (v1) After (v2)
result = await tools.call(name, args)str result.text or str(result)
conv.on_tool_result callback receives str callback now receives McpToolResult
models = await client.models.list()list result.models (access via .models)
StreamEvent["tool_end"]["result"]str .result is now McpToolResult

License

MIT

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

llm4agents_sdk-2.8.0.tar.gz (107.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

llm4agents_sdk-2.8.0-py3-none-any.whl (66.7 kB view details)

Uploaded Python 3

File details

Details for the file llm4agents_sdk-2.8.0.tar.gz.

File metadata

  • Download URL: llm4agents_sdk-2.8.0.tar.gz
  • Upload date:
  • Size: 107.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.10.12

File hashes

Hashes for llm4agents_sdk-2.8.0.tar.gz
Algorithm Hash digest
SHA256 05f68f3b5453dab719d6aa17a90c5fd2da69b0731033e037429c1b82e1080c83
MD5 25204f35717d73cf87c26a1fe985dad8
BLAKE2b-256 b93cfe3b27b04768912536c01b355b09b27396dfc6f5910884a57c9170859924

See more details on using hashes here.

File details

Details for the file llm4agents_sdk-2.8.0-py3-none-any.whl.

File metadata

  • Download URL: llm4agents_sdk-2.8.0-py3-none-any.whl
  • Upload date:
  • Size: 66.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.10.12

File hashes

Hashes for llm4agents_sdk-2.8.0-py3-none-any.whl
Algorithm Hash digest
SHA256 f19cfa1e52e9b66d51f031faa8a8977866a9c80fbf950079fa6cf60b177bc5f7
MD5 9694d128925726c01a8fc599c7b9a454
BLAKE2b-256 9dc0fbcc736060070139298a6eb8e62b681079aded426f66ff17027ab2c28b72

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page