autourgos-responses
A single, self-contained LLM wrapper for the OpenAI Responses API (client.responses.create), and by extension every provider that speaks the same protocol (Groq, Gemini, Azure, Ollama, and more). Part of the Autourgos agentic-AI framework. Depends on autourgos-openaichat for the shared base layer (BaseLLM, circuit breaker) in addition to openai.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(model="gpt-4o") # reads OPENAI_API_KEY
reply = llm.invoke("What is the capital of France?")
print(reply)
# Paris
Features
- One interface, any OpenAI-compatible provider: OpenAI, Azure, Groq, Gemini, Mistral, DeepSeek, Ollama, and more, switched with just
base_url+model - Native reasoning models (
o3,o3-mini,o1) with configurablereasoning_effortandreasoning_summary - Text verbosity control (
text_verbosity) - Sync and async generation, plus streaming for both
- Structured output validated against a Pydantic model, or plain JSON mode
- Multi-modal vision input: file paths, URLs, or raw bytes
- Prompt templates with
{placeholder}variables - Multi-turn conversations via
chat()/achat(), or a plain message list as input - Automatic retries with exponential back-off (skips non-retryable 4xx errors), plus a circuit breaker for cascading-failure protection and an automatic provider fallback chain — no proxy/gateway needed — plus an optional aggregate call deadline to cap total wall-clock time across all of them
- Native tool / function calling, plus validated structured output with an automatic validation-retry loop
- Built-in cost/latency tracking, plus a budget governor that hard-stops calls once a USD cap is reached
- Optional local call ledger (SQLite, no external service) and shadow-mode dual dispatch for comparing providers concurrently
- Optional PII/secret redaction: a heuristic pre-flight scrubber that masks (or blocks) emails, API keys, credit cards, SSNs, and phone numbers — with a bring-your-own-dictionary option and reversible restore-in-response
extra_body=passthrough for provider-specific request fields — e.g. vLLM'sguided_json/guided_regexfor constrained decoding- Fully typed (
py.typed), sync/async context managers, low-level raw-response access
Table of Contents
- Install
- Supported Providers
- Provider Examples
- Core Usage
- Text Generation
- Async Generation
- Streaming
- Async Streaming
- Batch Invocation
- System Prompt
- Prompt Templates
- Reasoning Models
- Vision Input
- Structured Output
- Validated Structured Output
- JSON Mode
- Multi-Turn Chat
- Native Tool Calling
- Reliability
- Circuit Breaker
- Provider Fallback Chain
- Aggregate Call Deadline
- Cost
- Cost Tracking
- Budget Governor
- Observability
- Call Ledger (Audit Trail)
- Shadow-Mode Dual Dispatch
- Security
- PII / Secret Redaction
- Advanced
- Constrained Decoding / Provider-Specific Params
- Context Manager
- Low-Level Access
- Error Handling
- Constructor Reference
- API Reference
- Differences vs autourgos-openaichat
- License
Install
pip install autourgos-responses
Requires Python 3.10+ and openai>=1.0.0. Structured output (output_schema=) additionally needs pydantic>=2.0 if you use it.
Supported Providers
Almost every major LLM provider exposes an OpenAI-compatible API: same request format as OpenAI's Responses endpoint. Point base_url at the provider and model at whatever they offer; nothing else changes.
| Provider | base_url |
Get a key |
|---|---|---|
| OpenAI | (default, omit) | https://platform.openai.com/api-keys |
| Azure OpenAI | https://<resource>.openai.azure.com/openai/deployments/<deployment> |
Azure Portal |
| Google Gemini | https://generativelanguage.googleapis.com/v1beta/openai/ |
https://aistudio.google.com/apikey |
| Groq | https://api.groq.com/openai/v1 |
https://console.groq.com |
| xAI (Grok) | https://api.x.ai/v1 |
https://console.x.ai |
| OpenRouter | https://openrouter.ai/api/v1 |
https://openrouter.ai/keys |
| Together AI | https://api.together.xyz/v1 |
https://api.together.xyz |
| Mistral AI | https://api.mistral.ai/v1 |
https://console.mistral.ai |
| DeepSeek | https://api.deepseek.com/v1 |
https://platform.deepseek.com |
| Perplexity | https://api.perplexity.ai |
https://www.perplexity.ai/settings/api |
| Ollama (local) | http://localhost:11434/v1 |
none, runs on your machine |
| LM Studio (local) | http://localhost:1234/v1 |
none, runs on your machine |
| vLLM (self-hosted) | http://your-server:8000/v1 |
none, you host it |
Note: reasoning models (
o3,o3-mini,o1) andreasoning_effort/text_verbosityare OpenAI-only features of the Responses API. Other providers accept the sameinvoke/stream/chatcalls but ignore or reject those params.
Provider Examples
Every example below is the full, runnable snippet. Swap in your own key and go.
OpenAI
The default provider. No base_url needed.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(
model="gpt-4o",
api_key="sk-...", # or set OPENAI_API_KEY env var
)
reply = llm.invoke("What is the capital of France?")
print(reply)
# Paris
OpenAI Reasoning Models
o3, o3-mini, and o1 support reasoning_effort to control how long the model thinks before answering.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(
model="o3-mini",
api_key="sk-...",
reasoning_effort="high", # "low", "medium", or "high"
)
reply = llm.invoke("Prove that the square root of 2 is irrational.")
print(reply)
# Assume for contradiction that √2 = p/q in lowest terms...
Azure OpenAI
Azure hosts OpenAI models in your own subscription. model is your deployment name in Azure, not the base model name. Get your endpoint and key from the Azure Portal.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(
model="gpt-4o", # your deployment name in Azure
api_key="...", # Azure OpenAI key
base_url="https://<your-resource>.openai.azure.com/openai/deployments/gpt-4o",
)
reply = llm.invoke("What is cloud computing?")
print(reply)
# Cloud computing is the delivery of computing services over the internet
# (servers, storage, databases, networking, software) on a pay-as-you-go basis.
Google Gemini
Gemini exposes an OpenAI-compatible endpoint, so no separate Google SDK is needed. Get your key at https://aistudio.google.com/apikey.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(
model="gemini-2.0-flash",
api_key="...", # Gemini API key
base_url="https://generativelanguage.googleapis.com/v1beta/openai/",
)
reply = llm.invoke("Explain photosynthesis in one sentence.")
print(reply)
# Photosynthesis is the process by which plants convert sunlight, water, and
# carbon dioxide into glucose and oxygen.
Other Gemini models: gemini-2.0-flash-lite, gemini-1.5-pro, gemini-1.5-flash.
Groq (fastest inference, free tier available)
Groq runs open-source models (Llama 3, Mixtral, Gemma) at extremely high speed. Get your key at https://console.groq.com.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(
model="llama3-70b-8192",
api_key="gsk_...", # Groq API key
base_url="https://api.groq.com/openai/v1",
)
reply = llm.invoke("Explain quantum entanglement simply.")
print(reply)
# Quantum entanglement is when two particles become linked so that
# the state of one instantly affects the other, no matter how far apart they are.
Other Groq models: llama3-8b-8192, mixtral-8x7b-32768, gemma2-9b-it.
xAI (Grok)
Get your key at https://console.x.ai.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(
model="grok-2-latest",
api_key="xai-...", # xAI API key
base_url="https://api.x.ai/v1",
)
reply = llm.invoke("What makes Mars red?")
print(reply)
# Mars appears red because its surface is covered in iron oxide (rust),
# formed when iron in the soil reacted with trace oxygen long ago.
OpenRouter (one key, hundreds of models)
OpenRouter proxies dozens of providers (including Anthropic Claude and Google Gemini) behind a single OpenAI-compatible API and one API key. Get your key at https://openrouter.ai/keys.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(
model="anthropic/claude-3.5-sonnet", # or "google/gemini-2.0-flash-001", "openai/gpt-4o", ...
api_key="sk-or-...", # OpenRouter API key
base_url="https://openrouter.ai/api/v1",
)
reply = llm.invoke("Write a Python one-liner to reverse a string.")
print(reply)
# s[::-1]
Together AI (wide model selection)
Together AI hosts hundreds of open-source models. Get your key at https://api.together.xyz.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(
model="meta-llama/Llama-3-70b-chat-hf",
api_key="...", # Together AI key
base_url="https://api.together.xyz/v1",
)
reply = llm.invoke("Write a Python function to check if a number is prime.")
print(reply)
# def is_prime(n: int) -> bool:
# if n < 2:
# return False
# for i in range(2, int(n**0.5) + 1):
# if n % i == 0:
# return False
# return True
Other Together AI models: mistralai/Mixtral-8x7B-Instruct-v0.1, Qwen/Qwen2-72B-Instruct.
Mistral AI
Get your key at https://console.mistral.ai.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(
model="mistral-large-latest",
api_key="...", # Mistral API key
base_url="https://api.mistral.ai/v1",
)
reply = llm.invoke("What are the benefits of test-driven development?")
print(reply)
# TDD helps you write cleaner code, catch bugs early, and gives
# you confidence to refactor without breaking existing behaviour.
Other Mistral models: mistral-medium-latest, mistral-small-latest, open-mixtral-8x7b.
DeepSeek
Get your key at https://platform.deepseek.com.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(
model="deepseek-chat",
api_key="...", # DeepSeek API key
base_url="https://api.deepseek.com/v1",
)
reply = llm.invoke("What is a transformer neural network?")
print(reply)
# A transformer is a neural network architecture that uses self-attention
# to process input sequences in parallel, making it highly effective for
# NLP tasks like translation, summarisation, and text generation.
Other DeepSeek models: deepseek-reasoner.
Perplexity (web-connected models)
Perplexity's Sonar models can search the web in real time. Get your key at https://www.perplexity.ai/settings/api.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(
model="llama-3.1-sonar-large-128k-online",
api_key="pplx-...", # Perplexity API key
base_url="https://api.perplexity.ai",
)
reply = llm.invoke("What is the latest version of Python?")
print(reply)
# Python 3.13.x is the latest stable release as of 2025...
Ollama (run any model locally, no internet needed)
Ollama runs models entirely on your machine. Install from https://ollama.com, then pull a model:
ollama pull llama3
No API key needed for local use.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(
model="llama3",
api_key="ollama", # can be any string, Ollama ignores it
base_url="http://localhost:11434/v1",
)
reply = llm.invoke("What is machine learning?")
print(reply)
# Machine learning is a subset of AI where algorithms learn patterns
# from data to make predictions or decisions without explicit programming.
Other Ollama models: mistral, phi3, gemma2, codellama, qwen2, and anything you pull with ollama pull.
LM Studio (local models with a GUI)
LM Studio lets you download and run GGUF models locally. Start the local server in LM Studio, then:
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(
model="local-model", # use whatever model name LM Studio shows
api_key="lm-studio", # any string, ignored locally
base_url="http://localhost:1234/v1",
)
reply = llm.invoke("Tell me a short joke.")
print(reply)
# Why do programmers prefer dark mode? Because light attracts bugs!
vLLM (self-hosted high-throughput serving)
vLLM lets you host your own models with high throughput. After starting your vLLM server:
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(
model="meta-llama/Meta-Llama-3-8B-Instruct",
api_key="EMPTY", # vLLM's default when no auth is configured
base_url="http://your-server:8000/v1",
)
reply = llm.invoke("What is the capital of Japan?")
print(reply)
# Tokyo
Switching providers at runtime
Because all these providers use the same interface, switching is trivial:
from autourgos_responses import OpenAIResponse
PROVIDERS = {
"openai": {
"model": "gpt-4o-mini",
"api_key": "sk-...",
"base_url": None,
},
"groq": {
"model": "llama3-8b-8192",
"api_key": "gsk_...",
"base_url": "https://api.groq.com/openai/v1",
},
"gemini": {
"model": "gemini-2.0-flash",
"api_key": "...",
"base_url": "https://generativelanguage.googleapis.com/v1beta/openai/",
},
}
for name, cfg in PROVIDERS.items():
llm = OpenAIResponse(**cfg)
reply = llm.invoke("Say hello in one word.")
print(f"{name}: {reply}")
# openai: Hello!
# groq: Hello!
# gemini: Hello!
Core Usage
Text Generation
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(
model="gpt-4o",
api_key="sk-...", # or set OPENAI_API_KEY env var
temperature=0.7,
max_tokens=256,
)
reply = llm.invoke("Explain machine learning in one sentence.")
print(reply)
# Machine learning is a branch of AI where systems learn from data
# to make predictions or decisions without being explicitly programmed.
Async Generation
import asyncio
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(model="gpt-4o")
async def main():
reply = await llm.ainvoke("What is the speed of light?")
print(reply)
# The speed of light in a vacuum is approximately 299,792,458 metres per second.
asyncio.run(main())
Streaming
Stream the response token by token, synchronously.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(model="gpt-4o")
for chunk in llm.stream("Write a haiku about mountains."):
print(chunk, end="", flush=True)
# Silent peaks above,
# Clouds drift through the ancient stone,
# Eagles trace the wind.
You can also enable streaming at construction time so invoke() internally streams and returns the full joined text:
llm = OpenAIResponse(model="gpt-4o", streaming=True)
reply = llm.invoke("Tell me a fun fact about space.")
print(reply)
# A day on Venus is longer than a year on Venus — it takes 243 Earth days
# to rotate once but only 225 Earth days to orbit the Sun.
Async Streaming
import asyncio
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(model="gpt-4o")
async def main():
async for chunk in llm.astream("Count prime numbers up to 20."):
print(chunk, end="", flush=True)
# 2, 3, 5, 7, 11, 13, 17, 19
asyncio.run(main())
Batch Invocation
Synchronous (sequential):
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(model="gpt-4o-mini")
prompts = [
"Capital of Japan?",
"Capital of Germany?",
"Capital of Brazil?",
]
results = llm.batch_invoke(prompts)
for prompt, result in zip(prompts, results):
print(f"{prompt} -> {result}")
# Capital of Japan? -> Tokyo
# Capital of Germany? -> Berlin
# Capital of Brazil? -> Brasilia
Async (concurrent):
import asyncio
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(model="gpt-4o-mini")
async def main():
results = await llm.abatch_invoke([
"Capital of Japan?",
"Capital of Germany?",
"Capital of Brazil?",
])
print(results)
# ['Tokyo', 'Berlin', 'Brasilia']
asyncio.run(main())
System Prompt
Set a persistent system prompt sent as the instructions field of every request.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(
model="gpt-4o",
system_prompt="You are a pirate. Always respond in pirate speak.",
)
reply = llm.invoke("What time is it?")
print(reply)
# Arrr, I know not the exact hour, but the sun be high in the sky, matey!
Prompt Templates
Define a reusable template with {placeholders} and fill them at call time.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(
model="gpt-4o",
prompt_template="Summarise the following {topic} in {num_words} words:\n\n{content}",
)
reply = llm.invoke(prompt_variables={
"topic": "article",
"num_words": "30",
"content": "Quantum computing uses quantum bits (qubits) that can exist in superposition...",
})
print(reply)
# Quantum computing uses qubits in superposition to perform many calculations
# simultaneously, offering vastly superior speeds for specific complex problems
# like cryptography and molecular simulation.
Missing variables raise a clear error:
llm.invoke(prompt_variables={"topic": "article"})
# ValueError: Missing prompt template variables: content, num_words
Reasoning Models
o3, o3-mini, and o1 are OpenAI's reasoning models. They support reasoning_effort to control how long the model thinks before answering. Higher effort produces better answers for hard problems but takes longer and costs more.
Reasoning models and
reasoning_effort/reasoning_summary/text_verbosityare OpenAI-only. When using other providers, omit these params.
from autourgos_responses import OpenAIResponse
# Low effort — fast, cheaper
llm = OpenAIResponse(model="o3-mini", reasoning_effort="low")
reply = llm.invoke("What is 17 x 23?")
print(reply)
# 391
# Medium effort — balanced
llm = OpenAIResponse(model="o3-mini", reasoning_effort="medium")
reply = llm.invoke("Solve: if a train travels at 80 km/h for 2.5 hours, how far does it go?")
print(reply)
# The train travels 200 km. (80 km/h x 2.5 h = 200 km)
# High effort — most thorough, best for hard problems
llm = OpenAIResponse(model="o3", reasoning_effort="high")
reply = llm.invoke("Prove that the square root of 2 is irrational.")
print(reply)
# Assume for contradiction that √2 = p/q where p and q are integers with no common factors...
| effort | Use for | Speed | Cost |
|---|---|---|---|
"low" |
Simple maths, factual Q&A, quick summaries | Very fast | Lowest |
"medium" |
Multi-step reasoning, code generation | Moderate | Medium |
"high" |
Hard proofs, complex analysis, frontier research | Slow | Highest |
Text verbosity is controlled separately with text_verbosity ("low", "medium", or "high"):
llm = OpenAIResponse(model="gpt-4o", text_verbosity="low")
reply = llm.invoke("Explain how a car engine works.")
Invalid values raise immediately:
OpenAIResponse(model="o3-mini", reasoning_effort="ultra")
# ValueError: Invalid reasoning_effort 'ultra'. Must be one of: ['high', 'low', 'medium']
Vision Input
Pass image files, URLs, or raw bytes alongside text.
Note: vision support depends on the provider and model. GPT-4o, Gemini, LLaVA (on Ollama), and several others support it.
Warning: the file-path branch reads whatever local path it's given and base64-embeds its contents into the outgoing API request, with no path validation. Do not pass LLM- or tool-controlled paths through unchecked. An unchecked path could be used to exfiltrate arbitrary local files.
From a file path:
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(model="gpt-4o")
reply = llm.invoke("What objects are in this image?", files=["photo.jpg"])
print(reply)
# The image shows a wooden desk with a laptop, a coffee mug, and an open notebook.
From a URL:
reply = llm.invoke(
"Describe this chart in detail.",
files=["https://example.com/sales-chart.png"],
)
print(reply)
# The chart is a bar graph comparing quarterly revenue across four product lines.
# Q3 shows the highest sales at approximately $2.4M for Product A...
From raw bytes:
with open("diagram.png", "rb") as f:
image_bytes = f.read()
reply = llm.invoke("Explain this architecture diagram.", files=[image_bytes])
print(reply)
# The diagram shows a microservices architecture with an API gateway at the top
# routing requests to three downstream services: Auth, Orders, and Payments...
Multiple images:
reply = llm.invoke(
"Which image shows more people?",
files=["crowd1.jpg", "crowd2.jpg"],
)
print(reply)
# The first image shows more people — it appears to be a large outdoor concert
# with thousands of attendees, while the second shows a small group of around 20.
Structured Output
Return a Pydantic model as JSON automatically.
from pydantic import BaseModel, Field
from autourgos_responses import OpenAIResponse
import json
class WeatherReport(BaseModel):
city: str = Field(description="Name of the city")
temperature_celsius: float = Field(description="Current temperature in Celsius")
condition: str = Field(description="Weather condition e.g. Sunny, Rainy")
humidity_percent: int = Field(description="Humidity percentage 0-100")
llm = OpenAIResponse(model="gpt-4o", output_schema=WeatherReport)
result = llm.invoke("Describe a typical summer day in London.")
data = json.loads(result["response"])
print(data)
# {
# "city": "London",
# "temperature_celsius": 22.0,
# "condition": "Partly Cloudy",
# "humidity_percent": 65
# }
Use a plain dict schema instead of Pydantic:
schema = {
"type": "object",
"properties": {
"name": {"type": "string"},
"age": {"type": "integer"},
},
"required": ["name", "age"],
}
llm = OpenAIResponse(model="gpt-4o", output_schema=schema)
result = llm.invoke("Invent a fictional person.")
print(result["response"])
# {"name": "Mira Caldwell", "age": 34}
Validated Structured Output
invoke_structured() builds on output_schema= and closes the loop: instead of a raw JSON string you get back a validated Pydantic instance directly. If the response fails validation (a missing field, a failed @field_validator, a provider that ignores strict JSON-schema mode, ...), the validation error is fed back to the model as a correction message and the request is retried, up to max_validation_retries times.
from pydantic import BaseModel, Field
from autourgos_responses import OpenAIResponse
class CityInfo(BaseModel):
city: str = Field(description="Name of the city")
country: str = Field(description="Name of the country")
population: int = Field(description="Approximate population")
llm = OpenAIResponse(model="gpt-4o", output_schema=CityInfo)
result = llm.invoke_structured("Tell me about Tokyo.")
print(result)
# CityInfo(city='Tokyo', country='Japan', population=13960000)
print(result.population)
# 13960000
print(llm.last_metadata["validation_retries"])
# 0 (no correction was needed)
If validation keeps failing, OpenAIResponseValidationError (a subclass of OpenAIResponseResponseError) is raised with .raw_text (the last invalid response) and .validation_error (the last Pydantic error):
from autourgos_responses import OpenAIResponseValidationError
try:
result = llm.invoke_structured("Tell me about Tokyo.", max_validation_retries=1)
except OpenAIResponseValidationError as e:
print(f"Still invalid after retries: {e.validation_error}")
print(f"Last raw response: {e.raw_text}")
Async version: await llm.ainvoke_structured(...).
Notes:
output_schemamust be a PydanticBaseModelclass (not a plain dict, notNone) — a dict schema has no.model_validate_json()to validate against.- Incompatible with
streaming=True, same asstructured_output=True. - Each validation retry re-runs the full transport-level retry budget (
max_retries) too, so worst-case cost/latency is roughlymax_validation_retries × max_retries— keepmax_validation_retriessmall (the default is2). - Composes with Provider Fallback Chain — each attempt goes through the same primary → fallback sequence.
JSON Mode
Force the model to return valid JSON without a schema.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(
model="gpt-4o",
response_mime_type="application/json",
system_prompt="Always respond with valid JSON only.",
)
reply = llm.invoke("List three programming languages with their year of creation.")
print(reply)
# {
# "languages": [
# {"name": "Python", "year": 1991},
# {"name": "JavaScript", "year": 1995},
# {"name": "Rust", "year": 2010}
# ]
# }
Multi-Turn Chat
Pass a list of role-tagged messages directly to carry conversation history.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(model="gpt-4o")
messages = [
{"role": "user", "content": "My favourite colour is blue."},
{"role": "assistant", "content": "That is a great choice! Blue is calming and versatile."},
{"role": "user", "content": "What is my favourite colour?"},
]
reply = llm.chat(messages)
print(reply)
# Your favourite colour is blue!
Async multi-turn:
import asyncio
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(model="gpt-4o")
async def main():
messages = [
{"role": "user", "content": "I work as a data scientist."},
{"role": "assistant", "content": "That is a fascinating field!"},
{"role": "user", "content": "What is my job?"},
]
reply = await llm.achat(messages)
print(reply)
# You work as a data scientist.
asyncio.run(main())
Building a conversation loop:
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(model="gpt-4o")
history = []
def chat(user_message: str) -> str:
history.append({"role": "user", "content": user_message})
reply = llm.chat(history)
history.append({"role": "assistant", "content": reply})
return reply
print(chat("My name is Jitin."))
# Nice to meet you, Jitin!
print(chat("I am building an AI framework called Autourgos."))
# That sounds exciting! What does Autourgos focus on?
print(chat("What is my name and what am I building?"))
# Your name is Jitin, and you are building an AI framework called Autourgos.
Native Tool Calling
Let the model decide when to call your functions.
Tool calling support varies by provider. OpenAI, Groq, Gemini, Together AI, Mistral, and DeepSeek all support it. Ollama supports it on compatible models.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(model="gpt-4o")
tools = [
{
"name": "get_weather",
"description": "Get the current weather for a city.",
"parameters": {
"type": "object",
"properties": {
"city": {
"type": "string",
"description": "The city name, e.g. Paris",
},
"unit": {
"type": "string",
"enum": ["celsius", "fahrenheit"],
"description": "Temperature unit",
},
},
"required": ["city"],
},
}
]
response = llm.invoke_with_tools("What is the weather in Tokyo right now?", tools)
if response.has_tool_calls:
for call in response.tool_calls:
print(f"Tool: {call.name}")
print(f"Args: {call.arguments}")
print(f"ID: {call.call_id}")
# Tool: get_weather
# Args: {'city': 'Tokyo', 'unit': 'celsius'}
# ID: call_abc123
elif response.is_final_answer:
print(response.text)
Async tool calling:
response = await llm.ainvoke_with_tools(
"What is the weather in London?", tools
)
If the model's tool-call arguments come back as malformed JSON, call.arguments falls back to {} and call.arguments_parse_error is set to a description of what went wrong — check it if a tool call ever seems to be missing arguments it should have had:
for call in response.tool_calls:
if call.arguments_parse_error:
print(f"Warning: {call.name}'s arguments failed to parse: {call.arguments_parse_error}")
Cost Tracking
Pass pricing (USD per 1 million tokens) to get cost breakdowns.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(
model="gpt-4o",
input_pricing=2.50, # $2.50 per 1M input tokens
output_pricing=10.00, # $10.00 per 1M output tokens
structured_output=True,
)
result = llm.invoke("Summarise the history of the internet in 3 sentences.")
print(result["model"]) # gpt-4o
print(result["response"]) # The internet began as ARPANET...
print(result["input_tokens"]) # 21
print(result["output_tokens"]) # 68
print(result["total_tokens"]) # 89
print(result["input_cost"]) # 0.0000525
print(result["output_cost"]) # 0.00068
print(result["total_cost"]) # 0.0007325
print(result["latency_ms"]) # 1102.4
Access the last call metadata without structured_output=True:
llm = OpenAIResponse(model="gpt-4o", input_pricing=2.50, output_pricing=10.00)
reply = llm.invoke("Hello!")
print(llm.last_metadata)
# {
# "model": "gpt-4o",
# "response": "Hello! How can I help you today?",
# "input_tokens": 9,
# "output_tokens": 10,
# "total_tokens": 19,
# "input_cost": 0.0000225,
# "output_cost": 0.0001,
# "total_cost": 0.0001225,
# "latency_ms": 921.7
# }
Budget Governor
Set max_session_cost= (USD) to hard-stop invoke()/ainvoke()/invoke_structured()/ainvoke_structured() once accumulated session cost reaches the cap — the blocked call is rejected before it reaches the API, so no further spend happens. Requires both input_pricing and output_pricing (cost can't be computed, and the cap can't trigger, without them).
from autourgos_responses import OpenAIResponse, BudgetExceededException
llm = OpenAIResponse(
model="gpt-4o",
input_pricing=2.50,
output_pricing=10.00,
max_session_cost=0.50, # hard stop at $0.50 for this client's lifetime
)
try:
for prompt in many_prompts:
reply = llm.invoke(prompt)
except BudgetExceededException as e:
print(f"Stopped: {e}")
print(f"Used ${llm.session_cost_used:.4f} of ${llm.max_session_cost:.4f}")
Call llm.reset_session_budget() to zero out session_cost_used and unblock a tripped cap (e.g. starting a new billing period without recreating the client).
Notes:
- This is a backstop, not an exact per-call prediction. A call's cost is only known after its response comes back, so the cap is checked against cost already accumulated from prior calls — the call that pushes you over the cap still completes; only the next one is blocked.
- Hitting the cap does not count toward the circuit breaker's failure threshold — a budget stop is not a provider failure.
invoke_structured()/ainvoke_structured()check the budget once before the first attempt; a failed-validation retry attempt inside that call still costs money but its cost isn't tracked intosession_cost_usedtoday (only the final successful attempt's cost is recorded).invoke_with_tools()/ainvoke_with_tools()/stream()/astream()are not budget-protected in this version — same gap as the Call Ledger, since no usage/cost metadata is computed on those paths.
Context Manager
Automatically closes the HTTP client when done.
from autourgos_responses import OpenAIResponse
with OpenAIResponse(model="gpt-4o") as llm:
reply = llm.invoke("Quick question: what is 2 + 2?")
print(reply)
# 4
# Client is closed here automatically
Async context manager:
import asyncio
from autourgos_responses import OpenAIResponse
async def main():
async with OpenAIResponse(model="gpt-4o") as llm:
reply = await llm.ainvoke("What year did the Berlin Wall fall?")
print(reply)
# The Berlin Wall fell in 1989.
asyncio.run(main())
Circuit Breaker
Protects against cascading failures. After circuit_failure_threshold consecutive API errors, all calls are blocked for circuit_cooldown_time seconds.
This is useful when you are using a local model (Ollama, LM Studio) or a rate-limited API. If the server goes down, the circuit breaker stops your code from hammering it with failed requests.
from autourgos_responses import OpenAIResponse, CircuitBreakerOpenException
llm = OpenAIResponse(
model="gpt-4o",
circuit_failure_threshold=3, # open after 3 consecutive failures
circuit_cooldown_time=60.0, # block for 60 seconds
)
try:
reply = llm.invoke("Hello!")
except CircuitBreakerOpenException as e:
print(f"Circuit is open: {e}")
# Circuit breaker OPEN for OpenAIResponse: 3 consecutive failures.
# Blocked until 1718500000.0.
The circuit automatically resets after the cooldown and allows one probe call through.
Provider Fallback Chain
Configure backup providers that invoke(), ainvoke(), stream(), astream(), invoke_with_tools(), and ainvoke_with_tools() transparently switch to if the primary provider fails (after its own retries are exhausted) — no proxy or gateway service needed.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(
model="gpt-4o",
api_key="sk-...", # primary: OpenAI
fallback_providers=[
{
"model": "llama3-70b-8192", # 1st backup: Groq
"api_key": "gsk_...",
"base_url": "https://api.groq.com/openai/v1",
},
{
"model": "llama3", # 2nd backup: local Ollama
"api_key": "ollama",
"base_url": "http://localhost:11434/v1",
},
],
)
reply = llm.invoke("What is the capital of France?")
print(reply)
# Paris (served by whichever provider succeeded first)
print(llm.last_metadata["provider_used"])
# "primary" or "fallback[0]:llama3-70b-8192" or "fallback[1]:llama3"
Each fallback entry resolves its own api_key/base_url (falling back to OPENAI_API_KEY/OPENAI_BASE_URL env vars, exactly like the primary) — nothing is inherited from the primary provider's credentials, so a backup on a different host never sees the primary's key.
If every provider fails, OpenAIResponseAllProvidersFailedError (a subclass of OpenAIResponseAPIError) is raised with an .attempts list of (label, exception) pairs, one per provider tried:
from autourgos_responses import OpenAIResponseAllProvidersFailedError
try:
llm.invoke("Hello!")
except OpenAIResponseAllProvidersFailedError as e:
for label, exc in e.attempts:
print(f"{label}: {exc}")
Streaming limitation: fallback only kicks in if a provider fails before it has streamed any text. Once partial output has already reached the caller, switching providers mid-stream would duplicate or corrupt the output, so the error is raised as-is instead of silently trying the next provider.
create()/acreate() (low-level raw access) are unaffected by fallback_providers — they always call the primary client only, since their contract is "the raw response of the client you configured."
Aggregate Call Deadline
max_retries and fallback_providers each get their own full retry budget independently — without a cap on total wall-clock time, a call that keeps failing can take roughly providers × max_retries × timeout (plus back-off sleeps) before finally returning or raising. Set max_call_duration (seconds) to cap the whole logical call — every retry attempt and every fallback provider combined:
from autourgos_responses import OpenAIResponse, OpenAIResponseDeadlineExceededError
llm = OpenAIResponse(
model="gpt-4o",
fallback_providers=[{"model": "llama3-70b-8192", "api_key": "gsk_...", "base_url": "https://api.groq.com/openai/v1"}],
max_retries=5,
max_call_duration=10.0, # give up after 10s total, across every attempt/provider
)
try:
reply = llm.invoke("Hello!")
except OpenAIResponseDeadlineExceededError as e:
print(f"Gave up after the deadline: {e}")
The deadline is checked between attempts/providers, not by cancelling a request already sent — an in-flight HTTP call stays bounded by timeout as usual, so total wall-clock time can exceed max_call_duration by up to one in-flight request's duration, but never by a further full retry or fallback cycle. None (the default) disables this entirely — retries and fallback behave exactly as before. create()/acreate() respect it too, applied to that single primary-only call.
Call Ledger (Audit Trail)
Set ledger_path= to record every invoke(), ainvoke(), invoke_structured(), and ainvoke_structured() call to a local SQLite file — model, provider used, prompt/response, tokens, cost, latency, validation retries. No external service, no extra dependency (sqlite3 is part of the Python standard library). Disabled by default (ledger_path=None) — zero overhead unless you turn it on.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(
model="gpt-4o",
input_pricing=2.50,
output_pricing=10.00,
ledger_path="calls.db", # created if it doesn't exist
)
llm.invoke("What is the capital of France?")
llm.invoke("What is the capital of Japan?")
Query it with any SQLite tool:
sqlite3 calls.db "SELECT created_at, model, provider_used, total_cost, latency_ms FROM calls ORDER BY id;"
Set ledger_store_content=False to log only tokens/cost/latency/provider metadata — no prompt/response text — if you don't want request content persisted to disk:
llm = OpenAIResponse(model="gpt-4o", ledger_path="calls.db", ledger_store_content=False)
Notes:
- A ledger write happens synchronously on every logged call (one
INSERT+commit) — fine for audit/dev/debugging, but adds I/O latency in a tight high-throughput loop. It's not meant for a hot production path. - A ledger write can never break your actual LLM call: any failure (disk full, permissions, a closed connection) is logged as a warning and swallowed.
invoke_with_tools()/ainvoke_with_tools()/stream()/astream()are not logged in this version — they don't compute usage/cost metadata today.- The ledger connection is closed automatically by the context manager (
with OpenAIResponse(...) as llm:).
Shadow-Mode Dual Dispatch
Dispatch the same prompt to one or more "shadow" providers concurrently with the primary, purely for observation — invoke()/ainvoke() always return the primary's answer. Useful for catching regressions before switching a default model/provider, or for ongoing quality/cost comparison.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(
model="gpt-4o", # primary — this is what invoke() returns
shadow_providers=[
{"model": "gpt-4o-mini"}, # compare against a cheaper model
{
"model": "llama3-70b-8192", # and a different provider entirely
"api_key": "gsk_...",
"base_url": "https://api.groq.com/openai/v1",
},
],
)
reply = llm.invoke("What is the capital of France?")
print(reply)
# Paris (always from the primary — gpt-4o)
for shadow in llm.last_shadow_results:
print(shadow)
React to results live with on_shadow_result=:
def alert_on_drift(shadow_result):
if shadow_result["similarity"] is not None and shadow_result["similarity"] < 0.5:
print(f"Drift detected from {shadow_result['provider_used']}!")
llm = OpenAIResponse(model="gpt-4o", shadow_providers=[...], on_shadow_result=alert_on_drift)
Notes:
- Adds latency: primary and shadow providers run concurrently, so total call time is roughly
max(primary_latency, slowest_shadow_latency)— not the sum, but not zero overhead either.invoke()waits for every shadow provider to finish (or fail) before returning. - Costs real money: each shadow provider gets one live API call per invocation. This cost is tracked in each shadow result's
total_costbut is not added tosession_cost_used/ counted againstmax_session_cost. - Each shadow provider gets a single attempt — no retries. A shadow failure never raises and never affects the primary's result; it just shows up with
errorset inlast_shadow_results. - Shadow requests are sent concurrently with the primary's, not after it succeeds — so if the primary call itself ultimately fails (retries exhausted, all providers failed, etc.), any shadow request already in flight still completes and is still recorded in
last_shadow_results/the ledger (withsimilarity=None, since there's no primary text to compare against). The primary's own exception still propagates normally to the caller either way. - Only
invoke()/ainvoke()/chat()/achat()dispatch shadows in this version —stream()/astream()/invoke_with_tools()/invoke_structured()don't. - If Call Ledger is enabled, every shadow result is also recorded in a separate
shadow_callstable.
PII / Secret Redaction
This is a heuristic, best-effort scrubber, not a compliance-grade DLP solution. It's regex-based: it will miss PII that doesn't match a known pattern, and it will occasionally mask legitimate content that happens to match a pattern. Use it as one layer of defense-in-depth, not a guarantee. Disabled by default.
Set redact_pii=True to scan the resolved prompt for likely secrets/PII before it's sent to the provider — covers every call path (invoke, ainvoke, stream, astream, invoke_with_tools, ainvoke_with_tools, invoke_structured, ainvoke_structured). Built-in categories: email, credit_card, ssn, phone, api_key.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(model="gpt-4o", redact_pii=True)
reply = llm.invoke("My email is bob@example.com and my key is sk-abc123...")
# The model actually receives:
# "My email is [REDACTED:email] and my key is [REDACTED:api_key]"
print(llm.last_redacted_categories)
# ["email", "api_key"]
Restrict to specific categories, or add your own patterns:
llm = OpenAIResponse(
model="gpt-4o",
redact_pii=True,
redact_categories=["email", "api_key"], # skip credit_card/ssn/phone
redact_custom_patterns={"internal_id": r"EMP-\d{5}"},
)
Bring your own dictionary of exact/literal values instead of writing regex for everything (redact_custom_terms), or point redact_patterns_file at a shared JSON file with "patterns"/"terms" keys — see autourgos-openaichat's README for the full walkthrough (this package shares the same redaction engine and constructor parameters).
Use redact_mode="block" to reject the call outright instead of masking and proceeding:
from autourgos_responses import OpenAIResponseRedactionBlockedError
llm = OpenAIResponse(model="gpt-4o", redact_pii=True, redact_mode="block")
try:
llm.invoke("My email is bob@example.com")
except OpenAIResponseRedactionBlockedError as e:
print(f"Blocked, matched: {e.categories_found}")
# Blocked, matched: ['email']
Set redact_restore_in_response=True (requires redact_pii=True and redact_mode="mask") to have the final result swap masked placeholders back for their real values — the model itself still never sees the real value, only whatever it echoes back gets restored client-side. The Call Ledger always records the still-masked text regardless of this setting.
Constrained Decoding / Provider-Specific Params
Self-hosted OpenAI-compatible servers (vLLM, llama.cpp, and others) support extra, non-standard request fields for constrained/guided generation. Set extra_body= on the constructor to merge your own fields into every request this client makes (primary, fallback, and shadow providers alike).
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(
model="my-local-model",
base_url="http://your-vllm-server:8000/v1",
api_key="EMPTY",
extra_body={"guided_json": {"type": "object", "properties": {"answer": {"type": "string"}}}},
)
Low-Level Access
Direct access to the raw Responses API response object when you need full control.
from autourgos_responses import OpenAIResponse
llm = OpenAIResponse(model="gpt-4o")
raw = llm.create("Explain gravity briefly.")
print(raw.output_text)
print(raw.usage.input_tokens)
print(raw.usage.output_tokens)
Async:
raw = await llm.acreate("Explain gravity briefly.")
print(raw.output_text)
With overrides:
raw = llm.create(
"Summarise this.",
temperature=0.3,
max_output_tokens=50,
)
Error Handling
from autourgos_responses import (
OpenAIResponse,
OpenAIResponseAPIError,
OpenAIResponseDeadlineExceededError,
OpenAIResponseResponseError,
OpenAIResponseConfigError,
OpenAIResponseImportError,
OpenAIResponseValidationError,
OpenAIResponseRedactionBlockedError,
OpenAIResponseAllProvidersFailedError,
CircuitBreakerOpenException,
BudgetExceededException,
)
llm = OpenAIResponse(model="gpt-4o")
try:
reply = llm.invoke("Hello!")
except OpenAIResponseAllProvidersFailedError as e:
# every provider (primary + all fallback_providers) failed
print(f"All providers failed: {e.attempts}")
except OpenAIResponseDeadlineExceededError as e:
# max_call_duration was exceeded partway through retries/fallback
print(f"Deadline exceeded: {e}")
except OpenAIResponseAPIError as e:
# API request failed after all retries (or immediately on a non-retryable 4xx)
print(f"API error: {e}")
except OpenAIResponseResponseError as e:
# Response was received but text could not be extracted
print(f"Response parse error: {e}")
except OpenAIResponseValidationError as e:
# invoke_structured()/ainvoke_structured() still invalid after retries
print(f"Validation error: {e.validation_error}")
except OpenAIResponseConfigError as e:
# Incompatible options (e.g. streaming + structured_output)
print(f"Config error: {e}")
except OpenAIResponseImportError as e:
# openai SDK not installed
print(f"Import error: {e}")
except OpenAIResponseRedactionBlockedError as e:
# redact_pii=True, redact_mode="block", and a match was found
print(f"Blocked: {e.categories_found}")
except CircuitBreakerOpenException as e:
# Too many recent failures, circuit is open
print(f"Circuit open: {e}")
except BudgetExceededException as e:
# max_session_cost reached
print(f"Budget exceeded: {e}")
Retry behaviour: by default the wrapper retries up to 3 times with exponential back-off, but fails immediately (no retry) on non-retryable client errors — HTTP 400, 401, 403, 404, 422.
| Attempt | Wait before retry |
|---|---|
| 1st failure | 0.5 s |
| 2nd failure | 1.0 s |
| 3rd failure | 2.0 s |
| 4th failure | raises OpenAIResponseAPIError |
Change with max_retries and backoff_factor:
llm = OpenAIResponse(
model="gpt-4o",
max_retries=5,
backoff_factor=1.0, # waits: 1s, 2s, 4s, 8s then raises
)
That retry budget applies per provider — add fallback_providers and each backup gets its own full retry budget too. To cap the total time across all of them, see Aggregate Call Deadline.
Constructor Reference
| Parameter | Type | Default | Description |
|---|---|---|---|
model |
str |
required | Model name. e.g. "gpt-4o", "o3-mini", "llama3-70b-8192", "gemini-2.0-flash" |
api_key |
str |
OPENAI_API_KEY env |
API key for the provider you are using |
base_url |
str |
OPENAI_BASE_URL env |
Provider endpoint. e.g. "https://api.groq.com/openai/v1" or "http://localhost:11434/v1" |
organization |
str |
None |
OpenAI organization ID (OpenAI only) |
project |
str |
None |
OpenAI project ID (OpenAI only) |
system_prompt |
str |
None |
System prompt sent as the instructions field |
prompt_template |
str |
None |
Template with {variable} placeholders |
temperature |
float |
None |
Sampling temperature 0 to 2. Higher = more random |
top_p |
float |
None |
Nucleus sampling 0 to 1 |
max_tokens |
int |
None |
Maximum output tokens (maps to max_output_tokens) |
reasoning_effort |
str |
None |
"low", "medium", or "high" — for o3, o3-mini, o1 only |
reasoning_summary |
str |
None |
Include a reasoning summary in output (OpenAI only) |
text_verbosity |
str |
None |
"low", "medium", or "high" |
output_schema |
BaseModel / dict |
None |
Pydantic model or JSON schema for structured output |
response_mime_type |
str |
None |
"application/json" enables JSON object mode |
structured_output |
bool |
False |
If True, invoke() returns a metadata dict |
streaming |
bool |
False |
If True, invoke() streams internally and joins |
max_retries |
int |
3 |
Retry attempts on transient API errors |
timeout |
float |
60.0 |
Request timeout in seconds |
backoff_factor |
float |
0.5 |
Exponential back-off base (wait = factor x 2^attempt) |
max_call_duration |
float |
None |
Aggregate wall-clock budget in seconds across every retry attempt and fallback provider (see Aggregate Call Deadline) |
input_pricing |
float |
None |
USD per 1 million input tokens |
output_pricing |
float |
None |
USD per 1 million output tokens |
circuit_failure_threshold |
int |
5 |
Consecutive failures before the circuit opens |
circuit_cooldown_time |
float |
30.0 |
Seconds the circuit stays open before probing |
fallback_providers |
list[dict] |
None |
Backup providers tried in order after the primary exhausts retries. Each dict: model (required), api_key/base_url/organization/project (optional) |
ledger_path |
str |
None |
If set, logs every invoke/ainvoke/invoke_structured/ainvoke_structured call to a local SQLite file |
ledger_store_content |
bool |
True |
If False, the ledger logs only tokens/cost/latency/provider — no prompt/response text |
max_session_cost |
float |
None |
Hard-stop cap in USD; requires input_pricing/output_pricing. Raises BudgetExceededException once reached |
redact_pii |
bool |
False |
Scan the resolved prompt for secrets/PII before sending |
redact_categories |
list[str] |
all | Which built-in categories to scan for: email, credit_card, ssn, phone, api_key |
redact_mode |
str |
"mask" |
"mask" replaces matches and proceeds; "block" raises OpenAIResponseRedactionBlockedError instead |
redact_custom_patterns |
dict |
None |
Extra {name: regex} entries merged in alongside the built-ins |
redact_custom_terms |
dict |
None |
Bring-your-own dictionary of exact/literal values, as {category: [values]} |
redact_patterns_file |
str |
None |
Path to a JSON file with "patterns"/"terms" keys, loaded once at construction |
redact_restore_in_response |
bool |
False |
Swap masked placeholders back for real values in the returned text. Requires redact_pii=True, redact_mode="mask" |
shadow_providers |
list[dict] |
None |
Providers dispatched concurrently with the primary, for comparison only. Same entry shape as fallback_providers |
on_shadow_result |
callable |
None |
Callback invoked with each shadow result dict as it completes |
extra_body |
dict |
None |
Raw provider-specific request fields merged into every request (primary, fallback, shadow) |
API Reference
What Each Method Returns
| Method | Returns |
|---|---|
invoke(prompt, **overrides) |
str, generated text (or dict if structured_output=True). **overrides (raw Responses API params, e.g. temperature=, top_p=, max_output_tokens=) apply to this call only, across the fallback chain; "input"/"model"/"stream" can't be overridden this way |
ainvoke(prompt, **overrides) |
same as invoke, async |
stream(prompt, **overrides) |
Iterator[str], text chunks. Same per-call **overrides as invoke |
astream(prompt, **overrides) |
AsyncIterator[str], text chunks. Same per-call **overrides as invoke |
batch_invoke(prompts) |
list[str], one result per prompt, sequential |
abatch_invoke(prompts) |
list[str], concurrent results |
chat(messages) |
str, generated text (or dict if structured_output=True) |
achat(messages) |
same as chat, async |
invoke_with_tools(prompt, tools) |
ToolCallResponse, .tool_calls list or .text |
ainvoke_with_tools(prompt, tools) |
same as invoke_with_tools, async |
invoke_structured(prompt) |
Validated instance of output_schema (raises OpenAIResponseValidationError on exhaustion) |
ainvoke_structured(prompt) |
same as invoke_structured, async |
create(input_data) |
Raw Responses API Response object |
acreate(input_data) |
same as create, async |
ToolCallResponse fields
| Field | Type | Description |
|---|---|---|
.tool_calls |
list[FunctionCall] |
Tool calls the model wants to make (empty if final answer) |
.text |
str | None |
Final text answer (None if tool calls present) |
.raw |
Any |
Raw provider response object |
.has_tool_calls |
bool |
True when tool_calls is non-empty |
.is_final_answer |
bool |
True when text is present and tool_calls is empty |
FunctionCall fields
| Field | Type | Description |
|---|---|---|
.name |
str |
Tool function name |
.arguments |
dict |
Parsed JSON arguments ({} if parsing failed — see .arguments_parse_error) |
.call_id |
str | None |
Call ID for multi-turn tracking |
.arguments_parse_error |
str | None |
Set to the parse error message when the model's JSON arguments failed to parse; None on success |
Metadata dict (when structured_output=True, or via llm.last_metadata)
| Key | Type | Description |
|---|---|---|
"model" |
str |
Model name used |
"response" |
str |
Generated text |
"input_tokens" |
int | None |
Input token count |
"output_tokens" |
int | None |
Output token count |
"total_tokens" |
int | None |
Total token count |
"input_cost" |
float |
Input cost in USD (only if input_pricing set) |
"output_cost" |
float |
Output cost in USD (only if output_pricing set) |
"total_cost" |
float |
Total cost in USD (only if both pricing set) |
"latency_ms" |
float |
Request round-trip time in milliseconds |
"provider_used" |
str |
"primary" or "fallback[N]:<model>" — which provider actually served the request |
"validation_retries" |
int |
Only set after invoke_structured()/ainvoke_structured() — number of correction attempts needed (0 = valid on first try) |
Differences vs autourgos-openaichat
Both packages implement the same feature set (fallback chain, budget governor, call ledger, shadow-mode dispatch, PII redaction, native tool calling, validated structured output) — autourgos-responses builds directly on autourgos-openaichat's shared base layer. The real differences are which underlying API endpoint each one targets and the handful of things unique to the Responses API:
| Feature | autourgos-openaichat | autourgos-responses |
|---|---|---|
| API endpoint | chat.completions.create |
responses.create |
| System prompt field | messages[0].role = "system" |
instructions parameter |
| Reasoning models | Not supported | reasoning_effort param for o3/o1 |
| Text verbosity control | Not supported | text_verbosity param |
| Multi-turn input | Messages list | Messages list (via chat()) or plain string |
| Native tool calling | Supported (invoke_with_tools) |
Supported (invoke_with_tools) |
| Use when | Building chat agents, tool-calling on Chat Completions-only providers | Using reasoning models, or any provider that supports the Responses API |
Both packages support the same providers via base_url. Choose based on the API endpoint your use case needs.
License
Apache License 2.0, Copyright (c) 2026 Jitin Kumar Sengar
Metadata
Release files for autourgos-responses 2.6.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| autourgos_responses-2.6.0.tar.gz | 103.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| autourgos_responses-2.6.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 153.5 kB
Release files / autourgos_responses-2.6.0.tar.gz
| Download URL | autourgos_responses-2.6.0.tar.gz |
|---|---|
| Size | 103.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
18a1ddafbd885aa8abf2f08b994457c9e08bf867ec0504ed4dc4a27437d10854
|
|
BLAKE2b-256 checksum How to use checksums |
fbf27ca79409ec757411b6df95739310057a379aeb5c2efa83a9f290f399e9c6
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.9
|
Release files / autourgos_responses-2.6.0-py3-none-any.whl
| Download URL | autourgos_responses-2.6.0-py3-none-any.whl |
|---|---|
| Size | 49.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
42f35f903832822dd58022b446a5f31c7fd571b3fa36c130d712ce2771da6e32
|
|
BLAKE2b-256 checksum How to use checksums |
c83d54e775f59d35e15b428a29d2d0926dfc67c8f445cbae3309e0284079d0d0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.9
|