A model-agnostic Python library with a unified API for multiple LLM providers
Project description
coffee_with_llm
A model-agnostic Python library providing a unified API for OpenAI, Anthropic Claude, and Google Gemini.
Features
- Model-agnostic: Automatically selects the appropriate provider based on model name
- Unified API: Same interface for OpenAI, Anthropic Claude, and Google Gemini
- Tool Calling: Full support for OpenAI's tool calling with multi-step execution
- Structured Outputs: Support for JSON schema and response formatting
- Caching: Built-in prompt caching for both providers
- Citations: Automatic citation injection for Google Gemini responses
- Reasoning: Support for OpenAI's reasoning models with effort control
Installation
pip install coffee_with_llm
Quick Start
import asyncio
from coffee_with_llm import AskLLM
async def main():
# Initialize with any model (model is required)
llm = AskLLM(model="gpt-5.4")
# Simple question
response = await llm.ask(
prompt="What is Python?",
system_instruct="You are a helpful assistant."
)
print(response.text)
# Use Google Gemini
llm_gemini = AskLLM(model="gemini-3.1-pro-preview")
response = await llm_gemini.ask(
prompt="Explain quantum computing",
system_instruct="You are a physics expert."
)
print(response.text)
asyncio.run(main())
Configuration
Set environment variables for API keys:
export OPENAI_API_KEY="your-openai-key"
export ANTHROPIC_API_KEY="your-anthropic-key"
export GOOGLE_API_KEY="your-google-key"
Usage Examples
Basic Usage
from coffee_with_llm import AskLLM
# Model parameter is required
llm = AskLLM(model="gpt-5.4")
response = await llm.ask(prompt="Hello, world!")
With System Instructions
response = await llm.ask(
prompt="Write a haiku about coding",
system_instruct="You are a creative poet."
)
Structured Outputs (JSON Schema)
response = await llm.ask(
prompt="Extract key information from: 'John Doe, age 30, works at Acme Corp'",
response_format={
"type": "json_schema",
"json_schema": {
"name": "person_info",
"schema": {
"type": "object",
"properties": {
"name": {"type": "string"},
"age": {"type": "number"},
"company": {"type": "string"}
}
}
}
}
)
Tool Calling (OpenAI, Anthropic, Google)
def get_weather(location: str) -> dict:
# Your tool implementation
return {"temperature": 72, "condition": "sunny"}
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get weather for a location",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string"}
}
}
}
}
]
async def execute_tool(name: str, args: dict) -> dict:
if name == "get_weather":
result = get_weather(args.get("location", ""))
return {"ok": True, "result": result}
return {"ok": False, "error": "Unknown tool"}
response = await llm.ask(
prompt="What's the weather in San Francisco?",
tools_schema=tools,
execute_tool_cb=execute_tool
)
Reasoning Models (OpenAI)
llm = AskLLM(model="gpt-5.4")
response = await llm.ask(
prompt="Solve this math problem: 2x + 5 = 15",
reasoning_effort="high"
)
Streaming
llm = AskLLM(model="gpt-5.4")
result = await llm.ask(prompt="Explain recursion in programming.", stream=True)
async for chunk in result:
print(chunk, end="", flush=True)
print(f"\nUsage: {result.usage.input_tokens} in, {result.usage.output_tokens} out")
Note: Streaming is not supported with
tools_schemaorresponse_format.result.usageis populated after iteration completes. Rate limits trigger retry before the first chunk.
Supported Models
OpenAI
gpt-5.4,gpt-5.4-pro(flagship, reasoning)gpt-5.3-instant,gpt-5-mini,gpt-5-nano(fast, cost-effective)gpt-4o,gpt-4o-mini- Any OpenAI model name
Anthropic Claude
claude-opus-4-6,claude-sonnet-4-6(latest)claude-haiku,claude-3-5-sonnet- Any Claude model name (claude-* prefix)
Google Gemini
gemini-3.1-pro-preview,gemini-3.1-flash(latest)gemini-2.5-pro,gemini-2.5-flash- Any Google Gemini model name
API Reference
AskLLM
__init__(*, model, config=None, ...)
Initialize the LLM client.
Parameters:
model(str): Model name (provider auto-detected, required)config(Config, optional): Config instance. If None, usesConfig.from_env()for API keysmin_delay_between_calls(float, optional): Min delay between API calls in seconds (default: 1.0)max_retries(int, optional): Max retries for rate limit errors (default: 3)request_timeout(float, optional): Request timeout in seconds (default: 60)google_explicit_cache(bool, optional): Enable Google context caching (default: True)google_inline_citations(bool, optional): Inject[cite: url]markers for Gemini grounding (default: True)
ask(...)
Generate a response from the LLM.
Parameters:
prompt(str): User prompt/questionsystem_instruct(str, optional): System instructionmessages(list, optional): Conversation historymax_tokens(int, optional): Maximum tokens to generatetemperature(float, optional): Sampling temperature (0-2)top_p(float, optional): Nucleus sampling parameterpresence_penalty(float, optional): Presence penalty (OpenAI only)reasoning_effort(str, optional): Reasoning effort level (OpenAI only)tools_schema(list, optional): Tool/function calling schema (OpenAI, Anthropic, Google)response_format(dict, optional): Response format specificationexecute_tool_cb(callable, optional): Tool execution callback (OpenAI, Anthropic, Google)tool_error_callback(callable, optional): Callback when tool returns ok=Falsemax_steps(int, optional): Maximum tool-calling steps (default: 24)max_effective_tool_steps(int, optional): Maximum effective tool steps (default: 12)force_tool_use(bool, optional): Force at least one tool call when tools provided (default: False)stream(bool, optional): When True, returnStreamResult(default: False)
Returns: AskResult – Object with .text (str) and .usage (TokenUsage). When stream=True, returns StreamResult – async iterable of text chunks; .usage populated after iteration completes.
Raises:
ValidationError: If prompt is empty or invalid parameters providedAPIError: If the API call failsConfigurationError: If API keys are missing or client initialization fails
Environment Variables
OPENAI_API_KEY: OpenAI API key (required for OpenAI models)ANTHROPIC_API_KEY: Anthropic API key (required for Claude models)GOOGLE_API_KEY: Google API key (required for Google models)
Google-specific options (google_explicit_cache, google_inline_citations) are passed as constructor params to AskLLM; see docstring.
License
MIT
Contributing
Contributions welcome! Please open an issue or submit a pull request.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file coffee_with_llm-0.1.0.tar.gz.
File metadata
- Download URL: coffee_with_llm-0.1.0.tar.gz
- Upload date:
- Size: 38.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.4
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
259c8da56afe418f327eda107b37f520b0a5d3cbdb1654c5b563afb692580fe7
|
|
| MD5 |
3a5f249775fdda9e39a70339fc64f89b
|
|
| BLAKE2b-256 |
79be9c394c5519b095a788f9c14af3027cbeb6b1e6f3b71a197b06a51125ee28
|
File details
Details for the file coffee_with_llm-0.1.0-py3-none-any.whl.
File metadata
- Download URL: coffee_with_llm-0.1.0-py3-none-any.whl
- Upload date:
- Size: 35.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.4
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
773edad93ce43994e414147cd2d010b92811fa3ad4a1ebe3fd3aaedd3975c90b
|
|
| MD5 |
2a6f27f78a8c747a3a8dd0890dd3a0f2
|
|
| BLAKE2b-256 |
17e784d38de0c3fd3cf95e77635496e0cc47560744de5e4dddf84b73aa758b63
|