Skip to main content

coffee_with_llm

A model-agnostic Python library providing a unified API for OpenAI, Anthropic Claude, Google Gemini, and Inception Mercury.

Features

  • Model-agnostic: Picks the provider from the model id, or from an explicit provider/model prefix (see below)
  • Unified API: Same interface for OpenAI, Anthropic Claude, Google Gemini, and Inception Mercury
  • Tool Calling: Full support for OpenAI's tool calling with multi-step execution
  • Structured Outputs: Support for JSON schema and response formatting
  • Caching: Google explicit context cache; Anthropic automatic prompt cache; OpenAI/Inception pass through cached_tokens when present
  • Attachments: Send PDFs and images as input with one Attachment type; each provider translates it to its own native format (Inception Mercury is text-only)
  • Citations: Automatic citation injection for Google Gemini responses
  • Grounded flows: Two-pass search → JSON (ask_with_grounded_json) or markdown (ask_with_grounded_markdown) with citation allowlists
  • Gemini Interactions API: Optional ask_interaction() for server-side multi-turn sessions
  • Extended thinking: Single reasoning_effort knob ("low" | "medium" | "high", plus "instant" on Inception) translates to OpenAI reasoning.effort, Anthropic adaptive output_config.effort (4.6+) or legacy thinking.budget_tokens, Google thinking_config, and Inception reasoning_effort

Installation

pip install coffee_with_llm

Quick Start

import asyncio
from coffee_with_llm import AskLLM

async def main():
    # Initialize with any model (model is required)
    llm = AskLLM(model="gpt-5.4")
    
    # Simple question
    response = await llm.ask(
        prompt="What is Python?",
        system_instruct="You are a helpful assistant."
    )
    print(response.text)
    
    # Use Google Gemini
    llm_gemini = AskLLM(model="gemini-3.1-pro-preview")
    response = await llm_gemini.ask(
        prompt="Explain quantum computing",
        system_instruct="You are a physics expert."
    )
    print(response.text)

    # Google Gemma (or any Google API model id): prefix selects Google; only the part
    # after "/" is sent to the API
    llm_gemma = AskLLM(model="google/gemma-2-9b-it")
    response = await llm_gemma.ask(prompt="Say hi in five words or fewer.")
    print(response.text)

asyncio.run(main())

Provider prefix (provider/model)

You can force the provider and pass the exact API model id with a single slash:

Prefix Provider Example
google/ Google google/gemma-2-9b-it
gemini/ Google gemini/gemini-2.5-flash
openai/ OpenAI openai/gpt-4o-mini
anthropic/ Anthropic anthropic/claude-sonnet-4-6
claude/ Anthropic claude/claude-sonnet-4-6
inception/ Inception inception/mercury-2

Only the segment after the first / is sent to the provider API (e.g. google/gemma-2-9b-it → API model gemma-2-9b-it). If the prefix is missing or unknown, the whole string is used for both routing and the API (legacy behavior: ids like gpt-4o-mini, claude-…, gemini-…, google-…, mercury-… still auto-route).

An empty id after the prefix (e.g. google/) is rejected with ValidationError.

Configuration

Copy .env.example to .env and set the keys you need. Config.from_env() loads .env automatically — no export required.

cp .env.example .env
OPENAI_API_KEY=...
ANTHROPIC_API_KEY=...
GOOGLE_API_KEY=...
INCEPTION_API_KEY=...

Smoke-test Inception wiring (live API)::

uv run scripts/test_inception_wiring.py

Live Gemini citation smoke tests (require GOOGLE_API_KEY in .env)::

uv run scripts/test_json_hook_citations.py      # curator JSON hooks
uv run scripts/test_interactions_api.py         # Interactions API two-turn
uv run scripts/test_markdown_web_citations.py   # orchestrator web + markdown

Each script logs per-stage latency and token usage. See CHANGELOG.md for release notes.

Grounded search + citations (Gemini)

When JSON output and Google Search are combined in one call, Gemini often omits grounding metadata. Use two-pass helpers instead:

from coffee_with_llm import AskLLM, ask_with_grounded_json, ask_with_grounded_markdown

llm = AskLLM(model="google/gemini-flash-latest")

# Curator-style JSON card hooks
result = await ask_with_grounded_json(
    llm,
    research_prompt="Research two cards about …",
    json_prompt="Convert your notes into a JSON array of cards …",
)

# Orchestrator-style markdown reply
web, markdown, allowed_urls = await ask_with_grounded_markdown(
    llm,
    web_query="What is new in …?",
    markdown_prompt="Write a markdown answer with cited bullets …",
)

Pass 1 searches the web; pass 2 formats output using only URLs from pass 1.

Gemini Interactions API

For server-side session state (multi-turn agent flows):

first = await llm.ask_interaction(
    prompt="Search and summarize …",
    system_instruct="Use web search.",
)
follow_up = await llm.ask_interaction(
    prompt="Turn that into JSON …",
    previous_interaction_id=first.interaction_id,
)

Requires google-genai ≥ 2.0.0. Default ask() still uses generateContent.

Usage Examples

Basic Usage

from coffee_with_llm import AskLLM

# Model parameter is required
llm = AskLLM(model="gpt-5.4")
response = await llm.ask(prompt="Hello, world!")

With System Instructions

response = await llm.ask(
    prompt="Write a haiku about coding",
    system_instruct="You are a creative poet."
)

Structured Outputs (JSON Schema)

response = await llm.ask(
    prompt="Extract key information from: 'John Doe, age 30, works at Acme Corp'",
    response_format={
        "type": "json_schema",
        "json_schema": {
            "name": "person_info",
            "schema": {
                "type": "object",
                "properties": {
                    "name": {"type": "string"},
                    "age": {"type": "number"},
                    "company": {"type": "string"}
                }
            }
        }
    }
)

Tool Calling (OpenAI, Anthropic, Google, Inception)

def get_weather(location: str) -> dict:
    # Your tool implementation
    return {"temperature": 72, "condition": "sunny"}

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get weather for a location",
            "parameters": {
                "type": "object",
                "properties": {
                    "location": {"type": "string"}
                }
            }
        }
    }
]

async def execute_tool(name: str, args: dict) -> dict:
    if name == "get_weather":
        result = get_weather(args.get("location", ""))
        return {"ok": True, "result": result}
    return {"ok": False, "error": "Unknown tool"}

response = await llm.ask(
    prompt="What's the weather in San Francisco?",
    tools_schema=tools,
    execute_tool_cb=execute_tool
)

Extended Thinking (provider-agnostic)

reasoning_effort is a single string — "low", "medium", or "high" (Inception also accepts "instant") — that each provider translates to its native extended-thinking config. Pass None (or omit) to disable. Unknown values are ignored with a warning.

Effort Anthropic (4.6+) Anthropic (legacy) / Google Inception
instant reasoning_effort=instant
low output_config.effort 1,024 token budget reasoning_effort=low
medium (adaptive thinking) 4,096 token budget reasoning_effort=medium
high 16,384 token budget reasoning_effort=high
# Same call works on OpenAI, Anthropic, Google, or Inception
llm = AskLLM(model="anthropic/claude-sonnet-4-6")
response = await llm.ask(
    prompt="Solve this math problem: 2x + 5 = 15",
    reasoning_effort="high",
)

Provider-specific notes:

  • OpenAI — passed through as reasoning.effort.
  • Anthropic — on Claude Opus/Sonnet 4.6+, Gen 5, and Mythos, sets thinking={"type": "adaptive"} and output_config={"effort": ...} when reasoning_effort is set. On Sonnet 5 with no effort, explicitly disables thinking (API default is on). On Fable/Mythos 5 with no effort, uses adaptive at low effort (thinking cannot be disabled). On older models, sets thinking={"type": "enabled", "budget_tokens": N} (legacy); then temperature is forced to 1, top_p / top_k are dropped, and max_tokens is widened to leave ~1024 tokens for the visible answer. Short model aliases such as sonnet-5 normalize to claude-sonnet-5. Sampling params are omitted on Opus 4.7+, Sonnet 5, and Fable/Mythos 5 (API returns 400 otherwise). With anthropic_prompt_cache=True (default), adds top-level cache_control={"type": "ephemeral"} for automatic multi-turn caching; TokenUsage.cached_tokens reflects cache_read_input_tokens; cache_creation_tokens reflects cache writes; both are included in cost_usd.
  • Google (Gemini 2.5+ / 3.x) — sets thinking_config with include_thoughts=False so only the final answer streams to the caller.
  • Inception (Mercury) — passed as reasoning_effort on Chat Completions (instant | low | medium | high). Supports tool calling and structured outputs; text-only (no attachments).

Attachments (PDFs and images)

Build one Attachment and it works on OpenAI, Anthropic, and Google — each translates it into its own native content part (Anthropic document/image blocks, OpenAI input_file/input_image parts, Google inline_data parts). Inception Mercury is text-only and raises ValidationError if attachments are passed.

from coffee_with_llm import AskLLM, Attachment

llm = AskLLM(model="claude-opus-4-8")  # or gpt-…, or gemini-…

result = await llm.ask(
    prompt="Which model wins on each benchmark? Read the charts.",
    attachments=[Attachment.from_path("model-card.pdf")],
)
print(result.text)

Construct from bytes when the file never touches disk:

Attachment(data=pdf_bytes, mime_type="application/pdf", filename="report.pdf")

Notes:

  • Input-only. Attachments are what the model reads; the response is always text.
  • Attached to the prompt turn, never to messages history.
  • Works with stream=True — the binary goes in, the answer streams out.
  • Supported types: application/pdf, image/png, image/jpeg, image/gif, image/webp. Anything else raises ValidationError before any network call.
  • Attachments are re-sent (and re-billed) on every request; providers do not index or persist them. For repeated questions about the same document, enable the provider's context/prompt cache.

Streaming

from coffee_with_llm import StreamTextDelta

llm = AskLLM(model="gpt-5.4")
result = await llm.ask(prompt="Explain recursion in programming.", stream=True)
async for chunk in result:
    if isinstance(chunk, StreamTextDelta):
        print(chunk.text, end="", flush=True)
print(f"\nUsage: {result.usage.input_tokens} in, {result.usage.output_tokens} out")
print(f"Prompt tokens (all buckets): {result.usage.prompt_tokens}")
print(f"Billable tokens: {result.usage.billable_tokens}")
print(result.usage.to_dict())

Events: Iteration yields StreamTextDelta, StreamToolCallStart, StreamToolArgumentsDelta, StreamToolCallEnd, and StreamStepBoundary (between tool rounds), then completes. Bare str from older mocks is accepted and normalized to StreamTextDelta.

Tools and schema: stream=True works with tools_schema and response_format when the provider supports them; you must pass execute_tool_cb whenever tools_schema is set.

Gemini: Streaming with custom tools uses function calling only (no Google Search grounding in the same streaming request).

Usage and cost: result.usage (including cost_usd) is set when the stream finishes normally. If you stop early (break), call await result.aclose() so usage can be filled from the best-effort StreamUsageSink when the provider reported partial usage.

Understanding token usage

TokenUsage exposes provider-native buckets plus computed totals:

Field Meaning
input_tokens Uncached input billed at the full input rate
cached_tokens Cache reads (discounted on Anthropic/OpenAI)
cache_creation_tokens Cache writes (Anthropic; billed at 125% of input)
output_tokens Generated output
total_tokens Legacy: input_tokens + output_tokens only
prompt_tokens All prompt-side tokens (input + cache read + cache write)
billable_tokens prompt_tokens + output_tokens
cost_usd Estimated USD (includes cache buckets when priced)

On Anthropic with prompt caching enabled (default), the first turn of a long system prompt often looks like input_tokens=2 with most prompt tokens in cache_creation_tokens — not a metering bug. Use usage.to_dict() for logging and dashboards.

payload = result.usage.to_dict()
# {"input_tokens": 2, "output_tokens": 8158, "cache_creation_tokens": 40000,
#  "prompt_tokens": 40002, "billable_tokens": 48160, "cost_usd": 0.28, ...}

Rate limits trigger retry before the first chunk.

Supported Models

OpenAI

  • gpt-5.4, gpt-5.4-pro (flagship, reasoning)
  • gpt-5.3-instant, gpt-5-mini, gpt-5-nano (fast, cost-effective)
  • gpt-4o, gpt-4o-mini
  • Any OpenAI model name

Anthropic Claude

  • claude-sonnet-5, claude-opus-4-8, claude-sonnet-4-6 (latest)
  • Short aliases after provider prefix: anthropic/sonnet-5, anthropic/fable-5
  • claude-haiku, claude-3-5-sonnet
  • Any Claude model name (claude-* prefix)

Google (Gemini, Gemma, …)

  • gemini-3.1-pro-preview, gemini-3.1-flash (latest)
  • gemini-2.5-pro, gemini-2.5-flash
  • Any Gemini id the API accepts (gemini-… or google-… still auto-routes to Google)
  • Gemma and other non-gemini Google ids: use the prefix form, e.g. google/gemma-2-9b-it, so routing hits Google and the API receives gemma-2-9b-it

Inception (Mercury)

  • mercury-2 (chat; also inception/mercury-2)
  • Short alias: mercurymercury-2
  • mercury-* ids auto-route to Inception
  • Text-only Chat Completions (tool calling + structured outputs + reasoning_effort)

API Reference

AskLLM

__init__(*, model, config=None, ...)

Initialize the LLM client.

Parameters:

  • model (str): Model name (required). Optional provider/model form (see Provider prefix); otherwise provider is inferred from the id (gpt-…, claude-…, gemini-…, google-…, mercury-…, etc.).
  • config (Config, optional): Config instance. If None, uses Config.from_env() for API keys
  • min_delay_between_calls (float, optional): Min delay between API calls in seconds (default: 1.0)
  • max_retries (int, optional): Max retries for rate limit errors (default: 3)
  • request_timeout (float, optional): Request timeout in seconds (default: 60)
  • google_explicit_cache (bool, optional): Enable Google context caching (default: True)
  • anthropic_prompt_cache (bool, optional): Enable Anthropic automatic prompt caching (default: True)
  • google_inline_citations (bool, optional): Inject [cite: url] markers for Gemini grounding (default: True)
  • google_attach_search_tool (bool, optional): When using Gemini with no custom tools, attach the Google Search tool (default: True). Ignored for non-Google models.
  • google_api_mode (str, optional): "generate_content" (default) or "interactions" for Gemini routing.

ask_interaction(...)

Gemini Interactions API only. Same general parameters as ask where applicable, plus previous_interaction_id for multi-turn sessions. Returns AskResult with interaction_id set.

ask(...)

Generate a response from the LLM.

Parameters:

  • prompt (str): User prompt/question
  • system_instruct (str, optional): System instruction
  • messages (list, optional): Conversation history
  • max_tokens (int, optional): Maximum tokens to generate
  • temperature (float, optional): Sampling temperature (0-2)
  • top_p (float, optional): Nucleus sampling parameter
  • presence_penalty (float, optional): Presence penalty (OpenAI only)
  • reasoning_effort (str, optional): Extended-thinking effort, "low" | "medium" | "high" (Inception also: "instant"). Provider-agnostic — see Extended Thinking.
  • tools_schema (list, optional): Tool/function calling schema (OpenAI, Anthropic, Google, Inception)
  • response_format (dict, optional): Response format specification
  • execute_tool_cb (callable, optional): Tool execution callback (OpenAI, Anthropic, Google, Inception)
  • tool_error_callback (callable, optional): Callback when tool returns ok=False
  • max_steps (int, optional): Maximum tool-calling steps (default: 24)
  • max_effective_tool_steps (int, optional): Maximum effective tool steps (default: 12)
  • force_tool_use (bool, optional): Force at least one tool call when tools provided (default: False)
  • stream (bool, optional): When True, return StreamResult (default: False)
  • attachments (list[Attachment], optional): PDFs or images for the model to read alongside prompt. Provider-agnostic — see Attachments.

Returns: AskResult – Object with .text (str) and .usage (TokenUsage). When stream=True, returns StreamResult – async iterable of stream events (see Streaming above); .usage after completion or aclose().

Raises:

  • ValidationError: If prompt is empty or invalid parameters provided
  • APIError: If the API call fails
  • ConfigurationError: If API keys are missing or client initialization fails

Environment Variables

Loaded from .env (see .env.example) or the process environment:

  • OPENAI_API_KEY: OpenAI API key (required for OpenAI models)
  • ANTHROPIC_API_KEY: Anthropic API key (required for Claude models)
  • GOOGLE_API_KEY: Google API key (required for Google models)
  • INCEPTION_API_KEY: Inception API key (required for Mercury models)

Provider options (google_explicit_cache, google_inline_citations, anthropic_prompt_cache) are passed as constructor params to AskLLM; see docstring.

License

MIT

Contributing

Contributions welcome! Please open an issue or submit a pull request.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

coffee_with_llm-0.8.0.tar.gz (90.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

coffee_with_llm-0.8.0-py3-none-any.whl (75.9 kB view details)

Uploaded Python 3

File details

Details for the file coffee_with_llm-0.8.0.tar.gz.

File metadata

  • Download URL: coffee_with_llm-0.8.0.tar.gz
  • Upload date:
  • Size: 90.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for coffee_with_llm-0.8.0.tar.gz
Algorithm Hash digest
SHA256 f34d182e6c3d8b0061210e04ea67c9dad0082845f058fc313401807633250020
MD5 fc66985c70cec5b26277736343a93d36
BLAKE2b-256 2a576a733874fa583af9dd93e9c191980241b986c9d86b0cea8bb0c93dfebdd2

See more details on using hashes here.

Provenance

The following attestation bundles were made for coffee_with_llm-0.8.0.tar.gz:

Publisher: publish.yml on paveenrajai/coffee-with-llm

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file coffee_with_llm-0.8.0-py3-none-any.whl.

File metadata

File hashes

Hashes for coffee_with_llm-0.8.0-py3-none-any.whl
Algorithm Hash digest
SHA256 f22c7ea123a609a0912473e215cb66e26b02bb28718723b042fe2e915f400ac2
MD5 79e45138e04b6874487de4f4705ebfd2
BLAKE2b-256 0fbbffc061c2ba519029c86246d1310a956043633bde2a82a951c3eff11ae327

See more details on using hashes here.

Provenance

The following attestation bundles were made for coffee_with_llm-0.8.0-py3-none-any.whl:

Publisher: publish.yml on paveenrajai/coffee-with-llm

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page