Skip to main content

Small, clean Python SDK for OpenAI-compatible LLM APIs.

Project description

📦 llm-sdk

Small Python SDK for OpenAI-compatible LLM APIs.

One file, clean API, boring on purpose. Use it with local servers, OpenAI-style endpoints, structured output, tool calls, vision inputs, and reasoning streams.

✨ Features

  • Sync and async clients
  • Streaming and non-streaming responses
  • Configurable retry with per-call override
  • OpenAI Chat Completions support
  • Optional OpenAI Responses API mode
  • Structured output from JSON schema or typed Python classes
  • Tool schema generation from Python callables
  • Vision input normalization from URL, path, base64, or PIL image
  • Thinking/reasoning token parsing
  • Lightweight verbose stats for streams

🚀 Get Started

Install the package directly from PyPI:

pip install llm-sdk-py

If you also need PIL image support:

pip install "llm-sdk-py[pillow]"

Alternatively, since it's designed to be simple, you can still just drop llm_sdk.py directly into your project!

from llm_sdk import LLM

llm = LLM(
    model="qwen3.6-27b",
    base_url="http://localhost:1234",
    api_key="lm-studio",
)

response = llm.response(input="Write a tiny haiku about fast code.")

print(response["answer"])

By default, base_url="http://localhost:1234/v1" and api_key="lm-studio", so local LM Studio-style servers work with very little setup.

All inference methods accept either input="..." for the common single-user-message case or a Chat Completions-style message list:

response = llm.response(messages=[
    {"role": "system", "content": "Be concise."},
    {"role": "user", "content": "Write a tiny haiku about fast code."},
])

📡 Streaming

for event in llm.stream_response(input="Explain adapters in one paragraph."):
    if event["type"] == "answer":
        print(event["content"], end="", flush=True)

Events are small dictionaries:

{"type": "answer", "content": "..."}
{"type": "reasoning", "content": "..."}
{"type": "tool_call", "content": {"id": "...", "name": "...", "arguments": {...}}}
{"type": "verbose", "content": {"tokens": 42, "tokens_per_second": 91.3, "latency": 0.2, "prompt_tokens": 10, "completion_tokens": 32, "total_tokens": 42}}
{"type": "final", "content": {"answer": "..."}}
{"type": "done"}

Use final=True if you also want a final aggregated response event.

⏱️ Async

import asyncio
from llm_sdk import LLM

async def main():
    async with LLM(model="gpt-5.5", api_key="sk-...", base_url="https://api.openai.com/v1", use_responses_api=True) as llm:
        response = await llm.async_response(input="Give me a crisp project name.")
        print(response["answer"])

asyncio.run(main())

🔄 Retry

Set max_retries globally or override it per call.

# global
llm = LLM(model="qwen3.6-27b", max_retries=5)
# per call
llm.response(input="...", max_retries=0)

📐 Structured Output

Pass a JSON schema or a typed class. Classes are converted into OpenAI-compatible JSON schema.

class Verdict:
    sentiment: str
    score: float
    tags: list[str]

result = llm.response(
    input="Review: fast, small, surprisingly nice.",
    output_format=Verdict,
)

print(result["answer"])

🛠️ Tools

Pass Python callables or already-built OpenAI tool definitions. The SDK exposes tool definitions and returns streamed/final tool calls.

It does not execute tools for you. You stay in control.

def search_docs(query: str, limit: int = 5) -> str:
    """Search internal docs."""
    return "..."

response = llm.response(
    input="Find the auth setup notes.",
    tools=[search_docs],
)

print(response.get("tool_calls", []))

👁️ Vision

Image content can be a URL, a local path, base64, or a PIL image.

response = llm.response([
    {
        "role": "user",
        "content": [
            {"type": "text", "text": "What is in this image?"},
            {"type": "image_path", "image_path": "photo.png"},
        ],
    }
])

Supported image forms include:

  • {"type": "image_url", "image_url": "https://..."}
  • {"type": "image_path", "image_path": "local-file.png"}
  • {"type": "image_base64", "image_base64": "..."}
  • {"type": "image_pil", "image_pil": image}

🔌 Responses API

Use use_responses_api=True for endpoints that prefer OpenAI's Responses API shape.

llm = LLM(
    model="gpt-5.5",
    api_key="sk-...",
    base_url="https://api.openai.com/v1",
    use_responses_api=True,
)

🧠 Reasoning Effort

Use reasoning_effort="high" to set models's reasoning effort.

response = llm.response(
    input="...",
    reasoning_effort="high"
)

⚙️ API

  • response(...) and stream_response(...)
  • async_response(...) and async_stream_response(...)
  • input="..." or messages=[...] for all inference methods
  • list_models(fallback=[...]) and async_list_models(fallback=[...])
  • max_retries=3 globally on LLM(...) or per call
  • reasoning_effort="low|medium|high" where supported
  • hide_thinking=False to stream/return reasoning content
  • CustomThinkingToken(...) for custom <think>-style parsing
  • verbose=True for token-ish stream stats
  • with LLM(...) as llm: / async with LLM(...) as llm: for cleanup

💡 Why

Most LLM wrappers either become frameworks or stay too close to raw HTTP. This sits in the middle: enough structure to be pleasant, little enough surface area to understand in one sitting.

📜 License

MIT

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

llm_sdk_py-1.2.0.tar.gz (21.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

llm_sdk_py-1.2.0-py3-none-any.whl (17.3 kB view details)

Uploaded Python 3

File details

Details for the file llm_sdk_py-1.2.0.tar.gz.

File metadata

  • Download URL: llm_sdk_py-1.2.0.tar.gz
  • Upload date:
  • Size: 21.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for llm_sdk_py-1.2.0.tar.gz
Algorithm Hash digest
SHA256 008432aeaba8b7313fd36f9262f9d9f76deb43b8a37a88a17ad6b6a020440822
MD5 0b657da5b5dfe1fcc47a38b38a8475b3
BLAKE2b-256 1f7ba810a0b7b0972c3a06d871c75eb17de1c1dd638444f6862b27a91af7b012

See more details on using hashes here.

Provenance

The following attestation bundles were made for llm_sdk_py-1.2.0.tar.gz:

Publisher: publish.yml on flgaertig/llm-sdk

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file llm_sdk_py-1.2.0-py3-none-any.whl.

File metadata

  • Download URL: llm_sdk_py-1.2.0-py3-none-any.whl
  • Upload date:
  • Size: 17.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for llm_sdk_py-1.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 ba35d6edf70b9d2f7ada44a140a05ae7c2859e2da341745b12566f3c4bc5e64b
MD5 16df1bf06c085e615427aac9648db5fa
BLAKE2b-256 8f1ab243e429b1973e501c25f6bd1db23bfa68ecbe336b516f16d6647729295e

See more details on using hashes here.

Provenance

The following attestation bundles were made for llm_sdk_py-1.2.0-py3-none-any.whl:

Publisher: publish.yml on flgaertig/llm-sdk

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page