Async Azure OpenAI toolkit — structured output, agent pipelines, vision, token tracking
Project description
llm_service
Async Python toolkit for Azure OpenAI (Foundry). One module the team imports everywhere — handles config, model detection, concurrency, structured output, agent pipelines, vision, token tracking.
Installation
pip install httpx pyyaml pydantic
Clone the repo and import directly, or copy llm_service/ into your project.
Quick start
# config.yaml
api_key: ${AZURE_OPENAI_KEY}
endpoint: https://your-resource.openai.azure.com
model_name: gpt-4.1
import asyncio
from llm_service import LLMConfig, LLMClient
async def main():
cfg = LLMConfig.from_yaml("config.yaml")
async with LLMClient(cfg) as llm:
answer = await llm.chat(
"Summarize this document.",
system="You extract key facts."
)
print(answer)
print(llm.usage.summary())
asyncio.run(main())
In a Jupyter notebook asyncio.run() is not needed — just await directly:
cfg = LLMConfig.from_yaml("config.yaml")
async with LLMClient(cfg) as llm:
answer = await llm.chat("Hello!")
Features
Batch (concurrent)
results = await llm.batch(
["What is ETL?", "What is RAG?", "What is a vector DB?"],
system="Answer in 1-2 sentences.",
)
Semaphore (default 8) controls how many requests fly in parallel.
Structured output (Pydantic)
from pydantic import BaseModel, Field
class Invoice(BaseModel):
number: str = Field(description="Invoice number")
total: float = Field(description="Gross amount")
currency: str = Field(description="Currency code")
invoice = await llm.structured(
f"Extract data from this invoice:\n\n{document}",
Invoice,
)
print(invoice.number, invoice.total, invoice.currency)
Two modes:
strict=True(default) — Azure constrains output to the schema server-sidestrict=False— JSON mode + post-validation with Pydantic
JSON output (no schema)
data = await llm.chat_json("Return top 3 cities in Poland as JSON with 'cities' key")
# → {"cities": [{"name": "Warsaw", ...}, ...]}
Pipeline (agent chains)
Multiple agents working on the same input, with dependencies:
from llm_service import Pipeline, Step
pipe = Pipeline(
# Layer 1 — run in parallel (no depends_on)
Step(name="analyst", system="You are a business analyst.", output=AnalystView),
Step(name="lawyer", system="You are a corporate lawyer.", output=LegalView),
# Layer 2 — waits for both
Step(
name="critic",
system="You are a critical reviewer.",
prompt="Review:\n\n{analyst}\n\n{lawyer}\n\nOriginal:\n{input}",
output=Verdict,
depends_on=["analyst", "lawyer"],
),
)
async with LLMClient(cfg) as llm:
result = await pipe.run(llm, input="Proposal to acquire startup ABC...")
print(result.output("critic"))
print(result.trace())
Pipeline completed in 4.2s
[OK] analyst: 1.8s, 450 tokens, $0.0018
[OK] lawyer: 2.1s, 520 tokens, $0.0022
[OK] critic: 1.9s, 380 tokens, $0.0014
Total: 1,350 tokens, $0.0054 USD
Pipelines are reusable — define once, run on many documents (including via asyncio.gather).
Vision / multimodal
# From file
answer = await llm.chat("Describe this image.", images=["photo.png"])
# Structured OCR from scan
invoice = await llm.structured(
"Extract invoice data from this scan.",
Invoice,
images=["scan.jpg"],
image_detail="high",
)
# Compare two images
result = await llm.structured(
"Compare these two images.",
ComparisonResult,
images=["v1.png", "v2.png"],
)
Accepts file paths (auto base64), URLs (passthrough), or raw bytes.
Token tracking and cost estimation
async with LLMClient(cfg) as llm:
await llm.batch(prompts)
print(llm.usage.summary())
Requests: 50
Tokens: 12,340 (prompt: 8,200, completion: 4,140)
reasoning: 1,200 (included in completion)
Cost: $0.0495 USD
Built-in pricing for gpt-4.1, gpt-4o, o-series, gpt-5.x families.
Model auto-detection
The client automatically adapts to the model:
| Model | temperature | system role | token param | reasoning_effort |
|---|---|---|---|---|
| gpt-4.1 | supported | system |
max_completion_tokens |
stripped |
| gpt-5.4-mini | stripped | system |
max_completion_tokens |
medium default |
| o3, o4-mini | stripped | developer |
max_completion_tokens |
medium default |
Unsupported params are silently dropped — no need for model-specific code.
Error handling
from llm_service import LLMError
try:
result = await llm.chat("...")
except LLMError as e:
print(e)
LLMError: Rate limit exceeded
HTTP 429
Azure code: RateLimitExceeded
Model: gpt-4.1
Hint: Rate limit — za duzo requestow. Zmniejsz concurrency lub dodaj retry.
Retries: 3
attempt 1: HTTP 429 after 1.2s
attempt 2: HTTP 429 after 0.8s
attempt 3: HTTP 429 after 0.9s
Retries (408, 429, 5xx, timeouts) with exponential backoff + Retry-After header support.
Configuration
YAML (recommended)
api_key: ${AZURE_OPENAI_KEY}
endpoint: https://your-resource.openai.azure.com
model_name: gpt-4.1
# All optional:
# temperature: 0.7
# max_tokens: 4096
# reasoning_effort: medium # for reasoning models: low|medium|high
# concurrency: 8 # max parallel requests
# timeout: 120 # seconds per request
# retries: 3
Supports ${ENV_VAR} and ${ENV_VAR:default} substitution.
Code
cfg = LLMConfig(
api_key=os.environ["AZURE_OPENAI_KEY"],
endpoint="https://your-resource.openai.azure.com",
model_name="gpt-4.1",
)
Only api_key, endpoint, model_name are required. Everything else has sensible defaults.
Project structure
llm_service/
__init__.py — public API exports
config.py — YAML config loader with env-var substitution
models.py — model capability registry (reasoning vs standard)
client.py — async httpx client, semaphore, retry, error handling
structured.py — Pydantic ↔ JSON schema, response parsing
pipeline.py — agent chains with DAG execution
usage.py — token counting, cost estimation
vision.py — image encoding for multimodal requests
config.example.yaml — config template
examples.py — usage examples for .py scripts
examples.ipynb — usage examples for Jupyter notebooks
Examples
See examples.py (scripts) and examples.ipynb (notebooks) for complete working examples covering all features.
Dependencies
httpx— async HTTP clientpyyaml— YAML config loadingpydantic— structured output validation
No dependency on the openai Python SDK. Direct HTTP to Azure OpenAI REST API.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file neuravec_llm-0.2.0.tar.gz.
File metadata
- Download URL: neuravec_llm-0.2.0.tar.gz
- Upload date:
- Size: 26.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e8203b3b3f7a32f4570717c4e997055007022f16819943afbaac22ebab3278f5
|
|
| MD5 |
76be26381a9ac05283027185b6d41d03
|
|
| BLAKE2b-256 |
6b6d2dd20de8ac1197bcacedd4de1946e1dc3ed1da9dccfdffba8825d2b6a995
|
File details
Details for the file neuravec_llm-0.2.0-py3-none-any.whl.
File metadata
- Download URL: neuravec_llm-0.2.0-py3-none-any.whl
- Upload date:
- Size: 26.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
29038f390a2c221284e9d1e2b066e5c9c261415e3fec39f8a6570f73f9d84160
|
|
| MD5 |
77f156d1b6db64a76ec0e3de3dad3649
|
|
| BLAKE2b-256 |
f7c788259e41033604f15f0f379f5892ab9c952bc10880c9c114e25f6a6ec9bb
|