Skip to main content

Language Model Development Kit

What it offers:

  • Simplest interface to call different Language Model APIs
  • Minimal dependencies: HTTP requests only, no third party packages
  • Streaming
  • Thinking / reasoning efforts
  • Comfy structured outputs via Pydantic models
  • Parallel completions
  • Unified HTTP error handling
  • Easy location config (for providers with multiple datacenters like AWS Bedrock, GCP Vertex and Azure)
  • Model fallbacks
  • Bring Your Own Key (for each provider)
  • Optional Telemetry following OpenTelemetry GenAI Semantic Conventions
  • In-process observation hook (observe()) to capture request/response pairs from wrapped code
  • Decision models (non-generative encoders) like Jev or Laya via decide(): label probabilities for typed questions

What it does NOT offer:

  • Tools / function calling / MCP
  • Agents (but you can build your own on top of this! See markov-agent)
  • Multimodality (only text-in, text-out)
  • Shady under-the-hood prompt modification (e.g. to force structured output)
  • API gateways

If you are looking for a more constrained but out-of-the-box agent interface, I'd recommend pydantic-ai or haystack-ai. If you are looking to keep granular control but extend on tools or multimodality, I'd recommend litellm or leveraging the OpenAI-compatible endpoints that providers normally set up. The closest to the "less intrusive path to unified LM calling" idea in lmdk is chatlas, but their main public interface is stateful (versus the full stateless design of lmdk). If you want a unified a token for all providers and are willing to give away telemetry data, check Gateways like Opper or openrouter.

Installation

uv add lmdk

Optional OpenTelemetry support:

uv add 'lmdk[telemetry]'

Usage

from lmdk import complete

model = "mistral:mistral-small-2603"
# supports locations as in "vertex:gemini-2.5-flash@europe-west4"
# or "bedrock:eu.anthropic.claude-opus-5@eu-central-1" (region inferred from the
# geo prefix when omitted)
Single prompt
response = complete(model=model, prompt="Tell me a joke")
Multi-turn conversation
messages = [
    UserMessage("My name is Alice."),
    AssistantMessage("Nice to meet you, Alice!"),
    UserMessage("What is my name?"),
]
response = complete(model=model, prompt=messages)
System prompt and generation kwargs
response = complete(
    model=model,
    prompt="Hi!",
    system_instruction="Talk like a pirate",
    generation_kwargs={"temperature": 0.9, "max_tokens": 10}
)
Streaming
token_iter = complete(model=model, prompt="Count from 1 to 5.", stream=True)
Model fallbacks
response = complete(model=["mistral:nonexistent-model", model], prompt="Hi")
# first request will raise NotFoundError bc model does not exist, second will work
Local models

Most local model servers (llama.cpp, vLLM, Ollama, LM Studio, etc.) expose an OpenAI-compatible /v1/chat/completions API. You can direct requests to any local or self-hosted endpoint using the local: provider prefix and providing the host/port via the @location suffix:

# Format: local:<model-id>@<host>[:<port>]
response = complete(
    model="local:Qwen3.6-27B-BF16@192.168.10.51:4000",
    prompt="Hello!"
)
  • If the location has no scheme, http:// is assumed.
  • If the server requires authentication, set the LOCAL_API_KEY environment variable, which is automatically sent as a Bearer token.
Structured output
class Ingredient(BaseModel):
    name: str
    quantity: int
    unit: str = ""

class Recipe(BaseModel):
    ingredients: list[Ingredient]

response = complete(model=model, prompt="How do I make cheescake?", output_schema=Recipe)
# response.parsed will have a Recipe instance
Reasoning / thinking

lmdk offers 4 thinking configurations for the time-test compute of recent LLMs. We map them to the configs available on the provider side. See the documentation on each provider for more info.

# "none" (default) | "low" | "medium" | "high"
response = complete(model=model, prompt="Solve this carefully...", thinking_effort="high")

# Works alongside structured output where the provider supports both:
response = complete(
    model=model,
    prompt="Plan a 3-day trip to Lisbon.",
    output_schema=Trip,
    thinking_effort="medium",
)

The usage of the thinking can be seen in the respective CompletionResponse.thinking and CompletionResponse.thinking_tokens fields.

Parallel calls
from lmdk import complete_batch

batch = complete_batch(model=model, prompt_list=["Greet in english", "Saluda en espanyol."])
# `batch` is a CompletionBatch. Iterate it to handle each outcome:
for result in batch:
    if isinstance(result, Exception):
        ...  # this prompt failed
    else:
        ...  # CompletionResponse

# Aggregates over successful responses:
batch.input_tokens, batch.output_tokens, batch.latency
batch.responses  # successes only
batch.errors     # exceptions only
Template Rendering
from lmdk import render_template

# Render a template string with variables
result = render_template(
    template="Hello, {{ name }}!",
    name="World"
)
# Output: "Hello, World!"

# Render a template from a jinja file
result = render_template(
    path="path/to/template.jinja2",
    name="World"
)
Observing wrapped code
from lmdk import observe

with observe() as obs:
    answer = my_function_that_calls_complete()

for record in obs.records:
    record.request    # CompletionRequest sent to the LM
    record.response   # CompletionResponse returned

Useful for tests, evals, and debug tooling where the wrapped function only returns its own result but you also want to inspect the underlying LM calls. Streaming completions are not recorded.

Decisions

Decision models are encoders: instead of generating text, they return a probability for each label of each question, in a single forward pass. Every question is evaluated independently against the same state, in one request.

from lmdk import Question, decide

response = decide(
    model="typesafe:jev-latest",  # or "laya:english@localhost:8000" for a self-hosted laya-serve
    state={"message": "I was billed twice for March, refund it or I cancel."},
    questions={
        "department": Question(
            text="Which team should handle `message`?",
            labels={"billing": "invoices, payments, refunds", "technical": "bugs, outages"},
        ),
        "urgency": Question(
            text="How urgent is `message`?",
            labels={"low": "can wait", "medium": "this week", "high": "today"},
            is_ordered=True,
        ),
        "churn": Question(
            text="Does the customer threaten to leave?",
            labels={"yes": "mentions cancelling or leaving", "no": "no such threat"},
        ),
    },
)
response.probabilities  # {"department": {"billing": 0.97, "technical": 0.03}, "urgency": {...}, ...}
  • A Question is its text plus labels, each with a description. Set is_ordered=True when the labels form a scale (in insertion order). Labels that are a yes/no (or true/false) pair use the provider's dedicated yes/no head.
  • state is a string or a dict; a dict keeps its structure, so questions can reference fields by path (`message`).
  • Model fallbacks work as in complete. Batching, telemetry and observe() cover completions only for now.

Telemetry

Telemetry is off by default and adds no required dependencies to the default install. To enable OpenTelemetry-based spans and metrics, install the optional extra and set LMDK_TELEMETRY:

uv add 'lmdk[telemetry]'
export LMDK_TELEMETRY=metadata  # spans/metrics without prompt text
# export LMDK_TELEMETRY=content  # also records prompt, system-instruction, and response text

We follows the experimental Gen AI semconv v1.41.0. We only instrument non-streaming responses for now.

lmdk only emits telemetry through the OpenTelemetry SDK. Your application owns exporter, processor, reader, collector endpoint, i.e.: you decide how and where to send the emitted traces.

Below are some minimal exporter setups. Call them once at process start before invoking complete / complete_batch.

Console (debugging)

Prints spans to stdout. Useful to verify instrumentation locally without any backend.

from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor, ConsoleSpanExporter


def configure_console_traces() -> None:
    provider = TracerProvider()
    provider.add_span_processor(BatchSpanProcessor(ConsoleSpanExporter()))
    trace.set_tracer_provider(provider)
Pydantic Logfire

Logfire installs itself as the global TracerProvider, so spans emitted by lmdk are forwarded automatically. Requires uv add logfire and a LOGFIRE_TOKEN.

import os
import logfire


def configure_logfire_traces() -> None:
    logfire.configure(
        token=os.environ["LOGFIRE_TOKEN"],
        service_name="my-app",
        # lmdk already controls prompt/response redaction via LMDK_TELEMETRY;
        # don't let Logfire second-guess scrubbing of content.
        scrubbing=False,
        send_to_logfire=True,
    )
Grafana (OTLP / Tempo)

Ship spans over OTLP to Grafana Cloud (or a self-hosted Tempo + OTel Collector). Requires uv add opentelemetry-exporter-otlp.

import os

from opentelemetry import trace
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor


def configure_grafana_traces() -> None:
    # For Grafana Cloud OTLP, set:
    #   OTEL_EXPORTER_OTLP_ENDPOINT=https://otlp-gateway-<region>.grafana.net/otlp
    #   OTEL_EXPORTER_OTLP_HEADERS=Authorization=Basic%20<base64(instanceID:token)>
    exporter = OTLPSpanExporter(
        endpoint=os.environ["OTEL_EXPORTER_OTLP_ENDPOINT"] + "/v1/traces",
    )
    provider = TracerProvider(resource=Resource.create({"service.name": "my-app"}))
    provider.add_span_processor(BatchSpanProcessor(exporter))
    trace.set_tracer_provider(provider)

Development

Structure

src/lmdk/
├── completion.py   # Entry points: complete, complete_batch
├── decision.py     # Entry point: decide
├── datatypes.py    # Common message, question and response schemas
├── provider.py     # Base Provider class and registry
├── providers/      # Concrete implementations (Mistral, Vertex, etc.)
├── errors.py       # Unified HTTP and API error handling
└── utils.py        # Shared helper functions

Tooling

We use just for development tasks. Use:

  • just install: Sync environment from the lockfile.
  • just format: Lints and formats with ruff.
  • just check-types: Static analysis with ty.
  • just check-complexity: Cyclomatic complexity checks with complexipy.
  • just test: Runs pytest with 90% coverage threshold.

See justfile for a complete list of dev commands.

Contribute

  1. Hooks: Install pre-commit hooks via just install-hooks. PRs will fail CI if linting/formatting is not applied.
  2. Issues: Open an issue first using the default template.
  3. PRs: Link your PR to the relevant issue using the PR template.

You can use just validate <model> (runs example.py) to verify which features run properly and which do not for a new provider / model. Not all of them have to pass to open a PR: some providers do not even support native structured output. Do at least the normal non-structured, non-streamed completion. The rest can raise NotImplementedError.

License

MIT

This inference package has been used to build:

Made with mold template

Metadata

Release files for lmdk 2.14.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for lmdk 2.14.0
File Size Uploaded
lmdk-2.14.0.tar.gz 37.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for lmdk 2.14.0
File Interpreter ABI Platform
lmdk-2.14.0-py3-none-any.whl Python 3 none any Details

Total release size: 87.9 kB

Release files / lmdk-2.14.0.tar.gz

Download URL lmdk-2.14.0.tar.gz
Size 37.6 kB
Tags Source
SHA-256 checksum
How to use checksums
1fec160bb2a38251c57d9ab7a6b6f6366ace6581d441b37c96685e553ca79cb5
BLAKE2b-256 checksum
How to use checksums
6a14feef5b1972ff98f50f1785c5ff125961557200ef7cc9f74b935eab6adf2f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / lmdk-2.14.0-py3-none-any.whl

Download URL lmdk-2.14.0-py3-none-any.whl
Size 50.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
cf106835c2b9d4a8a7dcd834705d38486c1149282e2750652d814db540318afa
BLAKE2b-256 checksum
How to use checksums
0cdee18f99e4b0e74027f2be5c1b3c70b7255123af6be883580df74f2470f8d3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

2.14.0 This release

2 release files

2.13.1

2 release files

2.13.0

2 release files

2.12.0

2 release files

2.10.3

2 release files

2.10.2

2 release files

2.10.1

2 release files

2.10.0

2 release files

2.9.1

2 release files

2.9.0

2 release files

2.8.1

2 release files

2.8.0

2 release files

2.7.0

2 release files

2.6.0

2 release files

2.5.1

2 release files

2.5.0

2 release files

2.4.0

2 release files

2.3.0

2 release files

2.2.0

2 release files

2.1.1

2 release files

2.1.0

2 release files

2.0.0

2 release files

1.8.0

2 release files

1.7.0

2 release files

1.6.1

2 release files

1.6.0

2 release files

1.5.0

2 release files

1.4.0

2 release files

1.3.1

2 release files

1.3.0

2 release files

1.2.1

2 release files

1.2.0

2 release files

1.1.2

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page