polars-llm
Call chat, TypeSafe decision, and embedding models from a Polars DataFrame, one row at a time, using native Polars expressions.
polars-llm registers an .llm namespace on Polars expressions so you can call any LangChain-supported chat model or embedding model on every row of a DataFrame — synchronously or asynchronously — and pipe the responses straight back into your data pipeline.
import polars as pl
import polars_llm # noqa: F401 — registers the `.llm` namespace
(
pl.DataFrame({"user_prompt": ["Summarise polars in one sentence."]})
.with_columns(
pl.col("user_prompt").llm.openai(model="gpt-4o-mini").alias("answer")
)
)
- Repository: https://github.com/diegoglozano/polars-llm
- Documentation: https://diegoglozano.github.io/polars-llm/
- PyPI: https://pypi.org/project/polars-llm/
Why polars-llm?
- Expression-native — works inside
with_columns,select, and any other Polars expression context. No Pythonforloops over rows, no notebook glue. - Sync and async — every provider verb has an
a-prefixed async sibling that fans out concurrently withasyncio.gatherand an optionalmax_concurrencycap. - Provider-agnostic clients — use
chat/achatandembed/aembedwith any LangChain-compatible client, including providers without a dedicated verb. - Per-row prompts and system messages — both the prompt and the system message can be Polars expressions, so you can build them from other columns.
- Structured outputs — pass a Pydantic model as
schema=to get a struct column back, parsed via LangChain'swith_structured_output. - Typed decisions — evaluate TypeSafe
Choice,Score, andNoulquestions together and receive probability-aware nested struct columns. - Embeddings, too —
openai_embedandgemini_embedreturnList[Float64]columns ready for vector search. - Top-K nearest-neighbour join —
df.ann.knn(other, on="vector", k=5)joins one DataFrame of embeddings against another, with a brute-force NumPy default and an optionalusearchHNSW backend for larger corpora. - Powered by LangChain — you get the same retries, batching, and observability primitives the rest of the LangChain ecosystem uses, plumbed straight into a DataFrame.
Common use cases:
- Summarise, classify, translate, or extract structured fields from a column of text.
- Score rows against a custom rubric using an LLM-as-judge.
- Classify, score, and gate rows with TypeSafe System One decisions and confidence values.
- Build embeddings for a corpus directly from a DataFrame, ready to write to a vector database.
- Mix LLM calls with the rest of your pipeline (joins, filters, group-bys) without leaving Polars.
Installation
polars-llm keeps its base install light. Pick the providers you need as extras:
# Just one provider
pip install "polars-llm[openai]"
pip install "polars-llm[anthropic]"
pip install "polars-llm[gemini]"
pip install "polars-llm[typesafe]"
# Top-K nearest-neighbour joins (adds usearch + numpy)
pip install "polars-llm[ann]"
# Or all of them
pip install "polars-llm[all]"
# uv
uv add "polars-llm[all]"
Requires Python 3.10+ and Polars 1.0+. Python 3.9 is no longer supported.
Authentication follows each provider's conventions — set OPENAI_API_KEY, ANTHROPIC_API_KEY, GOOGLE_API_KEY, or TYPESAFE_API_KEY as appropriate.
Quickstart
For a guided setup and more complete recipes, see Getting started and Examples.
1. Chat completion per row
import polars as pl
import polars_llm # noqa: F401
df = (
pl.DataFrame({"user_prompt": [
"What is the capital of Spain?",
"What is the capital of France?",
]})
.with_columns(
pl.col("user_prompt").llm.openai(model="gpt-4o-mini").alias("answer")
)
)
2. System prompt — literal or per-row
# Same system prompt for every row
pl.col("user_prompt").llm.anthropic(
model="claude-sonnet-4-6",
system="Answer in fewer than 10 words.",
)
# Per-row system prompt from another column
pl.col("user_prompt").llm.gemini(
model="gemini-2.5-pro",
system=pl.col("system_prompt"),
)
3. Async for throughput
The a-prefixed verbs run concurrently across the batch, capped at max_concurrency:
df.with_columns(
pl.col("user_prompt").llm.aopenai(
model="gpt-4o-mini",
max_concurrency=20,
).alias("answer")
)
4. Structured output with Pydantic
from pydantic import BaseModel
class Sentiment(BaseModel):
label: str # "positive" | "neutral" | "negative"
confidence: float
df.with_columns(
pl.col("review").llm.openai(
model="gpt-4o-mini",
schema=Sentiment,
).alias("sentiment")
).unnest("sentiment")
5. TypeSafe structured decisions
The source expression becomes the TypeSafe state. Ask several independent questions in one request and get a nested Polars struct containing typed answers, probabilities, and confidence:
from typesafe_sdk import Choice, Noul, Score
questions = {
"department": Choice(
instructions="Which team should handle this?",
criteria={
"returns": "Exchanges, wrong or damaged items",
"shipping": "Delivery status, delays, lost packages",
"billing": "Charges, invoices, payment problems",
},
),
"urgency": Score(
instructions="How urgent is this ticket?",
criteria=["Can wait", "Needs attention today", "Customer is blocked"],
),
"needs_reply": Noul(instructions="Does this ticket require a reply?"),
}
decisions = df.with_columns(
pl.col("ticket").llm.typesafe(questions=questions).alias("decision")
).unnest("decision")
# Async row evaluation with bounded concurrency:
pl.col("ticket").llm.atypesafe(questions=questions, max_concurrency=20)
Question dictionaries in the TypeSafe API request shape are accepted too, and the state expression may be a string, struct, or other JSON-compatible Polars value. TypeSafe defaults to jev-latest; pass model= to override it.
6. Embeddings
df.with_columns(
pl.col("text").llm.openai_embed(
model="text-embedding-3-small",
).alias("vector")
)
The generic verbs accept any compatible LangChain chat or embedding client,
so a provider does not need a dedicated polars-llm integration:
from langchain_ollama import ChatOllama, OllamaEmbeddings
chat = ChatOllama(model="llama3.2")
embeddings = OllamaEmbeddings(model="nomic-embed-text")
df.with_columns(
pl.col("text").llm.chat(client=chat).alias("answer"),
pl.col("text").llm.embed(client=embeddings).alias("vector"),
)
7. Top-K nearest-neighbour join
Once you have an embedding column on each side, df.ann.knn returns the k closest rows from other for every row of df:
import polars as pl
import polars_llm # noqa: F401 — registers the `.ann` namespace
queries = pl.DataFrame({
"q_id": ["q1", "q2"],
"vector": [[0.9, 0.1], [0.0, 1.0]],
})
docs = pl.DataFrame({
"doc_id": ["a", "b", "c"],
"vector": [[1.0, 0.0], [0.0, 1.0], [1.0, 1.0]],
})
queries.ann.knn(docs, on="vector", k=2)
# shape: (4, 5)
# ┌──────┬───────────┬────────┬──────┬──────────┐
# │ q_id ┆ vector ┆ doc_id ┆ rank ┆ score │
# ╞══════╪═══════════╪════════╪══════╪══════════╡
# │ q1 ┆ [0.9,0.1] ┆ a ┆ 0 ┆ 0.005… │
# │ q1 ┆ [0.9,0.1] ┆ c ┆ 1 ┆ 0.071… │
# │ q2 ┆ [0.0,1.0] ┆ b ┆ 0 ┆ 0.0 │
# │ q2 ┆ [0.0,1.0] ┆ c ┆ 1 ┆ 0.293… │
# └──────┴───────────┴────────┴──────┴──────────┘
backend="auto" (default) uses brute-force NumPy under ~50k rows and switches to usearch HNSW for larger corpora when the [ann] extra is installed. Force one with backend="brute" or backend="usearch". Pass flat=False to get a neighbors: List[Struct] column instead of a flat join. Lower score = closer match.
8. Retries, caching, metadata
pl.col("user_prompt").llm.aanthropic(
model="claude-sonnet-4-6",
retries=3,
backoff=0.5,
max_concurrency=10,
cache=True, # dedupe identical prompts within a batch
with_metadata=True, # struct {content, elapsed_ms, error}
)
API reference
All methods live under the .llm namespace on any Polars expression that resolves to a string column.
Chat verbs
| Method | Provider | Mode |
|---|---|---|
chat / achat |
Any compatible client | sync / async |
openai / aopenai |
OpenAI | sync / async |
anthropic / aanthropic |
Anthropic | sync / async |
gemini / agemini |
Google Gemini | sync / async |
Embedding verbs
| Method | Provider | Mode |
|---|---|---|
embed / aembed |
Any compatible client | sync / async |
openai_embed / aopenai_embed |
OpenAI Embeddings | sync / async |
gemini_embed / agemini_embed |
Google Gemini | sync / async |
Anthropic does not currently offer a first-party embeddings API.
Decision verbs
| Method | Provider | Mode |
|---|---|---|
typesafe / atypesafe |
TypeSafe System One | sync / async |
questions= is a mapping of answer names to TypeSafe Choice, Score, or Noul objects (or equivalent raw dictionaries). The default result is a nested Struct, with one field per question. with_metadata=True wraps it as Struct{answers, model, input_tokens, output_tokens, elapsed_ms, error}.
DataFrame .ann namespace
df.ann.knn(other, **kwargs) — top-K nearest-neighbour join between two DataFrames of vectors.
| Argument | Default | Notes |
|---|---|---|
on / left_on/right_on |
— | Vector column name(s). Use on= when both sides share a name, otherwise both *_on. |
k |
5 |
Number of neighbours per row. Clamped to len(other). |
metric |
"cosine" |
One of "cosine", "ip", "l2" (squared L2). Lower score = closer match. |
backend |
"auto" |
"auto" switches to usearch above ~50k right rows when installed; otherwise "brute". |
flat |
True |
True → len(df) * k rows. False → one row per query with a List[Struct] neighbors col. |
suffix |
"_right" |
Right-side column collision suffix (flat output only). |
rank_name/score_name |
"rank"/"score" |
Names of the added rank and distance columns. |
**backend_kwargs |
— | Forwarded to usearch.index.Index (connectivity, expansion_add, expansion_search, dtype, …). |
The vector columns must be List[Float32/64] or Array[Float32/64, dim], and dimensions must match between the two DataFrames.
Common arguments
All verbs are keyword-only and accept:
model(provider verbs, str) — model name forwarded to LangChain (e.g."gpt-4o-mini","claude-sonnet-4-6","gemini-2.5-pro").system(chat only) — literal string orpl.Exprfor a per-row system prompt.schema(chat only) — a Pydantic model class. Returns a struct column with the schema fields, viawith_structured_output.client— a pre-configured LangChain chat or embeddings instance. It is required by the generic verbs and skips the in-tree constructor when passed to a provider verb.retries(int, default 0) — retry on any exception raised by the provider call.backoff(float, default 0.0) — exponential backoff base (seconds).max_concurrency(async only, int) — cap on in-flight requests viaasyncio.Semaphore.cache(bool, default False) — memoise identical inputs within a batch.with_metadata(bool, default False) — return a struct column with timing and error metadata instead of just the content / vector.on_error("null" | "raise", default "null") — whenwith_metadata=False, what to do on errors."null"replaces failures withNoneand emits a warning;"raise"re-raises immediately.**model_kwargs— any additional keyword arguments forwarded to the underlying LangChain class (e.g.temperature=,max_tokens=,timeout=).
Return types
| Mode | Default dtype | With with_metadata=True |
|---|---|---|
Chat (no schema) |
Utf8 |
Struct{content: Utf8, elapsed_ms: Float64, error: Utf8} |
Chat (with schema) |
Struct{...} matching the Pydantic model |
Same struct; content JSON-serialised under content |
| Embeddings | List[Float64] |
Struct{vector: List[Float64], dim: Int64, elapsed_ms: Float64, error: Utf8} |
Tips and patterns
- Build prompts from columns with
pl.format("Translate to {}: {}", pl.col("language"), pl.col("text")). - Bring your own client with
chat(client=...)orembed(client=...); use theira-prefixed counterparts for concurrent execution. - Watch the warning — when a request fails and is silently nulled, polars-llm emits a
UserWarningso you don't ship a column of nulls by accident. Passwith_metadata=Trueto inspect per-row errors instead. - Combine with lazy frames — every verb is an expression, so it composes inside
LazyFrame.with_columns(...).
Contributing
Contributions are welcome — see CONTRIBUTING.md. Please open an issue before starting on larger changes.
License
MIT © Diego Garcia Lozano
Inspired by and patterned after polars-api.
Metadata
Release files for polars-llm 0.4.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| polars_llm-0.4.0.tar.gz | 34.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| polars_llm-0.4.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 57.8 kB
Release files / polars_llm-0.4.0.tar.gz
| Download URL | polars_llm-0.4.0.tar.gz |
|---|---|
| Size | 34.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
c010d4381b112b92c8cecaed2d37018d7621191d495d9aa96f2b50a6981f1713
|
|
BLAKE2b-256 checksum How to use checksums |
a4ac9adf940176e9193c3cd2825d90d9cd21430c2384f4e4feb9938580487121
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.14
|
Release files / polars_llm-0.4.0-py3-none-any.whl
| Download URL | polars_llm-0.4.0-py3-none-any.whl |
|---|---|
| Size | 23.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c40ac4713102eb397d1c4e7accf1729e2abee7e63e65d27ebc05cb63502c57ea
|
|
BLAKE2b-256 checksum How to use checksums |
b48e00a30919c745a7159ced2d5a18e03df2927e3da3339cc10a1c5ce9c22e10
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.14
|