Skip to main content

SilkLoom Core

A lightweight pandas accessor for batch LLM extraction.

DataFrame rows → Jinja2 prompt render → OpenAI-compatible API → repaired JSON → result DataFrame

One call — df.llm.extract(template) — concurrently sends every row to an LLM, parses the JSON response, and returns a DataFrame you can join back to the original.

Table of Contents

Install

pip install silkloom-core

# optional: progress bar
pip install silkloom-core[progress]

Importing silkloom_core registers the df.llm accessor on every DataFrame.

Quick Start

import pandas as pd
import silkloom_core

silkloom_core.configure(api_key="...", base_url="https://api.openai.com/v1")

df = pd.DataFrame({
    "title": ["A clear experiment", "A weak evaluation"],
    "abstract": ["Reliable and reproducible.", "Too small to conclude much."],
})

df = df.llm.extract(
    "Title: {{ title }}\nAbstract: {{ abstract }}\nReturn JSON with keys label and summary.",
    model="gpt-4o-mini",
    max_workers=8,
    json_mode=True,
)
# df now has: title, abstract, label, summary

extract() returns the original DataFrame with extracted columns appended. Pass join=False if you only want the extracted columns.

Configuration

SilkLoom supports three layers of client configuration, from broadest to narrowest.

Global Configuration

Configure once at the start of a script — every DataFrame can call extract() directly:

silkloom_core.configure(
    api_key="...",
    base_url="https://api.openai.com/v1",
)

# Or pass a pre-built client:
silkloom_core.configure(client=OpenAI(api_key="...", base_url="..."))

Per-DataFrame Setup

Override the global default for a specific DataFrame:

df.llm.setup(
    api_key="...",
    base_url="...",
    cache_path="special_cache.db",
)

# Chain directly into extract:
df.llm.setup(client=client).extract("...", model="gpt-4o-mini")

Per-Call Client

Override the client for a single extract() call — useful for mixing providers:

from openai import OpenAI

openai_client = OpenAI(api_key="...", base_url="https://api.openai.com/v1")
zhipu_client = OpenAI(api_key="...", base_url="https://open.bigmodel.cn/api/paas/v4")

silkloom_core.configure(client=openai_client)

# Same DataFrame, different providers:
df.llm.extract("...", model="gpt-4o")                          # → OpenAI
df.llm.extract("...", model="glm-4-flash", client=zhipu_client) # → Zhipu

Priority Chain

extract(client=...)  >  df.llm.setup(client=...)  >  silkloom_core.configure(client=...)  >  error

If no client is configured, extract() raises RuntimeError with a clear message.

Note: The SQLite cache file is created lazily — only on the first extract() call, never as a side effect of configure() or setup().

Extraction

Prompt Templates

Prompts use Jinja2 with StrictUndefined — typos in column names raise immediately instead of silently producing empty strings. Literal JSON braces ({ }) are safe; only {{ }} is treated as a template expression.

out = df.llm.extract(
    'Classify {{ text }} and return JSON like {"label": "positive", "score": 0.9}',
    model="gpt-4o-mini",
    temperature=0.1,
    max_workers=4,
    max_retries=2,
)

Result Columns

The returned DataFrame has the same index as the input. Column semantics:

Condition Column(s)
Model returns a JSON object Each key becomes a column
Model returns a non-object JSON value _llm_raw
JSON parse fails _llm_error + _llm_raw
API call fails after all retries _llm_error

Malformed JSON is repaired with json_repair before parsing.

Cache and Audit

Every API call is recorded in a SQLite database. Successful responses (ok = 1) are reused as cache hits on subsequent runs with the same request. Failed requests are also stored for debugging but are retried next time.

df.llm.setup(cache_path="cache/llm.sqlite").extract(...)

The cache table schema:

Column Type Description
cache_key TEXT PK SHA-256 of the full request
ok INTEGER 1 = success (cacheable), 0 = failure
model TEXT Model name used
messages_json TEXT Rendered messages array
params_json TEXT Request params (excluding model/messages)
request_json TEXT Full request payload
response TEXT Raw model response text
parsed_json TEXT Parsed result dict
error TEXT Error message (NULL on success)
attempts INTEGER Number of attempts made
created_at TEXT Row creation timestamp
updated_at TEXT Last update timestamp

To start fresh: delete the SQLite file or use a different cache_path.

Images

Pass image_column for local file paths, HTTP(S) URLs, or data:image/... URLs. Local files are auto-encoded as base64 data URLs with MIME detection. A single cell can hold a list of images for multi-image input.

out = df.llm.extract(
    "Extract fields from this receipt and return JSON.",
    image_column="receipt_path",
    model="gpt-4o-mini",
)

Rows with missing image values (NaN, None) fall back to text-only prompts.

Progress and Cancel

Progress bar — tqdm is used when verbose=True (default). If tqdm isn't installed, it degrades silently.

Callback — for UI integration:

def progress(done, total):
    print(f"{done}/{total}")

out = df.llm.extract("Analyze {{ text }}", progress_callback=progress)

Cancel — from another thread:

df.llm.cancel()

Queued work is cancelled where possible. Running rows stop before their next retry. Already-completed results are preserved in the returned DataFrame.

API Reference

silkloom_core.configure(...)

Parameter Type Default Description
api_key str | None None API key for OpenAI client
base_url str | None None API base URL
cache_path str | Path ".llm_cache.db" SQLite cache file path
client Any | None None Pre-built client (overrides api_key/base_url)
**client_options Extra kwargs for OpenAI()

df.llm.setup(...)

Same parameters as configure(). Returns self for chaining.

df.llm.extract(prompt_template, *, ...)

Parameter Type Default Description
prompt_template str Jinja2 template (required)
client Any | None None Per-call client override
image_column str | None None Column with image paths/URLs
system_prompt str | None "Please output valid JSON only." System message; None to omit
model str "gpt-4o-mini" Model name
max_workers int 4 Concurrent API call threads
json_mode bool False Set response_format={"type":"json_object"}
max_retries int 2 Retries on API error (exponential backoff)
join bool True If True, return original DataFrame with extracted columns appended; if False, return only extracted columns
progress_callback Callable[[int, int], None] | None None Called with (completed, total)
verbose bool True Show tqdm progress bar
**request_options Extra kwargs for chat.completions.create()

df.llm.cancel()

No parameters. Signals cancellation to all in-flight work.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

silkloom_core-7.1.0.tar.gz (16.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

silkloom_core-7.1.0-py3-none-any.whl (11.6 kB view details)

Uploaded Python 3

File details

Details for the file silkloom_core-7.1.0.tar.gz.

File metadata

  • Download URL: silkloom_core-7.1.0.tar.gz
  • Upload date:
  • Size: 16.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for silkloom_core-7.1.0.tar.gz
Algorithm Hash digest
SHA256 16e89047a51e723cf636872b6f74187ef58255850e04e41b9e5efd6ad6ceeee6
MD5 bdafcaacbac0e773c021d5a6094f01cf
BLAKE2b-256 cf4f2be4a76c957f65c5969c032ee64d3bfa57f2c8e2c90308ec28711b74ed33

See more details on using hashes here.

Provenance

The following attestation bundles were made for silkloom_core-7.1.0.tar.gz:

Publisher: publish.yml on LeLiu-GeoAI/silkloom-core

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file silkloom_core-7.1.0-py3-none-any.whl.

File metadata

  • Download URL: silkloom_core-7.1.0-py3-none-any.whl
  • Upload date:
  • Size: 11.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for silkloom_core-7.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 b5eea8ab86d3babb7350afdac523f07cdb6656b2099ee3e2878ccef632f6b7c2
MD5 26276700796a07b1c37537cd482bcc69
BLAKE2b-256 64a80131ec6e0721785932b40bd8b240fcc4ed4040236b1675ab7d6accda1fe8

See more details on using hashes here.

Provenance

The following attestation bundles were made for silkloom_core-7.1.0-py3-none-any.whl:

Publisher: publish.yml on LeLiu-GeoAI/silkloom-core

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page