Skip to main content

uselemma-tracing

HTTP tracing SDK for AI agents. The primary API sends trace payloads directly to Lemma over HTTP.

Installation

pip install uselemma-tracing

Quick Start

from uselemma_tracing import Lemma

lemma = Lemma(release="1.8.3")  # or set LEMMA_RELEASE

def run(trace):
    docs = search_docs(user_message)
    trace.record_tool(
        name="search_docs",
        input={"query": user_message},
        output=docs,
        tool_parameters={"query": "string"},
    )

    response = call_model(user_message, docs)
    trace.record_generation(
        name="draft-reply",
        input=response.messages,
        output=response.text,
        model="gpt-4o",
        llm_input_messages=[{"role": "user", "content": user_message}],
        llm_invocation_parameters={"temperature": 0.2},
    )

    return response.text

answer = lemma.trace(
    "support-agent",
    run,
    input=user_message,
    thread_id=conversation_id,
    user_id=user.id,
)

lemma.trace() measures the trace from callback start to completion. Use async_trace() for async callbacks.

Pass release (or set LEMMA_RELEASE) to stamp the running app version on every ingest payload. An explicit constructor value wins. Empty or invalid values are omitted.

Live Spans

def run(trace):
    span = trace.start_span(name="retrieve-context", input=query)
    try:
        docs = retrieve(query)
        span.end(output={"count": len(docs)})
        return docs
    except Exception as error:
        span.end(status="ERROR", error=error)
        raise

Live handles know their start time when created and their end time when .end() is called, so you usually do not pass duration_ms. Pass duration_ms only when replaying historical work or overriding the measured duration with a value from another timer.

For one-off records where you already measured the work, pass duration_ms on the record call:

trace.record_generation(
    name="answer",
    output=text,
    model="gpt-4o",
    duration_ms=measured_model_ms,
)

User-facing messaging tools

When a tool delivers the agent's response to the end user, pass the exact display text as user_facing_message. Lemma renders that text as an assistant message while preserving the complete tool input and output in the span detail:

tool_input = {
    "message": "Your order arrives Friday.",
    "send_as_voice_note": False,
    "should_terminate": True,
}

trace.record_tool(
    name="send_whatsapp",
    input=tool_input,
    output={"delivered": True},
    user_facing_message=tool_input["message"],
)

The tool's own schema can call the value message, text, body, or anything else. Lemma never guesses which input field the user saw. Omit user_facing_message for internal tools; their payload and rendering are unchanged.

The same handle pattern is available for tool calls and generations:

tool = trace.start_tool(name="search_docs", input={"query": query})
docs = search_docs(query)
tool.end(output=docs)

generation = trace.start_generation(name="answer", input=messages)
response = call_model(messages)
generation.end(output=response.text)

Sending a Trace You Built Yourself

trace() assumes the client owns the trace lifecycle within a single process. When the producer lives elsewhere — a cross-process buffer, a queue worker, a batch backfill — build a TraceContext yourself and deliver it with ingest():

from uselemma_tracing import Lemma, TraceContext

lemma = Lemma()

context = TraceContext(
    id=turn_id,  # stable id for this execution (use for retries)
    name=prompt,
    input=prompt,
    thread_id=conversation_id,
)
context.record_tool(name="search_docs", input=query, output=docs, duration_ms=25)
context.record_generation(name="answer", model="gpt-4o", output=final_answer)
context.output(final_answer)

lemma.ingest(context, started_at=started_at)

ingest() POSTs one payload. Deliver one complete trace when the execution (agent turn) finishes: root input/output, thread/user, and all child spans in one call. This is required — patching a trace over time is not currently supported.

ingest() is not an incremental merge API: omitted root fields do not preserve prior values, and after Lemma processes the trace once, a later re-delivery does not re-run issue extraction (occasional late child spans may still append to the tree for display). Retries of the same complete payload are safe — already-stored span IDs are skipped — so a failed send can be retried as-is. It raises on a non-2xx response and never mutates the trace's status.

Automatic delivery (trace / async_trace) fails open: a Lemma ingest 4xx/5xx or network error is logged in debug mode and dropped so it cannot fail the caller's application. LangChain and OpenAI Agents flush through this path. Use ingest() when you need a failed send to raise so you can retry.

One turn across processes

thread_id correlates turns of a conversation. It is not how you glue a host process and an E2B-style sandbox into one turn. The host mints a versioned context token, the child records a serializable journal without a Lemma API key, and the host applies the journal then ingest()s once.

import json
from uselemma_tracing import Lemma, attach_turn

lemma = Lemma()
turn = lemma.start_turn(
    "agent-turn",
    input=user_message,
    thread_id=conversation_id,
)
sandbox = turn.start_span(name="e2b-sandbox")
token = json.dumps(turn.export(parent_span_id=sandbox.id))
# pass token to the child on the existing channel, then:
turn.apply(child_journal)
sandbox.end()
turn.end(output=answer)  # strict ingest

In the child, do not construct Lemma and do not call /traces/ingest:

import json
import os
from uselemma_tracing import attach_turn

local = attach_turn(os.environ["LEMMA_TURN"])
local.record_tool(name="search_docs", input=query, output=docs)
print(json.dumps(local.records()))

The journal uses the same camelCase schema as the TypeScript SDK so a TS host can apply a Python child's dump (and the reverse). assemble_turn(token, journal) builds a TraceContext when the coordinator already has the dump; then call ingest() once. Re-applying the same journal is idempotent (stable span ids). If the sandbox dies before a clean dump, end the host sandbox span as ERROR; tools that started and never ended are left incomplete.

OpenAI Agents SDK

Install the OpenAI Agents extra and register the Lemma processor:

pip install "uselemma-tracing[openai-agents]" openai-agents
from agents import Agent, Runner
from uselemma_tracing import instrument_openai_agents

instrument_openai_agents()

agent = Agent(
    name="support-agent",
    instructions="Answer customer questions clearly and concisely.",
)

async def call_agent(user_message: str):
    result = await Runner.run(agent, user_message)
    return result.final_output

The processor creates one Lemma trace for each OpenAI Agents trace with root current-turn input, final output or terminal error, promoted thread_id / user_id, and wall-clock bounds from child spans. Generation/response spans become Lemma generations, function spans become Lemma tool spans, and parent IDs are preserved so tools stay nested under the generation or agent span that called them.

Pass OpenAI Agents group_id for thread_id and metadata user_id / userId for user_id. Call force_flush() / shutdown() to finalize open traces once.

Enable debug mode to validate live span shape while developing:

from uselemma_tracing import enable_debug_mode

enable_debug_mode()

Prompts, tool inputs, outputs, generated text, and error messages are always recorded — Lemma cannot show what a run consumed, produced, or why it failed without them.

LangChain and LangGraph

Install the optional integration dependency and pass langchain() as a callback handler. Each root run owns one Lemma trace with current-turn input, final output or root error, promoted thread_id / user_id, typed nested generations/tools/spans, and real wall-clock bounds. Call flush() / shutdown() to finalize open traces.

pip install "uselemma-tracing[langchain]" langchain-openai
from langchain_openai import ChatOpenAI
from uselemma_tracing import langchain

handler = langchain(
    agent_name="support-agent",
    thread_id_key="conversation_id",
    user_id_key="user_id",
)
model = ChatOpenAI(model="gpt-4o", callbacks=[handler])
response = model.invoke(
    user_message,
    config={"metadata": {"conversation_id": thread_id, "user_id": user_id}},
)
handler.flush()

langgraph() is the same LangChain callback adapter with a LangGraph default trace name (langgraph-agent):

pip install "uselemma-tracing[langgraph]"
from uselemma_tracing import langgraph

result = graph.invoke(
    {"input": user_message},
    {"callbacks": [langgraph(agent_name="support-graph")]},
)

Prompts, tool inputs, outputs, generated text, and error messages are always recorded.

Supported Contract Fields

Use native SDK keyword arguments for OpenInference-style fields:

  • LLM: llm_model_name, llm_provider, llm_system, llm_invocation_parameters, llm_input_messages, llm_output_messages, llm_tools, usage / input_tokens / output_tokens / cache and reasoning kwargs (omit when the provider did not supply them — never invent zeros), and prompt template fields
  • provenance: every span includes lemma.sdk.language and lemma.sdk.integration (manual by default; framework integrations override)
  • tools: tool_description, tool_parameters, user_facing_message
  • embeddings and rerankers: embedding_model_name, embedding_invocation_parameters, embedding_embeddings, reranker_model_name, reranker_input_documents, reranker_output_documents

Use attributes for raw attributes that do not yet have a native SDK keyword.

Configuration

Option Environment variable Default
api_key LEMMA_API_KEY Required
project_id LEMMA_PROJECT_ID Required
base_url none https://api.uselemma.ai

The SDK sends to {base_url}/traces/ingest.

You can pass configuration directly to the constructor instead of using environment variables:

lemma = Lemma(
    api_key="sk_...",
    project_id="proj_...",
    base_url="https://api.uselemma.ai",
)

Debug Mode

Debug mode logs trace starts, span starts, span completions, send attempts, and send results as they happen:

from uselemma_tracing import enable_debug_mode

enable_debug_mode()

You can also set LEMMA_DEBUG=1 (true also works). Use this when validating that spans are created in the expected order and the SDK is sending to the intended URL.

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

uselemma_tracing-7.11.0.tar.gz (34.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

uselemma_tracing-7.11.0-py3-none-any.whl (39.5 kB view details)

Uploaded Python 3

File details

Details for the file uselemma_tracing-7.11.0.tar.gz.

File metadata

  • Download URL: uselemma_tracing-7.11.0.tar.gz
  • Upload date:
  • Size: 34.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.12.9 {"installer":{"name":"uv","version":"0.12.9","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for uselemma_tracing-7.11.0.tar.gz
Algorithm Hash digest
SHA256 823ac099945a2097a91c722d3e5e00d2ac41f9af8f40fff1e15ef2011166413c
MD5 a3eca567d66bb9e2a8649694bbb2bad4
BLAKE2b-256 c82c6330e20045c9ea26aa2ae5bc10f4989afeaf1142277211086aafd620cac7

See more details on using hashes here.

File details

Details for the file uselemma_tracing-7.11.0-py3-none-any.whl.

File metadata

  • Download URL: uselemma_tracing-7.11.0-py3-none-any.whl
  • Upload date:
  • Size: 39.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.12.9 {"installer":{"name":"uv","version":"0.12.9","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for uselemma_tracing-7.11.0-py3-none-any.whl
Algorithm Hash digest
SHA256 aafa6ae395a4bccf5c70259c8f83ad85bf4ddbc99d8c6b9cc7d1748b3a63dde2
MD5 385a14e1c3132104bf7ade0570bb3e76
BLAKE2b-256 b1b4be6c2cf6e2f6b1d285f00e761a534b8b98c3e4f6192e33e86153d191ced9

See more details on using hashes here.

Release history Release notifications | RSS feed

7.11.2

2 files

7.11.1

2 files

This release

7.11.0 This release

2 files

7.10.3

2 files

7.10.2

2 files

7.10.1

2 files

7.10.0

2 files

7.8.0

2 files

7.7.2

2 files

7.7.1

2 files

7.7.0

2 files

7.6.0

2 files

7.5.0

2 files

7.4.2

2 files

7.4.1

2 files

7.4.0

2 files

7.3.0

2 files

7.2.0

2 files

7.1.0

2 files

7.0.0

2 files

6.0.0

2 files

5.0.0

2 files

4.2.0

2 files

4.1.0

2 files

4.0.1

2 files

4.0.0

2 files

3.0.6

2 files

3.0.5

2 files

3.0.4

2 files

3.0.3

2 files

3.0.2

2 files

3.0.1

2 files

3.0.0

2 files

2.17.0

2 files

2.16.0

2 files

2.14.1

2 files

2.14.0

2 files

2.13.0

2 files

2.12.0

2 files

2.11.0

2 files

2.10.0

2 files

2.9.0

2 files

2.8.0

2 files

2.7.0

2 files

2.6.0

2 files

2.5.0

2 files

2.4.0

2 files

2.3.0

2 files

2.2.0

2 files

2.1.0

2 files

2.0.0

2 files

1.1.0

2 files

1.0.2

2 files

1.0.1

2 files

1.0.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page