Skip to main content

langchain-content-normalizer

CI PyPI License: MIT Python

Normalize the messy content shapes produced by LangChain, MCP tools, Anthropic content blocks, and multimodal chat APIs.

The package has no runtime dependencies. It works by duck typing instead of importing LangChain or MCP classes.

What it solves

LLM agent stacks often receive content as one of many incompatible shapes:

Source Example shape Output
Classic chat "plain text" "plain text"
Anthropic blocks [{"type": "text", "text": "hi"}] "hi"
OpenAI Responses text [{"type": "output_text", "text": "hi"}] "hi"
Tool calls [{"type": "tool_use", ...}] skipped by default
MCP tool results [{"type": "tool_result", "content": [...]}] flattened text
MCP objects objects exposing .text extracted text
Message wrappers objects exposing .content recursively normalized

Install

uv add langchain-content-normalizer

Text normalization

from lc_content_normalizer import extract_text_content, normalize_tool_output

content = [
    {"type": "text", "text": "Reading logs..."},
    {"type": "tool_use", "name": "tail_logs", "input": {"service": "api"}},
]

assert extract_text_content(content) == "Reading logs..."
assert "tail_logs" in extract_text_content(content, skip_tool_use=False)
assert extract_text_content(content, separator="\n") == "Reading logs..."

safe_output = normalize_tool_output(huge_tool_payload, max_chars=50_000, separator="\n")

Vision format routing

from lc_content_normalizer import build_human_message_content, detect_vision_format

vision_format = detect_vision_format("anthropic", "claude-3-5-sonnet")
content = build_human_message_content(
    "Explain this alert screenshot",
    images=[{"data_url": "data:image/png;base64,...", "mime_type": "image/png"}],
    vision_format=vision_format,
)

detect_vision_format() returns:

Provider/model Format
anthropic native Anthropic image block with source.base64
ollama + known vision model marker (llava, bakllava, moondream, minicpm-v, qwen2-vl, llama3.2-vision, vision) OpenAI-compatible image_url block
ollama text-only model none, images are dropped
OpenAI-compatible providers OpenAI-compatible image_url block

Examples

  • examples/normalize_mcp_output.py shows how MCP-style tool results are flattened.
  • examples/build_vision_content.py shows provider-aware image block generation.

Roadmap

  • Add provider-specific adapters as content formats evolve.
  • Keep runtime dependencies at zero.

Strict mode

By default, unknown non-empty content is preserved with str(...) so tool output is not silently lost. Use strict mode when unknown shapes should fail fast:

from lc_content_normalizer import UnknownContentBlockError, extract_text_content

try:
    extract_text_content([{"type": "custom", "payload": "..."}], strict=True)
except UnknownContentBlockError:
    ...

Development

uv sync --dev
uv run ruff check .
uv run pytest
uv run python scripts/smoke.py
uv build

License

MIT

Metadata

Release files for langchain-content-normalizer 0.1.8

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for langchain-content-normalizer 0.1.8
File Size Uploaded
langchain_content_normalizer-0.1.8.tar.gz 15.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for langchain-content-normalizer 0.1.8
File Interpreter ABI Platform
langchain_content_normalizer-0.1.8-py3-none-any.whl Python 3 none any Details

Total release size: 22.8 kB

Release files / langchain_content_normalizer-0.1.8.tar.gz

Download URL langchain_content_normalizer-0.1.8.tar.gz
Size 15.7 kB
Tags Source
SHA-256 checksum
How to use checksums
c767be702678a7f4826ab1d7873c0c41899480dda2c758539d8840adb311764a
BLAKE2b-256 checksum
How to use checksums
f56df5590f93ee8ac4f1c17a44cbaa3b047bd8cac481554afcefe5acbd072f8e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jun 3, 2026.

Transparency log

Release files / langchain_content_normalizer-0.1.8-py3-none-any.whl

Download URL langchain_content_normalizer-0.1.8-py3-none-any.whl
Size 7.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
ebddbeeea9e9ce67d1da268632531ae10aea02690e7899a3c9957067aa62f096
BLAKE2b-256 checksum
How to use checksums
b03e36ac431e7a26e11c1f315bcf6122284a497a31da7585b0f83146b7dbb74a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jun 3, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.8 This release

2 release files

0.1.7

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page