langchain-content-normalizer
Normalize the messy content shapes produced by LangChain, MCP tools, Anthropic content blocks, and multimodal chat APIs.
The package has no runtime dependencies. It works by duck typing instead of importing LangChain or MCP classes.
What it solves
LLM agent stacks often receive content as one of many incompatible shapes:
| Source | Example shape | Output |
|---|---|---|
| Classic chat | "plain text" |
"plain text" |
| Anthropic blocks | [{"type": "text", "text": "hi"}] |
"hi" |
| OpenAI Responses text | [{"type": "output_text", "text": "hi"}] |
"hi" |
| Tool calls | [{"type": "tool_use", ...}] |
skipped by default |
| MCP tool results | [{"type": "tool_result", "content": [...]}] |
flattened text |
| MCP objects | objects exposing .text |
extracted text |
| Message wrappers | objects exposing .content |
recursively normalized |
Install
uv add langchain-content-normalizer
Text normalization
from lc_content_normalizer import extract_text_content, normalize_tool_output
content = [
{"type": "text", "text": "Reading logs..."},
{"type": "tool_use", "name": "tail_logs", "input": {"service": "api"}},
]
assert extract_text_content(content) == "Reading logs..."
assert "tail_logs" in extract_text_content(content, skip_tool_use=False)
assert extract_text_content(content, separator="\n") == "Reading logs..."
safe_output = normalize_tool_output(huge_tool_payload, max_chars=50_000, separator="\n")
Vision format routing
from lc_content_normalizer import build_human_message_content, detect_vision_format
vision_format = detect_vision_format("anthropic", "claude-3-5-sonnet")
content = build_human_message_content(
"Explain this alert screenshot",
images=[{"data_url": "data:image/png;base64,...", "mime_type": "image/png"}],
vision_format=vision_format,
)
detect_vision_format() returns:
| Provider/model | Format |
|---|---|
anthropic |
native Anthropic image block with source.base64 |
ollama + known vision model marker (llava, bakllava, moondream, minicpm-v, qwen2-vl, llama3.2-vision, vision) |
OpenAI-compatible image_url block |
ollama text-only model |
none, images are dropped |
| OpenAI-compatible providers | OpenAI-compatible image_url block |
Examples
examples/normalize_mcp_output.pyshows how MCP-style tool results are flattened.examples/build_vision_content.pyshows provider-aware image block generation.
Roadmap
- Add provider-specific adapters as content formats evolve.
- Keep runtime dependencies at zero.
Strict mode
By default, unknown non-empty content is preserved with str(...) so tool output is not silently lost. Use strict mode when unknown shapes should fail fast:
from lc_content_normalizer import UnknownContentBlockError, extract_text_content
try:
extract_text_content([{"type": "custom", "payload": "..."}], strict=True)
except UnknownContentBlockError:
...
Development
uv sync --dev
uv run ruff check .
uv run pytest
uv run python scripts/smoke.py
uv build
License
MIT
Metadata
Release files for langchain-content-normalizer 0.1.8
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| langchain_content_normalizer-0.1.8.tar.gz | 15.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| langchain_content_normalizer-0.1.8-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 22.8 kB
Release files / langchain_content_normalizer-0.1.8.tar.gz
| Download URL | langchain_content_normalizer-0.1.8.tar.gz |
|---|---|
| Size | 15.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
c767be702678a7f4826ab1d7873c0c41899480dda2c758539d8840adb311764a
|
|
BLAKE2b-256 checksum How to use checksums |
f56df5590f93ee8ac4f1c17a44cbaa3b047bd8cac481554afcefe5acbd072f8e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jun 3, 2026.
Transparency logRelease files / langchain_content_normalizer-0.1.8-py3-none-any.whl
| Download URL | langchain_content_normalizer-0.1.8-py3-none-any.whl |
|---|---|
| Size | 7.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
ebddbeeea9e9ce67d1da268632531ae10aea02690e7899a3c9957067aa62f096
|
|
BLAKE2b-256 checksum How to use checksums |
b03e36ac431e7a26e11c1f315bcf6122284a497a31da7585b0f83146b7dbb74a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jun 3, 2026.
Transparency log