Skip to main content

langchain-urlpipe

Let your LangChain agents read any web page, JavaScript sites included, and load pages into your RAG pipeline as clean Markdown.

URLpipe renders each page in real Chrome before it answers, so a single-page app or a docs site built with JavaScript comes back with its content, where a plain HTTP fetch gets an empty shell. This package gives you four LangChain tools, a toolkit and a document loader, built on the official urlpipe Python client.

Install

pip install langchain-urlpipe

Python 3.10 or later.

Set up your API key

Create a project at urlpipe.dev and copy its API key. The Free plan gives you 1,000 credits a month, no card needed.

export URLPIPE_API_KEY=your_project_key

Every tool, the toolkit and the loader read URLPIPE_API_KEY, or take api_key="…" directly.

Tools

Tool Name the model sees What it returns Credits
UrlpipeReadPage urlpipe_read_page the page's main content as Markdown 1
UrlpipePageMetadata urlpipe_page_metadata title, description, language, main image, author, publication date, feed 5
UrlpipeConsoleErrors urlpipe_console_errors the errors, warnings and uncaught exceptions the page logs 1
UrlpipeLighthouse urlpipe_lighthouse Lighthouse scores (0–100) and key metrics, on mobile or desktop 2

Each takes a url and an optional max_age: how fresh a stored copy must be, in seconds or as a duration such as "1 hour". Asking for the same page again within 7 days is served from storage and costs nothing; max_age=0 always loads it fresh.

from langchain_urlpipe import UrlpipeReadPage, UrlpipeLighthouse

UrlpipeReadPage().invoke({"url": "https://example.com"})
# '# Example Domain\n\nThis domain is for use in …'

UrlpipeLighthouse().invoke({"url": "https://example.com", "device": "desktop"})
# {'url': 'https://example.com', 'device': 'desktop',
#  'scores': {'performance': 100, 'accessibility': 88, 'best-practices': 100, 'seo': 90},
#  'metrics': {'first-contentful-paint': {'value': '0.3 s', 'score': 100}, …}}

In an agent

from langchain.agents import create_agent
from langchain_urlpipe import UrlpipeReadPage, UrlpipePageMetadata

agent = create_agent(
    "anthropic:claude-sonnet-5",
    tools=[UrlpipeReadPage(), UrlpipePageMetadata()],
    system_prompt="You answer questions about web pages. Read a page before you describe it.",
)

result = agent.invoke(
    {"messages": [{"role": "user", "content": "What plans does https://urlpipe.dev/pricing list?"}]}
)
print(result["messages"][-1].content)

Or bind the tools to any chat model that supports tool calling:

from langchain.chat_models import init_chat_model
from langchain_urlpipe import UrlpipeReadPage

read_page = UrlpipeReadPage()
model = init_chat_model("openai:gpt-4.1-mini").bind_tools([read_page])

message = model.invoke("Summarise https://example.com in one sentence.")
for call in message.tool_calls:
    print(read_page.invoke(call))  # a ToolMessage with the page's Markdown

When URLpipe cannot load a page, or your credits run out, the tool returns the API's explanation to the model as the tool result (a ToolMessage with status="error"), so the agent can try another page or tell you why. Pass handle_tool_error=False to have a ToolException raised instead.

Toolkit

UrlpipeToolkit hands an agent all four tools at once, sharing one client:

from langchain.agents import create_agent
from langchain_urlpipe import UrlpipeToolkit

tools = UrlpipeToolkit().get_tools()
agent = create_agent("anthropic:claude-sonnet-5", tools=tools)

agent.invoke(
    {"messages": [{"role": "user", "content": "Why is https://example.com slow on mobile?"}]}
)

Document loader

UrlpipeLoader turns a list of URLs into Documents, one per page, with the page as Markdown:

from langchain_urlpipe import UrlpipeLoader

loader = UrlpipeLoader(
    ["https://urlpipe.dev/docs", "https://urlpipe.dev/pricing"],
    page_options={"block_cookie_banners": True},
)
docs = loader.load()

docs[0].metadata
# {'source': 'https://urlpipe.dev/docs', 'operation': 'markdown', 'token': '…', 'cache': 'miss'}
  • operation="markdown" (the default, 1 credit a page), "html" for the HTML after JavaScript ran (1 credit), or "summarize" for an AI summary in Markdown (17 credits).
  • max_age and page_options (wait_for_selector, delay, block_ads, block_cookie_banners, remove_selectors) go with every page.
  • continue_on_failure=True logs and skips a page that fails instead of raising its urlpipe.UrlpipeError.
  • lazy_load() yields pages one at a time as they arrive; alazy_load() and aload() load up to max_concurrency pages at once (3 by default) and keep the order of your URLs.

For RAG

Markdown keeps the page's headings, so you can split on them and keep each chunk's section as metadata:

from langchain_text_splitters import MarkdownHeaderTextSplitter, RecursiveCharacterTextSplitter
from langchain_urlpipe import UrlpipeLoader

docs = UrlpipeLoader(["https://urlpipe.dev/docs"]).load()

by_heading = MarkdownHeaderTextSplitter(
    headers_to_split_on=[("#", "h1"), ("##", "h2"), ("###", "h3")]
)
sized = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=100)

chunks = []
for doc in docs:
    for section in by_heading.split_text(doc.page_content):
        section.metadata.update(doc.metadata)  # keep source, token, cache
        chunks.extend(sized.split_documents([section]))

# add `chunks` to any vector store

langchain-text-splitters is a separate install: pip install langchain-text-splitters.

Configuration

The tools, the toolkit and the loader take the same connection arguments:

UrlpipeReadPage(
    api_key="…",                                     # default: URLPIPE_API_KEY
    client_kwargs={"timeout": 120, "max_retries": 3},  # passed to urlpipe.Client
)

client_kwargs accepts base_url, timeout, max_retries and wait_timeout, as documented for the urlpipe client. To reuse a client you already have, or send requests through your own httpx client, pass client=urlpipe.Client(…) and async_client=urlpipe.AsyncClient(…). Without async_client, each async call opens its own connection, so the tools work from any event loop.

License

MIT © Aliat Partner S.L.

Release files for langchain-urlpipe 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for langchain-urlpipe 0.1.1
File Size Uploaded
langchain_urlpipe-0.1.1.tar.gz 15.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for langchain-urlpipe 0.1.1
File Interpreter ABI Platform
langchain_urlpipe-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 29.6 kB

Release files / langchain_urlpipe-0.1.1.tar.gz

Download URL langchain_urlpipe-0.1.1.tar.gz
Size 15.7 kB
Tags Source
SHA-256 checksum
How to use checksums
1daa5a8d48ed2a47937092fbe6b6d66d66737b63fa26fc3538444f7957721bdd
BLAKE2b-256 checksum
How to use checksums
3ba10d84a10d31ac677ed9215feef3615d25fe9a80dfed9b7938eae844bfb84b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / langchain_urlpipe-0.1.1-py3-none-any.whl

Download URL langchain_urlpipe-0.1.1-py3-none-any.whl
Size 13.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5bc9d8ffd559ea464094e1c51264befb97b3bed33837f96d2b4af27079a366ce
BLAKE2b-256 checksum
How to use checksums
25f5ef05d7d34839396ca50cd7aa5ffce53ac2e4fbd8defbcdca522722906dbf
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page