langchain-urlpipe
Let your LangChain agents read any web page, JavaScript sites included, and load pages into your RAG pipeline as clean Markdown.
URLpipe renders each page in real Chrome before it answers, so a single-page app or a docs site built with JavaScript comes back with its content, where a plain HTTP fetch gets an empty shell. This package gives you four LangChain tools, a toolkit and a document loader, built on the official urlpipe Python client.
Install
pip install langchain-urlpipe
Python 3.10 or later.
Set up your API key
Create a project at urlpipe.dev and copy its API key. The Free plan gives you 1,000 credits a month, no card needed.
export URLPIPE_API_KEY=your_project_key
Every tool, the toolkit and the loader read URLPIPE_API_KEY, or take api_key="…" directly.
Tools
| Tool | Name the model sees | What it returns | Credits |
|---|---|---|---|
UrlpipeReadPage |
urlpipe_read_page |
the page's main content as Markdown | 1 |
UrlpipePageMetadata |
urlpipe_page_metadata |
title, description, language, main image, author, publication date, feed | 5 |
UrlpipeConsoleErrors |
urlpipe_console_errors |
the errors, warnings and uncaught exceptions the page logs | 1 |
UrlpipeLighthouse |
urlpipe_lighthouse |
Lighthouse scores (0–100) and key metrics, on mobile or desktop |
2 |
Each takes a url and an optional max_age: how fresh a stored copy must be, in seconds or as a duration such as "1 hour". Asking for the same page again within 7 days is served from storage and costs nothing; max_age=0 always loads it fresh.
from langchain_urlpipe import UrlpipeReadPage, UrlpipeLighthouse
UrlpipeReadPage().invoke({"url": "https://example.com"})
# '# Example Domain\n\nThis domain is for use in …'
UrlpipeLighthouse().invoke({"url": "https://example.com", "device": "desktop"})
# {'url': 'https://example.com', 'device': 'desktop',
# 'scores': {'performance': 100, 'accessibility': 88, 'best-practices': 100, 'seo': 90},
# 'metrics': {'first-contentful-paint': {'value': '0.3 s', 'score': 100}, …}}
In an agent
from langchain.agents import create_agent
from langchain_urlpipe import UrlpipeReadPage, UrlpipePageMetadata
agent = create_agent(
"anthropic:claude-sonnet-5",
tools=[UrlpipeReadPage(), UrlpipePageMetadata()],
system_prompt="You answer questions about web pages. Read a page before you describe it.",
)
result = agent.invoke(
{"messages": [{"role": "user", "content": "What plans does https://urlpipe.dev/pricing list?"}]}
)
print(result["messages"][-1].content)
Or bind the tools to any chat model that supports tool calling:
from langchain.chat_models import init_chat_model
from langchain_urlpipe import UrlpipeReadPage
read_page = UrlpipeReadPage()
model = init_chat_model("openai:gpt-4.1-mini").bind_tools([read_page])
message = model.invoke("Summarise https://example.com in one sentence.")
for call in message.tool_calls:
print(read_page.invoke(call)) # a ToolMessage with the page's Markdown
When URLpipe cannot load a page, or your credits run out, the tool returns the API's explanation to the model as the tool result (a ToolMessage with status="error"), so the agent can try another page or tell you why. Pass handle_tool_error=False to have a ToolException raised instead.
Toolkit
UrlpipeToolkit hands an agent all four tools at once, sharing one client:
from langchain.agents import create_agent
from langchain_urlpipe import UrlpipeToolkit
tools = UrlpipeToolkit().get_tools()
agent = create_agent("anthropic:claude-sonnet-5", tools=tools)
agent.invoke(
{"messages": [{"role": "user", "content": "Why is https://example.com slow on mobile?"}]}
)
Document loader
UrlpipeLoader turns a list of URLs into Documents, one per page, with the page as Markdown:
from langchain_urlpipe import UrlpipeLoader
loader = UrlpipeLoader(
["https://urlpipe.dev/docs", "https://urlpipe.dev/pricing"],
page_options={"block_cookie_banners": True},
)
docs = loader.load()
docs[0].metadata
# {'source': 'https://urlpipe.dev/docs', 'operation': 'markdown', 'token': '…', 'cache': 'miss'}
operation="markdown"(the default, 1 credit a page),"html"for the HTML after JavaScript ran (1 credit), or"summarize"for an AI summary in Markdown (17 credits).max_ageandpage_options(wait_for_selector,delay,block_ads,block_cookie_banners,remove_selectors) go with every page.continue_on_failure=Truelogs and skips a page that fails instead of raising itsurlpipe.UrlpipeError.lazy_load()yields pages one at a time as they arrive;alazy_load()andaload()load up tomax_concurrencypages at once (3 by default) and keep the order of your URLs.
For RAG
Markdown keeps the page's headings, so you can split on them and keep each chunk's section as metadata:
from langchain_text_splitters import MarkdownHeaderTextSplitter, RecursiveCharacterTextSplitter
from langchain_urlpipe import UrlpipeLoader
docs = UrlpipeLoader(["https://urlpipe.dev/docs"]).load()
by_heading = MarkdownHeaderTextSplitter(
headers_to_split_on=[("#", "h1"), ("##", "h2"), ("###", "h3")]
)
sized = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=100)
chunks = []
for doc in docs:
for section in by_heading.split_text(doc.page_content):
section.metadata.update(doc.metadata) # keep source, token, cache
chunks.extend(sized.split_documents([section]))
# add `chunks` to any vector store
langchain-text-splitters is a separate install: pip install langchain-text-splitters.
Configuration
The tools, the toolkit and the loader take the same connection arguments:
UrlpipeReadPage(
api_key="…", # default: URLPIPE_API_KEY
client_kwargs={"timeout": 120, "max_retries": 3}, # passed to urlpipe.Client
)
client_kwargs accepts base_url, timeout, max_retries and wait_timeout, as documented for the urlpipe client. To reuse a client you already have, or send requests through your own httpx client, pass client=urlpipe.Client(…) and async_client=urlpipe.AsyncClient(…). Without async_client, each async call opens its own connection, so the tools work from any event loop.
Links
- The LangChain integration on urlpipe.dev: https://urlpipe.dev/integrations/langchain
- API docs: https://urlpipe.dev/docs
- MCP server, for using URLpipe from AI assistants: https://github.com/URLpipe/mcp
License
MIT © Aliat Partner S.L.
Release files for langchain-urlpipe 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| langchain_urlpipe-0.1.0.tar.gz | 15.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| langchain_urlpipe-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 29.0 kB
Release files / langchain_urlpipe-0.1.0.tar.gz
| Download URL | langchain_urlpipe-0.1.0.tar.gz |
|---|---|
| Size | 15.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
760b51632fdde4562ca4384b6df15e2f5401ab01dd2fb936083563e22f0b6549
|
|
BLAKE2b-256 checksum How to use checksums |
368505bf4f846a7f0d29f8bda21507d41d95d417d451e7f7f60f2888098114d8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency logRelease files / langchain_urlpipe-0.1.0-py3-none-any.whl
| Download URL | langchain_urlpipe-0.1.0-py3-none-any.whl |
|---|---|
| Size | 13.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0562dae3ec52e3a607d9e88e22f37e0f8fbc437f16363112529e6104f7b39fd7
|
|
BLAKE2b-256 checksum How to use checksums |
290df257617fb7f0debda0b1743efdeac01ca9e013054482cfb7f5157f22161a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency log