langchain-ferrolabsai
LangChain integration for Ferro Labs AI Gateway — route LangChain chat, streaming, tool-calling, structured-output, and embedding workloads across 30 LLM providers through a single OpenAI-compatible endpoint, with automatic fallback, load balancing, budgets, and observability.
Compatibility: langchain-ferrolabsai 0.2.x ↔ ferrolabsai ≥ 0.3.0; requires ai-gateway ≥ v1.4.0 (contract-tested against v1.4.5); langchain-core ≥ 0.3 (tested on 1.x).
Install
pip install langchain-ferrolabsai
Quick start
Chat
from langchain_ferrolabsai import FerroChatModel
from langchain_core.messages import HumanMessage
llm = FerroChatModel(
model="gpt-4o",
base_url="http://localhost:8080", # any Ferro Labs AI Gateway instance
api_key="fgw_...",
)
response = llm.invoke([HumanMessage(content="Hello, world")])
print(response.content)
print(response.response_metadata["provider"]) # which provider answered
print(response.response_metadata["trace_id"]) # gateway X-Request-ID
print(response.response_metadata.get("gateway_overhead_ms")) # gateway's own overhead
The gateway-derived fields on response_metadata are model, id,
trace_id, provider, and gateway_overhead_ms, with absent values stripped
(LangChain adds its own, such as finish_reason).
trace_id is the join key for client.admin.logs.list() and for the
gateway's observability exporters (LangSmith, Langfuse, Phoenix, …).
Swap providers without changing the model class — Ferro auto-routes by model name:
claude = FerroChatModel(model="claude-3-5-sonnet-20241022", base_url="...", api_key="...")
gemini = FerroChatModel(model="gemini-2.5-flash", base_url="...", api_key="...")
Streaming
for chunk in llm.stream([HumanMessage(content="Tell me a story")]):
print(chunk.content, end="", flush=True)
Async
response = await llm.ainvoke([HumanMessage(content="Hello")])
async for chunk in llm.astream([HumanMessage(content="Tell me a story")]):
print(chunk.content, end="", flush=True)
Tool calling / LangGraph agents
from langchain_core.tools import tool
@tool
def add(a: int, b: int) -> int:
"""Add two integers."""
return a + b
agent_llm = llm.bind_tools([add])
response = agent_llm.invoke([HumanMessage(content="What is 4 + 7?")])
print(response.tool_calls)
Structured output
Uses the OpenAI-style response_format={"type": "json_schema", ...} path, so
it works with every provider the gateway can translate it for (see
client.capabilities() on the core SDK).
from pydantic import BaseModel
class Answer(BaseModel):
city: str
population: int
structured = llm.with_structured_output(Answer)
print(structured.invoke("Largest city in France?")) # Answer(city='Paris', population=...)
# include_raw=True → {"raw": AIMessage, "parsed": Answer | None, "parsing_error": ...}
Embeddings
from langchain_ferrolabsai import FerroEmbeddings
embed = FerroEmbeddings(model="text-embedding-3-small", base_url="...", api_key="...")
vectors = embed.embed_documents(["hello", "world"])
query_vec = embed.embed_query("hello")
# async
vectors = await embed.aembed_documents(["hello", "world"])
Legacy LLM interface
from langchain_ferrolabsai import FerroLLM
llm = FerroLLM(model="gpt-4o", base_url="...", api_key="...")
print(llm.invoke("Write a haiku about gateways"))
Why use this instead of ChatOpenAI(base_url=...)?
ChatOpenAI pointed at a Ferro Labs gateway works as a drop-in. This package adds:
provider,trace_id, andgateway_overhead_msonresponse_metadata(andtrace_idon the first streamed chunk) — read from the gateway's real response headers, no guessing.- Typed gateway errors from the core SDK:
FerroBudgetExceededError(402),FerroPermissionError(403),FerroRateLimitError.retry_after, withRetry-After-aware retries on 429/5xx. - Async parity (
ainvoke,astream,aembed_*) overAsyncFerroClient.
Status & roadmap
0.2.0 adds the async surface and with_structured_output() on top of
ferrolabsai 0.3. See CHANGELOG.md.
Related
ferrolabsai— the core Python SDK this package wraps.- Ferro Labs AI Gateway — the open-source gateway server.
ai-gateway-cookbook— runnable recipes (start withpython/02-langgraph-multi-provider-agent).- Documentation
License
Apache-2.0
Metadata
Release files for langchain-ferrolabsai 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| langchain_ferrolabsai-0.2.0.tar.gz | 16.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| langchain_ferrolabsai-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 29.5 kB
Release files / langchain_ferrolabsai-0.2.0.tar.gz
| Download URL | langchain_ferrolabsai-0.2.0.tar.gz |
|---|---|
| Size | 16.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8e3a7ad4ac01fd246dbef6b7c1fbd5b1512f28ba915b6362e53175e26aa34cd7
|
|
BLAKE2b-256 checksum How to use checksums |
5a8ae0e7cf019c6f7650d570250122bfd3f66e9dd7ce5745ef45f94dc4106ff4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 29, 2026.
Transparency logRelease files / langchain_ferrolabsai-0.2.0-py3-none-any.whl
| Download URL | langchain_ferrolabsai-0.2.0-py3-none-any.whl |
|---|---|
| Size | 12.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
1e51cacb579ead92f58af2474ea62a0676caba3acd576e280c3b764bdd653155
|
|
BLAKE2b-256 checksum How to use checksums |
ead7bcad5ad779c30377c99a64322750950148f2d5f86b83bdf0a19e5c7d4abd
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 29, 2026.
Transparency log