Skip to main content

Production HTTP server for LangChain/LangGraph agents — OpenAI-compatible API, streaming, session memory and agent discovery

Project description

Fast LangChain Server

Production HTTP server for LangChain/LangGraph agents — OpenAI-compatible API, streaming, session memory, authentication, middleware, and agent discovery.

License: MIT Python 3.11+ PyPI version

Overview

Fast LangChain Server transforms any LangChain/LangGraph agent into a production-ready HTTP service. Deploy with a single line of code and get streaming responses, automatic session management, pluggable authentication, a composable middleware chain, and built-in observability.

Features

  • OpenAI-Compatible API — Drop-in compatible with OpenAI clients (/v1/chat/completions)
  • Streaming — Real-time token streaming via Server-Sent Events (SSE)
  • Session Memory — Conversation history with local or Redis backends
  • Authentication — Pluggable AuthProvider: API keys, JWT Bearer tokens, or compose multiple via |
  • Middleware — Composable chain with built-ins: AuthMiddleware, TimingMiddleware, RateLimitMiddleware
  • Authorization — Per-endpoint AuthCheck rules with built-ins: require_scopes, allow_own_session, etc.
  • Lifespan — Composable startup/shutdown lifecycle via @lifespan decorator and | operator
  • A2A Protocol — Agent-to-Agent JSON-RPC 2.0 with autonomous execution support
  • OpenTelemetry — Distributed tracing with W3C TraceContext propagation
  • Production Ready — Health checks, structured logging, Docker-ready

Quick Start

Installation

pip install fast-langchain-server

Minimal example

# agent.py
from langchain.agents import create_agent
from langchain_openai import ChatOpenAI
from langchain_core.tools import tool
from fast_langchain_server import serve
import os

@tool
def add(a: float, b: float) -> float:
    """Add two numbers."""
    return a + b

# Create model (with auto-extracted config from environment)
model = ChatOpenAI(
    model=os.getenv("MODEL_NAME"),
    api_key=os.getenv("MODEL_API_KEY"),
    base_url=os.getenv("MODEL_API_URL"),
)

agent = create_agent(model=model, tools=[add])

# serve() now auto-extracts model config from the agent
app = serve(agent)

Configure environment variables in .env:

AGENT_NAME=my-agent
MODEL_NAME=gpt-4o
MODEL_API_URL=https://api.openai.com/v1
MODEL_API_KEY=sk-...

Run:

uvicorn agent:app

CLI

# Auto-discover the agent in agent.py and start the server
fast-langchain-server run agent.py

# Explicit attribute
fast-langchain-server run agent.py:app

Auto-extracted model configuration

The serve() function automatically extracts model configuration from your LangChain agent. This eliminates the need to redundantly specify model settings:

from langchain_openai import ChatOpenAI
from fast_langchain_server import serve

# Model configured once
model = ChatOpenAI(
    model="gpt-4o",
    api_key="sk-...",
    base_url="https://api.openai.com/v1",
)

agent = create_agent(model=model, tools=[...])

# serve() extracts everything automatically
app = serve(agent)  # No additional config needed!

Environment variables are read as fallback:

MODEL_NAME=gpt-4o
MODEL_API_URL=https://api.openai.com/v1
MODEL_API_KEY=sk-...
AGENT_NAME=my-agent  # Optional; auto-generated if missing

API Endpoints

Method Endpoint Description
GET /health Liveness probe
GET /ready Readiness probe
POST /v1/chat/completions OpenAI-compatible chat (streaming + non-streaming)
GET /.well-known/agent.json A2A agent discovery card
GET /memory/sessions List active sessions
DELETE /memory/sessions/{id} Delete a session
POST / A2A JSON-RPC 2.0 (when TASK_MANAGER_TYPE=local)

Chat completions

# Non-streaming
curl -X POST http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [{"role": "user", "content": "What is 2+2?"}],
    "model": "agent"
  }'
# Streaming
curl -N -X POST http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [{"role": "user", "content": "Tell me a joke"}],
    "stream": true
  }'

Session continuity — pass session_id in the body or via X-Session-ID header to continue a conversation:

curl -X POST http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"messages": [{"role": "user", "content": "And 3+3?"}], "session_id": "abc123"}'

Authentication

Add token verification to every request with any AuthProvider:

from fast_langchain_server import (
    create_agent_server,
    AuthMiddleware,
    EnvAPIKeyProvider,
    JWTProvider,
)

server = create_agent_server(tools=[...])

# Option A: API keys from environment variable AGENT_API_KEYS=sk-a,sk-b
server.add_middleware(AuthMiddleware(provider=EnvAPIKeyProvider()))

# Option B: JWT Bearer tokens validated against a JWKS endpoint
server.add_middleware(AuthMiddleware(
    provider=JWTProvider(
        jwks_url="https://auth.example.com/.well-known/jwks.json",
        audience="my-agent",
    )
))

# Option C: compose both (JWT first, API key fallback)
server.add_middleware(AuthMiddleware(
    provider=JWTProvider(...) | EnvAPIKeyProvider()
))

Requests must include the token in one of these headers:

Authorization: Bearer <token>
X-API-Key: <token>

The following endpoints are excluded from auth by default: /health, /ready, /.well-known/agent.json

Available providers

Provider Description
APIKeyProvider(keys) Static dict {"key": "owner"}
EnvAPIKeyProvider(env_var) Comma-separated keys from env var (default: AGENT_API_KEYS)
JWTProvider(jwks_url, audience) Validates Bearer tokens against JWKS. Requires PyJWT[crypto]
MultiAuth(*providers) Tries providers in order — created automatically via |

Custom provider

from fast_langchain_server import AuthProvider, AuthToken

class MyDatabaseProvider(AuthProvider):
    async def verify_token(self, token: str) -> AuthToken | None:
        user = await db.lookup_token(token)
        if user:
            return AuthToken(subject=user.id, scopes=user.scopes, raw=token)
        return None

server.add_middleware(AuthMiddleware(provider=MyDatabaseProvider()))

Middleware

Middleware intercepts requests before and after the agent runs. Add them with add_middleware() — they execute in the order added (first = outermost).

from fast_langchain_server import (
    create_agent_server,
    AuthMiddleware,
    TimingMiddleware,
    RateLimitMiddleware,
    EnvAPIKeyProvider,
)

server = create_agent_server(tools=[...])
server.add_middleware(TimingMiddleware())
server.add_middleware(AuthMiddleware(provider=EnvAPIKeyProvider()))
server.add_middleware(RateLimitMiddleware(max_rpm=60))

Built-in middlewares

Middleware Description
AuthMiddleware(provider) Verifies tokens and sets ctx.auth_token
TimingMiddleware(log_level) Logs elapsed time per request
RateLimitMiddleware(max_rpm) Token-bucket rate limiter per session (default: 60 req/min)

Custom middleware

from fast_langchain_server import AgentMiddleware

class LoggingMiddleware(AgentMiddleware):
    async def on_request(self, ctx, call_next):
        print(f"Request: session={ctx.session_id} input={ctx.user_input[:50]}")
        result = await call_next(ctx)
        print(f"Done: session={ctx.session_id}")
        return result

server.add_middleware(LoggingMiddleware())

Hooks available:

Hook When it runs
on_request(ctx, call_next) Once per chat request — wraps the full lifecycle
on_agent_run(ctx, call_next) Immediately before/after the LangGraph agent runs

Authorization

Separate from authentication — controls what an authenticated user can do.

from fast_langchain_server import (
    AuthorizationMiddleware,
    require_scopes,
    allow_own_session,
    all_of,
)

server.add_middleware(AuthorizationMiddleware({
    "/v1/chat/completions": require_scopes("chat"),
    "/memory/sessions":     require_scopes("admin"),
}))

Built-in checks

Check Description
require_scopes(*scopes) All listed scopes must be present in the token
allow_any_authenticated() Any valid token is sufficient
allow_own_session() Token subject must own the requested session
deny_all() Always denies (maintenance mode)
all_of(*checks) AND — all checks must pass
any_of(*checks) OR — at least one check must pass

Custom check

from fast_langchain_server.authorization import AuthContext

def my_ip_allowlist(ctx: AuthContext) -> bool:
    # ctx.token carries the verified AuthToken
    return ctx.token and ctx.token.subject in ALLOWED_SUBJECTS

Lifespan

Manage startup and shutdown of resources with composable lifecycle hooks.

from fast_langchain_server import create_agent_server, lifespan, DEFAULT_LIFESPAN

@lifespan
async def db_lifespan(server):
    server.lifespan_context["db"] = await connect_db()
    yield {}
    await server.lifespan_context["db"].close()

@lifespan
async def cache_lifespan(server):
    server.lifespan_context["cache"] = await connect_redis()
    yield {}
    await server.lifespan_context["cache"].aclose()

# Compose with | — enters left-to-right, exits right-to-left (LIFO)
server = create_agent_server(
    tools=[...],
    lifespan=DEFAULT_LIFESPAN | db_lifespan | cache_lifespan,
)

Access the context at runtime:

db = server.lifespan_context["db"]

DEFAULT_LIFESPAN includes OpenTelemetry init, startup/shutdown logging, the autonomous loop launcher, and graceful shutdown of memory and task manager.


AgentContext

Every request carries an AgentContext through the middleware chain and into the agent execution layer. Middlewares can read and enrich it:

ctx.session_id      # resolved session identifier
ctx.request_id      # UUID for this specific request
ctx.user_input      # last user message
ctx.model           # model name from the request
ctx.headers         # lowercased HTTP headers dict
ctx.otel_context    # W3C TraceContext for distributed tracing
ctx.auth_token      # AuthToken set by AuthMiddleware (or None)
ctx.endpoint        # HTTP path ("/v1/chat/completions")

ctx.set_meta("key", value)    # store custom data for downstream use
ctx.get_meta("key", default)  # retrieve it

await ctx.emit_progress("tool_call", "search_web")  # push SSE progress event

Configuration

All settings are driven by environment variables (or .env file):

Variable Required Default Description
AGENT_NAME Agent identifier
MODEL_API_URL LLM API base URL (OpenAI, Ollama, vLLM…)
MODEL_NAME Model name (gpt-4o, llama3.2…)
MODEL_API_KEY not-needed API key for the model endpoint
MODEL_TEMPERATURE 0.7 Sampling temperature
MODEL_MAX_TOKENS (none) Max tokens per response
AGENT_DESCRIPTION "AI Agent" Human-readable description
AGENT_INSTRUCTIONS System prompt
AGENT_PORT 8000 HTTP port
AGENT_LOG_LEVEL INFO Log level
AGENT_ACCESS_LOG false Enable uvicorn access log
MEMORY_ENABLED true Enable session memory
MEMORY_TYPE local local | redis | null
MEMORY_REDIS_URL For Redis Redis connection URL
MEMORY_CONTEXT_LIMIT 20 Messages to load per request
MEMORY_MAX_SESSIONS 1000 Max sessions in local memory
AGENT_API_KEYS Comma-separated API keys (used by EnvAPIKeyProvider)
TASK_MANAGER_TYPE none none | local (enables A2A)
AUTONOMOUS_GOAL Goal for autonomous execution loop
AUTONOMOUS_INTERVAL_SECONDS 0 Interval between autonomous runs
OTEL_ENABLED false Enable OpenTelemetry
OTEL_SERVICE_NAME OTel service name
OTEL_EXPORTER_OTLP_ENDPOINT OTel collector endpoint

Docker

# Build
docker build -t fast-langchain-server .

# Run
docker run -p 8000:8000 \
  -e AGENT_NAME=my-agent \
  -e MODEL_API_URL=https://api.openai.com/v1 \
  -e MODEL_NAME=gpt-4o \
  -e MODEL_API_KEY=sk-... \
  -e AGENT_API_KEYS=my-secret-key \
  -v $(pwd)/agent.py:/app/agent.py \
  fast-langchain-server

Architecture

HTTP Request
     │
     ▼
 [FastAPI]  — builds AgentContext(session_id, request_id, user_input, headers…)
     │
     ▼
 [AuthMiddleware]          verify token → ctx.auth_token
     │
     ▼
 [AuthorizationMiddleware] check scopes per endpoint
     │
     ▼
 [RateLimitMiddleware]     token bucket per session
     │
     ▼
 [TimingMiddleware]        measure elapsed time
     │
     ▼
 [AgentServer._run_agent / _stream_response]
     │
     ▼
 [LangGraph agent.ainvoke / astream]
     │
     ▼
 [Memory backend]          save messages
     │
     ▼
 HTTP Response / SSE stream

Module map:

Module Purpose
server.py AgentServer, create_agent_server, serve
context.py AgentContext — per-request state object
auth.py AuthProvider, APIKeyProvider, EnvAPIKeyProvider, JWTProvider, MultiAuth
middleware.py AgentMiddleware, AuthMiddleware, TimingMiddleware, RateLimitMiddleware
authorization.py AuthorizationMiddleware, require_scopes, allow_own_session, …
lifespan.py @lifespan, ComposedLifespan, DEFAULT_LIFESPAN
memory.py LocalMemory, RedisMemory, NullMemory
a2a.py A2A protocol, LocalTaskManager, JSON-RPC handlers
telemetry.py OpenTelemetry init and trace context propagation
cli.py fast-langchain-server CLI

Development

# Install dev dependencies
make dev

# Run tests (185 tests)
make test

# Lint
make lint

# Run locally with hot-reload
make run

License

MIT — see LICENSE for details.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

fast_langchain_server-0.3.0.tar.gz (81.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

fast_langchain_server-0.3.0-py3-none-any.whl (46.9 kB view details)

Uploaded Python 3

File details

Details for the file fast_langchain_server-0.3.0.tar.gz.

File metadata

  • Download URL: fast_langchain_server-0.3.0.tar.gz
  • Upload date:
  • Size: 81.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for fast_langchain_server-0.3.0.tar.gz
Algorithm Hash digest
SHA256 946160e76b1aec40e143240aab0f5a0e15bd3477a0833a7bb6cdd8abd84212e4
MD5 c82b3d2a8f9dfdfc0b5a14e9cd5756e1
BLAKE2b-256 d01f4d9ec4217822f6f7f34f21a99aac1f0106bbf7349f7887ffb9d011dbd6b0

See more details on using hashes here.

File details

Details for the file fast_langchain_server-0.3.0-py3-none-any.whl.

File metadata

File hashes

Hashes for fast_langchain_server-0.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 4f42e1610ed0d372f3a98cf7ebdbf82aa389010b32f4ce34889599da95ec5eef
MD5 26d2904f6b812b1bdeae5410e92acf6c
BLAKE2b-256 3bbe1b4d940fa1129cfe3c2ce38eaee215ee8a5efe6eb18beef2fc79b7bb79f4

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page