ray-serve-algenta
Deploy Algenta's /mcp surface behind Ray Serve:
ray_serve_algenta.deployment.AlgentaMCPProxy, a byte-transparent HTTP reverse proxy run as a
KubeRay RayService, plus a real multi-replica
conformance test suite.
This package is not shaped like its pydantic-ai-algenta / langchain-algenta /
haystack-algenta siblings, on purpose. None of those wrap an MCP client and expose typed
tools to a model-calling framework -- this one wraps nothing tool-shaped at all. Ray Serve is a
general-purpose model/service-serving layer; the only thing worth building here is proof that
Algenta's own /mcp route can sit safely behind it with real horizontal scale-out. What this
package ships:
ray_serve_algenta.deployment.AlgentaMCPProxy-- a@serve.deployment-decorated,@serve.ingress-fronted class that forwards every request at/mcp(method, headers, query string, body -- all of it, streamed in both directions) to an operator-configured upstream Algenta engine. It never parses JSON-RPC content, never names an MCP tool, and never makes a routing decision based on what a request contains -- see Why a raw reverse proxy, not an MCP-aware client/server pair below.manifests/rayservice.yaml-- a ready-to-adapt KubeRayRayServicemanifest:autoscaling_configwith a floor of 2 replicas and a ceiling of 8, an upstream URL sourced from a KubernetesSecret(never a literal endpoint), and inline comments on every prerequisite this repository does not set up for you (an image, the KubeRay operator, the Secret itself).- A real multi-replica conformance test suite (
tests/) that starts an actualray.init()+serve.run(...)of the realAlgentaMCPProxydeployment with 2 live replicas, in front of a real stub Algenta MCP server, and proves two things over real HTTP -- see Testing this package.
Prerequisites
- A running self-hosted Algenta engine, reachable over HTTP, with its
/mcpendpoint URL at hand (for examplehttp://localhost:8000/mcp). This package never talks to any Algenta-hosted cloud service -- see Self-hosted-first below. If you don't have an engine running yet, skip to Try it locally: it stands up a real stub engine for you, with nothing external to configure. - Python 3.10 or newer.
- No Algenta API key or credential is required by this package itself.
AlgentaMCPProxyforwards whatever headers a caller sends it, unchanged, straight to your upstream URL -- if your engine requires authentication, that's configured between the caller and the engine (or its ingress), not here. Seemanifests/rayservice.yaml's header comment for the Kubernetes/KubeRay equivalent.
Install
pip install ray-serve-algenta
Unlike litellm-algenta, where the config-generation code works standalone without litellm
itself installed, ray[serve] / fastapi / httpx are real, non-optional runtime dependencies
here -- ray_serve_algenta.deployment cannot be imported, let alone deployed, without them. See
pyproject.toml's dependency comment for a real, independently-reproduced gap
in ray[serve]'s own published extra (jinja2 is imported unconditionally by ray.serve's
internals but not declared as one of ray[serve]'s own dependencies as of ray==2.58.0) that
this package pins around so you don't have to rediscover it yourself.
Why Algenta's /mcp transport makes this worth building at all
Algenta's own apps/mcp_server/README.md (in thyn-ai/algenta, the engine's source repository)
states plainly: "The canonical /mcp route implements stateless Streamable HTTP." Verified by
reading that file directly during this package's design, not assumed from a changelog or an older
memory of the protocol. Stateless means no server-side session is pinned to the connection that
created it -- a caller's n-th request does not need to land on the same process, replica, or
even the same physical machine as its n-1-th. That is precisely the property that makes
horizontal scale-out safe without extra machinery: no sticky-session load balancer configuration,
no shared session store between replicas, nothing for AlgentaMCPProxy to coordinate. Every
replica this package runs is fungible by construction, because the upstream it forwards to was
already fungible.
(One correction worth being explicit about: an earlier internal note referenced a "2026-07-28"
MCP spec revision as already shipped engine-side. Read directly from source during this package's
design, the real, currently-live protocol tag on thyn-ai/algenta's /mcp route is 2025-11-25
-- the statelessness claim above is independently verified against that same file and holds
regardless of which spec revision is in effect, so this package does not depend on the
"2026-07-28" figure at all; it just doesn't repeat it.)
Self-hosted-first
AlgentaMCPProxy forwards to your own self-hosted Algenta engine, resolved per replica, in
order, from:
upstream_base_url=passed tobuild_app(...)(an escape hatch -- production deployments should use option 2 below),- the
ALGENTA_BASE_URLenvironment variable (whatmanifests/rayservice.yamlsets, from aSecret, on every head and worker pod), http://localhost:8000/mcp(the same self-hosted-first fallback every in-process package in this repository uses -- seepydantic_ai_algenta.toolset.DEFAULT_ALGENTA_BASE_URL-- kept for consistency with every other package's documented default, not because it's safe to leave unset here specifically:serve runbinds Ray Serve's own HTTP proxy tolocalhost:8000by default too, so running this package locally withALGENTA_BASE_URLunset makes it forward every request to itself instead of to an engine. Always setALGENTA_BASE_URLexplicitly to your engine's real address before running this package -- see Quick start (local) below. Meaningless inside a container regardless, where nothing is listening on its own loopback).
Never an Algenta-hosted default, at any layer.
Try it locally (no live engine required)
Nothing above requires a real Algenta engine to see working -- this package's own test suite
already includes a real stub one (tests/stub_server.py, a genuine fastmcp.FastMCP server
running in stateless_http=True mode, the same setting the real engine's /mcp route uses).
examples/try_it_locally.py starts that stub, deploys the real
AlgentaMCPProxy in front of it, and makes one real MCP tool call through the whole path --
nothing mocked, no external network, nothing to configure:
cd python
uv sync --package ray-serve-algenta --all-extras
uv run python ray-serve-algenta/examples/try_it_locally.py
Expect a few seconds of genuine Ray/Serve startup logging, followed by:
Stub Algenta engine listening at http://127.0.0.1:54798/mcp (stands in for your real one)
AlgentaMCPProxy is up at http://127.0.0.1:8000/mcp, forwarding to the stub above
Real MCP response, round-tripped through the real proxy:
{'dataset': 'try-it-locally', 'request_id': 'demo-1', 'rows': [{'value': 1}, {'value': 2}]}
(The stub's own port is assigned by your OS and will differ each run -- only the final response matters.)
That response traveled through the real AlgentaMCPProxy code over real HTTP, exactly like it
would against your own engine -- only the endpoint it forwarded to is a stub. Once you have a real
self-hosted Algenta engine running, move to Quick start (local) below and
point ALGENTA_BASE_URL at it instead.
Quick start (local)
# Your own self-hosted Algenta engine's MCP endpoint -- must be a DIFFERENT host:port from the
# one this proxy binds below. `serve run` starts Ray Serve's HTTP proxy on its own default,
# localhost:8000; pointing ALGENTA_BASE_URL at that same address makes this proxy forward every
# request to itself instead of to your engine. Substitute wherever your engine actually listens.
export ALGENTA_BASE_URL="http://localhost:9000/mcp"
serve run ray_serve_algenta.deployment:app # binds Ray Serve's default HTTP port, 8000
ray_serve_algenta.deployment:app is bound with the package's default autoscaling_config
(min_replicas=2, max_replicas=8, target_ongoing_requests=10) -- serve run will start 2
replicas immediately. Point any MCP client at http://localhost:8000/mcp; it talks to it exactly
like the real engine's own /mcp endpoint, because every byte this proxy sees is exactly what it
forwards.
Quick start (Kubernetes / KubeRay)
kubectl create secret generic algenta-mcp-upstream \
--from-literal=base-url="http://algenta-engine.your-namespace.svc.cluster.local:8000/mcp"
kubectl apply -f manifests/rayservice.yaml
See the manifest's own header comment for the prerequisites it does not set up for you (the KubeRay operator, a container image with this package installed).
Why a raw reverse proxy, not an MCP-aware client/server pair
Every other Python package in this repository sits inside an agent process, translating between
a model-calling framework's own tool-calling shape and Algenta's MCP tool surface --
pydantic-ai-algenta wraps pydantic_ai.mcp.MCPToolset, haystack-algenta wraps Haystack's own
MCPToolset, and so on. Ray Serve is not a model-calling framework and has no tool-calling shape
of its own to translate into -- it is infrastructure for running a service at scale. The only
honest thing to build here is what an infrastructure layer actually needs: a deployment that gets
the bytes from a caller to the upstream engine and back, correctly, under real concurrent load,
across real replicas. Teaching this proxy to parse JSON-RPC, recognize execute_decision, or
apply the tool-profile contract would not make it more capable -- it would just be a second,
parallel (and, given scripts/check-no-engine-dependency.py, forbidden) reimplementation of logic
the connected engine already owns, sitting in the one place in this repository's whole design that
was supposed to stay a dumb pipe. Tool-profile enforcement for traffic that happens to pass through
this proxy remains entirely the connected engine's job, exactly as it is for direct-to-engine MCP
traffic that never goes through Ray Serve at all.
Why no algenta-sdk dependency
Same reasoning as litellm-algenta and llamaindex-algenta: every package
in this repository may depend on at most one Algenta-owned thing, the published algenta-sdk
client -- but only if something in the package would actually use it. AlgentaMCPProxy forwards
raw HTTP bytes with httpx; it never constructs an MCP client, never calls a tool, and has no use
for an SDK object of any kind. Declaring algenta-sdk anyway, unused, would repeat exactly the
leftover-placeholder-dependency pattern an adversarial review is on record catching elsewhere in
this repository's history (see litellm-algenta's own README section of the same name).
Testing this package
The conformance suite (tests/test_multi_replica_conformance.py) starts a real ray.init() local
cluster and a real serve.run(...) of the real AlgentaMCPProxy deployment with 2 live replicas,
in front of a real stub Algenta MCP server (tests/stub_server.py, a real fastmcp.FastMCP
server on a real HTTP socket, run in stateless_http=True mode to match the real engine's own
documented behavior) -- never a mocked deployment and never a mocked upstream. It proves, over
real HTTP with a real fastmcp.Client:
- No cross-replica state leakage -- many real MCP tool calls, each carrying a unique
request_id, fired concurrently and interleaved through the proxy; every response must echo back exactly its own call's arguments. - Real multi-replica routing, not just multi-replica configuration -- every response carries
an
X-Algenta-Ray-Replicaheader naming which replica served it (seedeployment.py'sREPLICA_HEADER); the test asserts more than one distinct replica actually appears across a batch of calls, so a routing bug that always happened to answer from replica 0 would fail this check even though it would pass check 1. - The documented
ALGENTA_BASE_URLenv-var resolution path, exercised end to end rather than only through theupstream_base_url=escape hatch the other two tests use for convenience.
A fourth, fast, no-Ray-runtime-needed test (test_default_autoscaling_config_has_a_floor_of_at_ least_two_replicas) guards the shipped autoscaling_config default itself -- the other three
tests all pin num_replicas=2 explicitly via build_app(..., num_replicas=2) for CI speed and
determinism (real autoscaler up/down-scaling behavior is slow and nondeterministic to assert on in
CI, and asserting on it would not add anything to what this package actually needs to prove), so
none of them exercise the real default on its own.
cd python
uv sync --all-packages --all-extras
uv run pytest ray-serve-algenta -v
Each Ray-backed test takes real wall-clock seconds (ray.init/serve.run/serve.shutdown/
ray.shutdown each cost several seconds) -- assertions are bundled into as few full start/stop
cycles as the suite can honestly get away with, the same tradeoff litellm-algenta's real-proxy
conformance suite documents for its own, structurally similar, external-process test cost.
No GPU is required anywhere in this package or its test suite -- AlgentaMCPProxy does no
computation of its own.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file ray_serve_algenta-0.1.2.tar.gz.
File metadata
- Download URL: ray_serve_algenta-0.1.2.tar.gz
- Upload date:
- Size: 24.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.14.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
84329c3600112bd75487101005a626b0cc92b7e17e6b85e49c2b2a67cb024c1d
|
|
| MD5 |
b90e8b62764d7e93cdd4da4183a6d6f0
|
|
| BLAKE2b-256 |
b7b3a5d585e5fc8fef879dbd438684c8454f00f024b7579b9ec9601565d4f414
|
File details
Details for the file ray_serve_algenta-0.1.2-py3-none-any.whl.
File metadata
- Download URL: ray_serve_algenta-0.1.2-py3-none-any.whl
- Upload date:
- Size: 17.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.14.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2f00259f4e65a2292cd1afeb12acd2465c3425641bfd805e730c40becf446110
|
|
| MD5 |
e21420cfb22e4a65f94746ba28635c12
|
|
| BLAKE2b-256 |
c13b5d9bb095f227f636134a6cd179c760cc82a9d6e578bc4f602a8f4642d0ed
|