Skip to main content

Part of AgentSave — the Python SDK that cuts AI agent token costs ~23%. This repo (agentsave-inferroute) is an Enterprise-tier component and requires a valid AgentSave Enterprise license key. Only deploy this if you operate a vLLM or sGLang inference cluster and hold an Enterprise license.

agentsave-inferroute — PPD Inference Router

CI PyPI License: Proprietary Python Enterprise

agentsave-inferroute is a FastAPI HTTP proxy sidecar that implements PPD (append-Prefill-Decode) routing for multi-turn LLM agent workloads. It classifies each request as a Turn 1 (prefill-heavy) or Turn 2+ (decode-heavy) conversation turn, then dispatches it to the backend optimized for that access pattern — targeting a ~68% TTFT reduction for Turn 2+ requests over a uniform-routing baseline.


What is PPD Routing?

Modern LLM inference backends are tuned differently depending on the memory access pattern of the workload:

  • Turn 1 (prefill-heavy): The model processes a long prompt (system prompt, tool schemas, initial user message) for the first time. The computation is dominated by the attention prefill pass over a large KV cache miss. A prefill-optimized backend (higher parallelism, larger chunked prefill batch) handles this efficiently.

  • Turn 2+ (append-heavy / decode-heavy): The conversation history is already cached. The model appends only the new user message and generates the next assistant turn. The computation is dominated by autoregressive decoding over a warm KV cache. A decode-optimized backend (larger batch, speculative decoding, lower prefill overhead) handles this efficiently.

PPD routing separates these two request classes at the proxy layer and sends each to its optimal backend. The routing decision is made per request with zero latency overhead — the classifier is a single pass over the messages[] array.

The ~68% Turn 2+ TTFT reduction figure is a static architectural estimate for v0.1.0, derived from the PPD routing design and the static improvement estimates embedded in the scorer (ttft_improvement = 68.0 ms, tpot_degradation = 5.0 ms). It has not been measured on a real vLLM/SGLang cluster — real instrumentation against live inference backends is future work. Treat this as a starting point to validate and tune for your own hardware, model, and traffic mix, not a benchmarked result.


Architecture

flowchart LR
    Agent["AI Agent\n(any framework)"]
    IR["agentsave-inferroute\n:8080\n/v1/chat/completions"]
    CL["Turn Classifier\nTURN1 / TURN2+"]
    SC["PPD Scorer\nroute_score = w_ttft·Δttft − w_tpot·Δtpot"]
    PB["Prefill Backend\n(vLLM / sGLang)\nTurn 1"]
    DB["Decode Backend\n(vLLM / sGLang)\nTurn 2+"]
    MT["/metrics\n(Prometheus-compatible)"]

    Agent -->|POST /v1/chat/completions| IR
    IR --> CL
    CL --> SC
    SC -->|score ≤ 0 or TURN1| PB
    SC -->|score > 0 and TURN2+| DB
    IR -->|fire-and-forget| MT

The proxy is transparent to the upstream agent — it accepts and returns standard OpenAI-compatible /v1/chat/completions JSON, including streaming ("stream": true).

Classification rule: if messages[] contains any entry with "role": "assistant", the request is TURN2_PLUS; otherwise it is TURN1.

Scoring rule: route_score = PPD_W_TTFT × ttft_improvement − PPD_W_TPOT × tpot_degradation. With default weights (0.7 / 0.3) and static estimates (68.0 ms / 5.0 ms), score = 46.1, so all TURN2_PLUS requests are decode-routed.


Quick Start

Prerequisites

  • Docker
  • A running vLLM or sGLang cluster with two endpoints: one prefill-optimized, one decode-optimized
  • A valid AgentSave Enterprise license key (set as AGENTSAVE_TOKEN)

Run as a Docker sidecar

docker run -d \
  --name inferroute \
  -p 8080:8080 \
  -e BACKEND_URL=http://your-primary-backend:8000 \
  -e BACKEND_TYPE=vllm \
  -e AGENTSAVE_TOKEN=your-enterprise-license-key \
  ghcr.io/aks-builds/agentsave-inferroute:latest

Point your agent at http://localhost:8080 instead of your inference backend directly. All /v1/chat/completions calls are automatically classified and routed.

Verify it is running

curl http://localhost:8080/health
# {"status":"ok","backend_type":"vllm","version":"0.1.0"}

Configuration

All configuration is via environment variables. No config file is required.

Variable Required Default Description
BACKEND_URL No http://localhost:8000 Base URL of the primary (fallback) backend. Used for Turn 1 requests or when decode routing is not configured separately.
BACKEND_TYPE No vllm Inference backend adapter to use. Accepted values: vllm, sglang.
AGENTSAVE_TOKEN Yes "" Your AgentSave Enterprise license key (a JWT with tier: "enterprise", signed by AgentSave). Validated at startup and enforced on every /v1/chat/completions request — missing, malformed, expired, or non-Enterprise tokens get a 403. Also used as the Bearer token when posting metrics to the AgentSave dashboard.
AGENTSAVE_METRICS_URL No (disabled) Full URL to POST routing metrics to (e.g. https://app.agentsave.ai/ingest). Metrics are silently skipped if this is unset.
PPD_W_TTFT No 0.7 Weight applied to TTFT improvement in the PPD scoring formula. Controls how strongly TTFT gains influence the route decision.
PPD_W_TPOT No 0.3 Weight applied to TPOT degradation penalty in the PPD scoring formula. Must satisfy PPD_W_TTFT + PPD_W_TPOT = 1.0 for standard operation.

Backend adapter signal details

Backend Decode routing signal
vllm Adds HTTP header X-Route-Type: decode to the proxied request
sglang Appends query parameter router_prefix=decode to the proxied request URL

API Endpoints

POST /v1/chat/completions

Accepts any OpenAI-compatible chat completions request body. Classifies the turn, scores the routing decision, and proxies the request to the appropriate backend.

  • Requires a valid Enterprise license. If AGENTSAVE_TOKEN is missing, malformed, expired, or not tier "enterprise", this endpoint returns 403 with a JSON {"detail": "..."} body explaining the problem, and never contacts the backend. /health still works unauthenticated for orchestrator liveness checks.
  • Streaming ("stream": true) is fully supported via StreamingResponse.
  • The request body and all headers are forwarded to the upstream backend unchanged, except for the routing signal injected by the adapter.
  • Upstream timeout: 120 seconds.

Example:

curl -X POST http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Meta-Llama-3-8B-Instruct",
    "messages": [
      {"role": "user", "content": "Summarize the quarterly report."}
    ]
  }'

GET /health

Liveness probe. Always returns HTTP 200 — regardless of license state, so orchestrators can still tell the container is alive — with a JSON body including the configured backend type, router version, and current license status.

{
  "status": "ok",
  "backend_type": "vllm",
  "version": "0.1.0",
  "license": {"valid": true, "tier": "enterprise", "org": "Acme Corp", "error": null}
}

GET /metrics

Prometheus-compatible metrics endpoint. Reports routing decisions and backend activity for observability integration.


Testing

The test suite covers the classifier, scorer, dispatcher, adapters, metrics emitter, and full end-to-end routing paths.

# Install development dependencies
pip install -e ".[dev]"

# Run the full suite
pytest tests/ -v

77 tests across 10 files, verified on Python 3.11, 3.12, and 3.13. CI runs the full suite on every push and pull request.

Test file Coverage area Tests
test_classifier.py Turn type detection from message history 8
test_scoring.py PPDWeights validation + PPDScorer routing decisions 11
test_adapters_vllm.py vLLM adapter request construction and decode signaling 7
test_adapters_sglang.py sGLang adapter request construction and decode signaling 9
test_dispatcher.py Dispatch routing logic (prefill vs. decode paths) 4
test_metrics.py Metrics emission, auth header, error handling 7
test_app.py FastAPI routes, health endpoint, HTTP layer 7
test_integration.py End-to-end routing with live classifier and scorer 6
test_license.py Enterprise-license JWT validation (valid/expired/malformed/wrong-tier) 9
test_app_license.py /v1/chat/completions license enforcement, /health license reporting 9

Contributing

This is a proprietary Enterprise component. External contributions are not accepted at this time. If you are an Enterprise customer and have found a bug or have a feature request, open an issue or contact your AgentSave account team.

For the open-source SDK, see aks-builds/agentsave.


License

Proprietary. Use of this software requires a valid AgentSave Enterprise license. See LICENSE for terms.


agentsave-inferroute v0.1.0 — part of the AgentSave ecosystem

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

agentsave_inferroute-0.1.1.tar.gz (23.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

agentsave_inferroute-0.1.1-py3-none-any.whl (18.9 kB view details)

Uploaded Python 3

File details

Details for the file agentsave_inferroute-0.1.1.tar.gz.

File metadata

  • Download URL: agentsave_inferroute-0.1.1.tar.gz
  • Upload date:
  • Size: 23.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.15

File hashes

Hashes for agentsave_inferroute-0.1.1.tar.gz
Algorithm Hash digest
SHA256 b884263a9b08a992a9900d988f25942ca715e440a1382c4864c684b3649b10c4
MD5 84840e7911cd99609c466809754a9599
BLAKE2b-256 bbe5baafd3e438ea9daa4610f820103de7d44d5c3f867cf22271ac578a8a94bf

See more details on using hashes here.

File details

Details for the file agentsave_inferroute-0.1.1-py3-none-any.whl.

File metadata

File hashes

Hashes for agentsave_inferroute-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 416a403626177353ded559bfadd80b15c96e632bf5c373afb1275cfb6322d807
MD5 e789c45157985108b31b9747fbdb9fb7
BLAKE2b-256 9e27691eaa8c22aba91f94cd6caaab20fb14dfefbb745a4dc90d90cd3dcf3056

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page