Skip to main content

datarobot-fastrag

FastRAG is an async-native rewrite of DRUM built on FastAPI. It serves DataRobot custom models using the same custom.py hook interface as DRUM, but handles each request on an async event loop instead of blocking a thread. For I/O-bound LLM workloads (vector DB + LLM calls), this typically gives 3–5× higher throughput at the same concurrency level.

Existing DRUM custom.py files work without modification — sync hooks are run in a thread pool automatically.

Features

  • async def chat() / async def score() run natively on the event loop
  • Sync hooks offloaded to a ThreadPoolExecutor (drop-in compatible with existing models)
  • OpenAI-compatible chat completions API (/v1/chat/completions, streaming included)
  • Predict/score API (/predict/)
  • OpenTelemetry instrumentation built in
  • Deployment prediction stats (Total Predictions and errors) reported to DataRobot
  • LLM safety guardrails via datarobot-moderations

Not implemented

  • Artifact guessing
  • Transform and custom tasks API
  • directAccess routes

Installation

FastRAG requires Python 3.12 or later.

pip install datarobot-fastrag

Or using uv:

uv add datarobot-fastrag

For local development from source:

git clone https://github.com/datarobot-oss/datarobot-fastrag
cd datarobot-fastrag
uv pip install -e .

Writing a custom LLM model

FastRAG uses the same custom.py hook convention as DRUM. For LLM models, you implement chat(). Making it async is what unlocks the throughput benefit.

Step 1 — create your model directory

my_model/
├── custom.py
└── model-metadata.yaml

model-metadata.yaml:

name: My LLM model
type: inference
targetType: textgeneration

Step 2 — implement custom.py

The minimal interface for a chat model:

# custom.py
import httpx

async def load_model(code_dir: str):
    # Return anything — it is passed as `model` to every hook.
    # Initialise your clients here (LLM, vector DB, etc).
    return httpx.AsyncClient()

async def chat(completion_create_params: dict, model, **kwargs):
    messages = completion_create_params["messages"]
    user_prompt = next(m["content"] for m in reversed(messages) if m["role"] == "user")

    # Both awaits release the event loop, so other requests run concurrently.
    # db_results = await model.post("http://vector-db/search", json={"q": user_prompt})
    # answer = await model.post("http://llm/generate", json={"prompt": ...})

    return {
        "id": "chatcmpl-1",
        "object": "chat.completion",
        "created": 0,
        "model": completion_create_params["model"],
        "choices": [{
            "index": 0,
            "message": {"role": "assistant", "content": f"Echo: {user_prompt}"},
            "finish_reason": "stop",
        }],
        "usage": {"prompt_tokens": 1, "completion_tokens": 1, "total_tokens": 2},
    }

Step 3 — run locally

fastrag server --code-dir ./my_model

Test it:

curl -X POST http://localhost:8080/v1/chat/completions/ \
  -H "Content-Type: application/json" \
  -d '{"model": "my-model", "messages": [{"role": "user", "content": "hello"}]}'

Or with the OpenAI Python client:

from openai import AsyncOpenAI

client = AsyncOpenAI(base_url="http://localhost:8080/v1", api_key="unused")
response = await client.chat.completions.create(
    model="my-model",
    messages=[{"role": "user", "content": "hello"}],
)
print(response.choices[0].message.content)

Step 4 — deploy on DataRobot

Upload custom.py and model-metadata.yaml as a custom model in the DataRobot UI, and select the [GenAI] Python 3.12 with Moderations execution environment. When the GENAI_RAG_FRAG_RUNNER platform flag is enabled for your org, the environment starts FastRAG automatically; otherwise it falls back to DRUM.

All supported hooks

Hook Signature Notes
load_model (code_dir: str) -> Any Return value becomes model in all other hooks. Runs once at startup.
init (code_dir: str) -> None Side-effecting setup (logging, connections). Runs before load_model.
chat (completion_create_params: dict, model: Any, **kwargs) -> dict | Iterator OpenAI chat completions. Return a dict or a (sync/async) generator for streaming.
score (data: pd.DataFrame, model: Any, **kwargs) -> pd.DataFrame Tabular predictions.
score_unstructured (data: Any, model: Any, **kwargs) -> Any Raw bytes in/out.
get_supported_llm_models (model: Any) -> list[Model] Populates /v1/models.

All hooks can be async def or plain def. Sync hooks run in a thread pool.

Streaming

Return a generator from chat() that yields ChatCompletionChunk objects (or plain dicts). Both sync generators and async generators work:

async def chat(completion_create_params, model, **kwargs):
    if completion_create_params.get("stream"):
        async def gen():
            for token in ["Hello", " world"]:
                yield {"choices": [{"delta": {"content": token}, "finish_reason": None, "index": 0}]}
            yield {"choices": [{"delta": {}, "finish_reason": "stop", "index": 0}]}
        return gen()
    # non-streaming fallback ...

A complete working example (including streaming) is at tests/models/python3_dummy_chat/custom.py.

Configuration

Configuration can be provided via CLI arguments or environment variables:

CLI Argument Environment Variable Description
--code-dir CODE_DIR Directory containing custom.py and model files.
--address ADDRESS Host and port to bind to (e.g., 0.0.0.0:8080).
--max-workers MAX_WORKERS Number of worker processes.
--verbose VERBOSE Enable verbose logging.

Prediction stats reporting

FastRAG reports chat-completion prediction counts to DataRobot using the same env vars as DRUM (EXTERNAL_WEB_SERVER_URL / API_TOKEN / DEPLOYMENT_ID / MODEL_ID; MLOPS_* aliases are accepted). Each chat request counts as 1, and 4xx/5xx are reported as userError / systemError with 0 predictions. Records are queued and POSTed in batches; reporting never blocks a request. Reporting is off only when those credentials are missing (typical for local runs).

Responses carry an X-Drum-Version header (1.17.12, overridable with FASTRAG_DRUM_VERSION) whenever reporting is on. DataRobot's predictions gateway reads that header to decide whether the model reports its own chat monitoring; with no header it assumes any 5xx went unreported and files a record of its own that carries one prediction, so a failed chat request would land twice in Total Requests and once in Total Predictions. The header is omitted when reporting is off, which keeps the gateway's coverage for deployments that report nothing.

Structured /predict/ requests are deliberately not reported. Those go through DataRobot's prediction API, which counts the scored rows itself and feeds them into deployment stats.

Development

The project uses uv for dependency management and a Makefile for common tasks.

To run the tests:

make test

To run with coverage:

make cov

To format the code:

make fmt

To clean up build and test artifacts:

make clean

Release files for datarobot-fastrag 0.2.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for datarobot-fastrag 0.2.2
File Size Uploaded
datarobot_fastrag-0.2.2.tar.gz 44.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for datarobot-fastrag 0.2.2
File Interpreter ABI Platform
datarobot_fastrag-0.2.2-py3-none-any.whl Python 3 none any Details

Total release size: 78.7 kB

Release files / datarobot_fastrag-0.2.2.tar.gz

Download URL datarobot_fastrag-0.2.2.tar.gz
Size 44.7 kB
Tags Source
SHA-256 checksum
How to use checksums
e9075da8e1ab00e4c9b6c6828e73a9f26204453986b4c3117f9577ab0e53c53c
BLAKE2b-256 checksum
How to use checksums
f1eb69d9cfb1966998c24b40c0c78f9dc42ce8bf966f89017d7b57bf0930ba22
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.9.30 {"installer":{"name":"uv","version":"0.9.30","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / datarobot_fastrag-0.2.2-py3-none-any.whl

Download URL datarobot_fastrag-0.2.2-py3-none-any.whl
Size 34.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
9e1ae0a3c16d8721369b57cd759897a6438f45bea01cba74af44d5250c7e4547
BLAKE2b-256 checksum
How to use checksums
652d6d2eb6944ebee0cbc373996460915df4bbc4360dc715c93fcee36b5cd56c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.9.30 {"installer":{"name":"uv","version":"0.9.30","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

0.2.7

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

This release

0.2.2 This release

2 release files

0.2.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page