datarobot-fastrag
FastRAG is an async-native rewrite of DRUM built on FastAPI. It serves DataRobot custom models using the same custom.py hook interface as DRUM, but handles each request on an async event loop instead of blocking a thread. For I/O-bound LLM workloads (vector DB + LLM calls), this typically gives 3–5× higher throughput at the same concurrency level.
Existing DRUM custom.py files work without modification — sync hooks are run in a thread pool automatically.
Features
async def chat()/async def score()run natively on the event loop- Sync hooks offloaded to a
ThreadPoolExecutor(drop-in compatible with existing models) - OpenAI-compatible chat completions API (
/v1/chat/completions, streaming included) - Predict/score API (
/predict/) - OpenTelemetry instrumentation built in
- Deployment prediction stats (Total Predictions and errors) reported to DataRobot
- LLM safety guardrails via datarobot-moderations
Not implemented
- Artifact guessing
- Transform and custom tasks API
directAccessroutes
Installation
FastRAG requires Python 3.12 or later.
pip install datarobot-fastrag
Or using uv:
uv add datarobot-fastrag
For local development from source:
git clone https://github.com/datarobot-oss/datarobot-fastrag
cd datarobot-fastrag
uv pip install -e .
Writing a custom LLM model
FastRAG uses the same custom.py hook convention as DRUM. For LLM models, you implement chat(). Making it async is what unlocks the throughput benefit.
Step 1 — create your model directory
my_model/
├── custom.py
└── model-metadata.yaml
model-metadata.yaml:
name: My LLM model
type: inference
targetType: textgeneration
Step 2 — implement custom.py
The minimal interface for a chat model:
# custom.py
import httpx
async def load_model(code_dir: str):
# Return anything — it is passed as `model` to every hook.
# Initialise your clients here (LLM, vector DB, etc).
return httpx.AsyncClient()
async def chat(completion_create_params: dict, model, **kwargs):
messages = completion_create_params["messages"]
user_prompt = next(m["content"] for m in reversed(messages) if m["role"] == "user")
# Both awaits release the event loop, so other requests run concurrently.
# db_results = await model.post("http://vector-db/search", json={"q": user_prompt})
# answer = await model.post("http://llm/generate", json={"prompt": ...})
return {
"id": "chatcmpl-1",
"object": "chat.completion",
"created": 0,
"model": completion_create_params["model"],
"choices": [{
"index": 0,
"message": {"role": "assistant", "content": f"Echo: {user_prompt}"},
"finish_reason": "stop",
}],
"usage": {"prompt_tokens": 1, "completion_tokens": 1, "total_tokens": 2},
}
Step 3 — run locally
fastrag server --code-dir ./my_model
Test it:
curl -X POST http://localhost:8080/v1/chat/completions/ \
-H "Content-Type: application/json" \
-d '{"model": "my-model", "messages": [{"role": "user", "content": "hello"}]}'
Or with the OpenAI Python client:
from openai import AsyncOpenAI
client = AsyncOpenAI(base_url="http://localhost:8080/v1", api_key="unused")
response = await client.chat.completions.create(
model="my-model",
messages=[{"role": "user", "content": "hello"}],
)
print(response.choices[0].message.content)
Step 4 — deploy on DataRobot
Upload custom.py and model-metadata.yaml as a custom model in the DataRobot UI, and select the [GenAI] Python 3.12 with Moderations execution environment. When the GENAI_RAG_FRAG_RUNNER platform flag is enabled for your org, the environment starts FastRAG automatically; otherwise it falls back to DRUM.
All supported hooks
| Hook | Signature | Notes |
|---|---|---|
load_model |
(code_dir: str) -> Any |
Return value becomes model in all other hooks. Runs once at startup. |
init |
(code_dir: str) -> None |
Side-effecting setup (logging, connections). Runs before load_model. |
chat |
(completion_create_params: dict, model: Any, **kwargs) -> dict | Iterator |
OpenAI chat completions. Return a dict or a (sync/async) generator for streaming. |
score |
(data: pd.DataFrame, model: Any, **kwargs) -> pd.DataFrame |
Tabular predictions. |
score_unstructured |
(data: Any, model: Any, **kwargs) -> Any |
Raw bytes in/out. |
get_supported_llm_models |
(model: Any) -> list[Model] |
Populates /v1/models. |
All hooks can be async def or plain def. Sync hooks run in a thread pool.
Streaming
Return a generator from chat() that yields ChatCompletionChunk objects (or plain dicts). Both sync generators and async generators work:
async def chat(completion_create_params, model, **kwargs):
if completion_create_params.get("stream"):
async def gen():
for token in ["Hello", " world"]:
yield {"choices": [{"delta": {"content": token}, "finish_reason": None, "index": 0}]}
yield {"choices": [{"delta": {}, "finish_reason": "stop", "index": 0}]}
return gen()
# non-streaming fallback ...
A complete working example (including streaming) is at tests/models/python3_dummy_chat/custom.py.
Configuration
Configuration can be provided via CLI arguments or environment variables:
| CLI Argument | Environment Variable | Description |
|---|---|---|
--code-dir |
CODE_DIR |
Directory containing custom.py and model files. |
--address |
ADDRESS |
Host and port to bind to (e.g., 0.0.0.0:8080). |
--max-workers |
MAX_WORKERS |
Number of worker processes. |
--verbose |
VERBOSE |
Enable verbose logging. |
Prediction stats reporting
FastRAG reports chat-completion prediction counts to DataRobot using the same env
vars as DRUM (EXTERNAL_WEB_SERVER_URL / API_TOKEN / DEPLOYMENT_ID /
MODEL_ID; MLOPS_* aliases are accepted). Each chat request counts as 1, and
4xx/5xx are reported as userError / systemError with 0 predictions. Records
are queued and POSTed in batches; reporting never blocks a request. Reporting is
off only when those credentials are missing (typical for local runs).
Responses carry an X-Drum-Version header (1.17.12, overridable with
FASTRAG_DRUM_VERSION) whenever reporting is on. DataRobot's predictions gateway
reads that header to decide whether the model reports its own chat monitoring; with
no header it assumes any 5xx went unreported and files a record of its own that
carries one prediction, so a failed chat request would land twice in Total Requests and once in Total Predictions. The header is omitted when reporting is
off, which keeps the gateway's coverage for deployments that report nothing.
Structured /predict/ requests are deliberately not reported. Those go
through DataRobot's prediction API, which counts the scored rows itself and feeds
them into deployment stats.
Development
The project uses uv for dependency management and a Makefile for common tasks.
To run the tests:
make test
To run with coverage:
make cov
To format the code:
make fmt
To clean up build and test artifacts:
make clean
Release files for datarobot-fastrag 0.2.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| datarobot_fastrag-0.2.3.tar.gz | 44.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| datarobot_fastrag-0.2.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 78.7 kB
Release files / datarobot_fastrag-0.2.3.tar.gz
| Download URL | datarobot_fastrag-0.2.3.tar.gz |
|---|---|
| Size | 44.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
5c94aae1169c82f539e71dd181bcbc769a9c98f72c6ab9bcb9c4a168e23d5844
|
|
BLAKE2b-256 checksum How to use checksums |
610f12d8453ffba6d6956feefc9461b6391d35e644efbbb51b3954cdea3887cf
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.9.30 {"installer":{"name":"uv","version":"0.9.30","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / datarobot_fastrag-0.2.3-py3-none-any.whl
| Download URL | datarobot_fastrag-0.2.3-py3-none-any.whl |
|---|---|
| Size | 34.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
def92756e644e78dec35e4e120cdbbcf34c38c7784f548bfd62d034d5985eead
|
|
BLAKE2b-256 checksum How to use checksums |
a67f147e97bef0f48c0c59c13aea2c97a98ca0d4fd25ee06dab8804385e91a58
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.9.30 {"installer":{"name":"uv","version":"0.9.30","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|