Skip to main content

ratel-ai

Context engineering for Python agents.

DocsGitHubDiscord

PyPI GitHub stars MIT license

ratel-ai retrieves the tools and skills relevant to each agent turn instead of sending the full catalog to the model. It bundles Ratel's Rust engine in-process: BM25 by default, with configurable semantic and hybrid retrieval available when needed. The default and local-model paths require no API key, vector database, or service. Installing a published package on a supported prebuilt target also requires no Rust toolchain.

Use ToolCatalog for ranked tools with sync or async handlers and SkillCatalog for ranked Markdown playbooks loaded on demand. Expose search_capabilities_tool, invoke_tool_tool, and get_skill_content_tool so an agent can discover tools and skills, invoke tools, and load full skill instructions. Tools from existing MCP servers can be ingested into the tool catalog with the mcp extra. Experimental — facts: the opt-in ratel_ai.experimental namespace adds FactCatalog for constant grounding content (a shop's address, a brand's voice). See Facts below. This API may change or be removed without a major version bump.

Semantic and hybrid retrieval use a configurable embedding model (ADR 0012), set per catalog via the embedding argument: the built-in default, a HuggingFace repo or local directory (in-process), or an OpenAI-compatible endpoint (OpenAI, Ollama, TEI, vLLM).

For semantic or hybrid retrieval, register() folds embedding in: it accepts one tool or a whole batch and embeds on a worker thread, so model loading, HTTP, and inference never block the asyncio loop or hold the GIL — and embedding errors surface right at register():

async def retrieve(tools):
    catalog = ToolCatalog(method="semantic", embedding={"ollama": "nomic-embed-text"})
    await catalog.register(tools)                              # embeds the batch here
    return await catalog.search_async("deploy the service", 5)

register() is async for every method (BM25 too); search() stays synchronous for BM25 only, and search_async() covers all three. To change the endpoint's model or vector dimension, construct a new catalog and re-register.

A SkillCatalog also takes a whole reloaded catalog at once with replace_all(), for a source that fetches the full set rather than individual changes (ADR 0015). The batch is the catalog: ids missing from it are removed, including ones registered in-process, so a host that mixes local and remote skills composes the batch itself. It mutates in place, so every holder of the catalog sees the reload without being rebuilt.

outcome = await catalog.replace_all([*local_skills, *await fetch_remote_skills()])
print(f"reload: +{outcome.added} -{outcome.removed} ~{outcome.updated}")

The corpus swap is the synchronous half of that call, so the counts are already final when it returns — read them without awaiting and a reload whose embedding pass fails still reports what it changed:

reload = catalog.replace_all(batch)  # corpus is live; counts are final
try:
    await reload  # drives the embedding pass
except EmbedderError:
    log.warning("applied +%d -%d, embeddings pending", reload.added, reload.removed)

Only new and re-worded skills are embedded — reloading an unchanged catalog costs no embedding calls — and a reload that races an in-flight operation — dense work, but also an ordinary BM25 search_async holding the read lock — raises rather than applying half of itself.

Install

pip install ratel-ai
# MCP ingestion: pip install 'ratel-ai[mcp]'

Quickstart

Save as quickstart.py, then run python quickstart.py:

import asyncio
from ratel_ai import ExecutableTool, ToolCatalog

async def main():
    catalog = ToolCatalog()
    await catalog.register(
        ExecutableTool(
            id="get_weather",
            name="get_weather",
            description="Get the current weather for a city.",
            input_schema={"properties": {"city": {"type": "string"}}},
            output_schema={"type": "object"},
            execute=lambda args: {"forecast": f"Sunny in {args['city']}"},
        )
    )

    hit = catalog.search("What is the weather in Rome?", 1)[0]
    print(await catalog.invoke(hit.tool_id, {"city": "Rome"}))


asyncio.run(main())

Continue with the Python guide, capability tools, API reference, or the Pydantic AI example.

Facts (experimental)

Tools and skills are pulled — a query ranks them and only the winners reach the model. Facts are the opposite: constant content the agent should always work from (a shop's address, hours, a brand's voice), pushed into the context and deduplicated so it is injected once rather than every turn.

Facts live in the opt-in ratel_ai.experimental namespace and may change without a major version bump. Registering one is like a skill, plus a pin tier:

from ratel_ai.experimental import Fact, FactCatalog, Pin

facts = FactCatalog()
await facts.register([
    Fact(
        id="shop-address",
        name="shop address & hours",
        description="where the shop is and when it's open",
        body="Fade & Blade — 12 Baker Street, London. Open Mon–Sat 9am–7pm.",
        pin=Pin.ALWAYS,       # every turn, regardless of the query
    ),
    Fact(
        id="cancellation",
        name="cancellation policy",
        description="cancelling or rescheduling a booking, and refunds",
        body="Cancel at least 24h ahead for a full refund; same-day is a 50% fee.",
        pin=Pin.RETRIEVED,    # only when the turn's query ranks it in (default)
    ),
])

Then pick one of two injection modes per turn.

ground() — persist into your stored history. Returns only the facts not already present; render each body verbatim and keep it in the messages you save. It takes a list of per-message strings — flatten multi-part content yourself, and note that a bare str is rejected (it is itself a Sequence[str], so it would be iterated character by character):

def text_of(message: dict) -> str:
    content = message["content"]
    if isinstance(content, str):
        return content
    return "\n".join(part.get("text", "") for part in content)  # multi-part content

result = await facts.ground(user_text, [text_of(m) for m in messages])
for item in result.inject:
    messages.append({"role": "system", "content": item.body})  # verbatim — presence is the dedupe

Turn 1 injects the address; turn 2 sees it in the transcript and injects nothing. It re-injects only when the body is gone (compaction) or was edited — item.reason is "never" / "evicted" / "mutated".

ground_snapshot() — per call, nothing stored. Returns the full applicable set every time; put it in the request you're about to send and discard it:

snapshot = await facts.ground_snapshot(user_text)
payload = [{"role": "system", "content": f.body} for f in snapshot] + messages

Use ground() for a long-lived agent whose messages you persist; ground_snapshot() for one-shot or stateless calls, or to keep injected content out of your stored history.

Facts are host-driven: the model-facing search_capabilities tool is unchanged and never returns facts — you decide what is true and inject it, rather than letting the model discover it. Every decision is traced (fact_inject with its reason, fact_inject_skip, fact_snapshot), so the skip rate — the tokens you saved — is measurable. See ADR-0017.

Telemetry export is optional. With the otlp extra installed, configure_telemetry() reads RATEL_OTLP_ENDPOINT (falling back to the superseded RATEL_URL, which warns) and RATEL_API_KEY, wires trace and Logs exporters, and returns a shutdown handle. It exports only gen_ai.*/ratel.* signal spans and EventRecords by default — export_all_spans=True widens spans only. Message/tool content stays off by default; opt in with capture_content/include_span_and_events (see the telemetry guide for the capture modes and their privacy implications). Hosts that already own OpenTelemetry providers add both ratel_span_processor and ratel_log_record_processor instead.

Package layout: ratel_ai/ is the Python surface, native/ contains the PyO3 binding, and tests/ exercises both. For local development, create .venv with uv, install maturin, pytest, pytest-asyncio, ruff, and mypy, then run .venv/bin/maturin develop and .venv/bin/pytest.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

ratel_ai-0.8.0.tar.gz (232.3 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

ratel_ai-0.8.0-cp39-abi3-win_amd64.whl (4.0 MB view details)

Uploaded CPython 3.9+Windows x86-64

ratel_ai-0.8.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (4.1 MB view details)

Uploaded CPython 3.9+manylinux: glibc 2.17+ x86-64

ratel_ai-0.8.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (3.9 MB view details)

Uploaded CPython 3.9+manylinux: glibc 2.17+ ARM64

ratel_ai-0.8.0-cp39-abi3-macosx_11_0_arm64.whl (3.6 MB view details)

Uploaded CPython 3.9+macOS 11.0+ ARM64

ratel_ai-0.8.0-cp39-abi3-macosx_10_12_x86_64.whl (3.9 MB view details)

Uploaded CPython 3.9+macOS 10.12+ x86-64

File details

Details for the file ratel_ai-0.8.0.tar.gz.

File metadata

  • Download URL: ratel_ai-0.8.0.tar.gz
  • Upload date:
  • Size: 232.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for ratel_ai-0.8.0.tar.gz
Algorithm Hash digest
SHA256 2f16ef745517f30f7be4025c161f73150e4d656e5c13a691a8e12ed7a5a8bb21
MD5 4e075e8a1b8f9375eea6e36579cd5c8e
BLAKE2b-256 c9518442d13acbd17024ffb7105169d95ac373a0dfd883315fba3d7da5e4f859

See more details on using hashes here.

Provenance

The following attestation bundles were made for ratel_ai-0.8.0.tar.gz:

Publisher: release.yml on ratel-ai/ratel

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ratel_ai-0.8.0-cp39-abi3-win_amd64.whl.

File metadata

  • Download URL: ratel_ai-0.8.0-cp39-abi3-win_amd64.whl
  • Upload date:
  • Size: 4.0 MB
  • Tags: CPython 3.9+, Windows x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for ratel_ai-0.8.0-cp39-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 042c720eb1f7ecf1f7a7366f43889a7a495bcd125bf519f3b725edc64910749f
MD5 8cf42954be671b38eafe84e26f7c52d9
BLAKE2b-256 671dd012e404e777b33cbec01689473949ac8edc5a58393d4f90881bbf994ac8

See more details on using hashes here.

Provenance

The following attestation bundles were made for ratel_ai-0.8.0-cp39-abi3-win_amd64.whl:

Publisher: release.yml on ratel-ai/ratel

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ratel_ai-0.8.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for ratel_ai-0.8.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 ebf72f751c763006dab7fc43940b70eb58810eeded3fc2990f11112e798cb9ab
MD5 d0e1225123627720d47648887ebd1538
BLAKE2b-256 67431e5342bda3351e46f32c81b0894ea5be372f9cdf0565cb5d1ea660b485b0

See more details on using hashes here.

Provenance

The following attestation bundles were made for ratel_ai-0.8.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: release.yml on ratel-ai/ratel

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ratel_ai-0.8.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for ratel_ai-0.8.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 dfcfccee9d215cf4353ef77260e26a46fd91148e686d9fd7289ce6bc9aab6ef7
MD5 c373217fec60673c6e6acc98e484fee5
BLAKE2b-256 d6b24613421fda5367e65f0984c89e2805283504c20488110d20e0e7cf94a84f

See more details on using hashes here.

Provenance

The following attestation bundles were made for ratel_ai-0.8.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl:

Publisher: release.yml on ratel-ai/ratel

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ratel_ai-0.8.0-cp39-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for ratel_ai-0.8.0-cp39-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 821f849787de1c4a441e02016d4ea450ca1ff657420a7e69a2bbfb82cfd15c8e
MD5 76898023a3e801c470c15ae78059821b
BLAKE2b-256 2ddd8a2c2a1307687d9ae673d11bf34ee42f9f9552ae2d36f234aca5b1000ab9

See more details on using hashes here.

Provenance

The following attestation bundles were made for ratel_ai-0.8.0-cp39-abi3-macosx_11_0_arm64.whl:

Publisher: release.yml on ratel-ai/ratel

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ratel_ai-0.8.0-cp39-abi3-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for ratel_ai-0.8.0-cp39-abi3-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 f7585166c741827d026c93a68e489af1376ad1e0396579e6486859d0c461e114
MD5 a2055e284c27a2c3c4221b483c775bd0
BLAKE2b-256 9dbd7956e46c9ea1b8707e075477590b869de7e42ca40e6ffd3ca7aca8023820

See more details on using hashes here.

Provenance

The following attestation bundles were made for ratel_ai-0.8.0-cp39-abi3-macosx_10_12_x86_64.whl:

Publisher: release.yml on ratel-ai/ratel

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.12.0

6 files

0.11.0

6 files

0.10.0

6 files

0.9.0

6 files

This release

0.8.0 This release

6 files

0.7.0

6 files

0.6.0

6 files

0.5.2

6 files

0.5.1

6 files

0.5.0

6 files

0.4.2

6 files

0.4.1

6 files

0.4.0

6 files

0.3.0

6 files

0.2.0

6 files

0.1.6

6 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page