Skip to main content
Strake Logo

Strake

The AI Data Layer

License PRs Welcome Docs


Strake is the AI Data Layer. Not a query tool. Not a RAG pipeline. The sandboxed execution environment where agents meet your data and return answers, not rows.

Built on Apache Arrow DataFusion, Strake enables AI agents to discover, query, and process data across your entire stack (PostgreSQL, Snowflake, S3, and more) without the need for data movement or ETL. Give AI agents structured access to your entire data stack safely.

📚 Full Documentation: Check out the complete documentation for installation, architecture, and API references.


Key Features

  • MCP-Native Discovery: Built for the Model Context Protocol. Your agents immediately discover your entire data catalog and schemas.
  • Run Python, Not Prompts: Every agent execution runs inside strict native OS sandboxes for performance, or ephemeral MicroVMs for hardware-level isolation.
  • Zero-Copy Federation: Query Postgres, S3, Local Files, REST, gRPC, and more simultaneously with Pushdown optimization via Apache Arrow.
  • Read-Only by Default: Strict read-only enforcement by default to protect your upstream data sources.
  • Developer First: Built for engineers shipping agents to production. Type-safe configuration, rich CLI tooling, and local development workflows.
  • Python Native: Zero-copy integration with Pandas and Polars via PyO3.
  • GitOps Native: Manage your data mesh configuration as code. Version control your sources, policies, and metrics.
  • Observability: Built-in OpenTelemetry tracing and Prometheus metrics.
  • Enterprise Capabilities: OIDC Authentication, Row-Level Security, and Data Contracts (Enterprise Edition).

Code Mode: Don't Compute in Context

Most agents fail by swallowing thousands of raw SQL rows. Strake's Code Mode lets them process data in Python where it lives, inside a secure sandbox, sending only the parsed results that matter to the LLM.

import strake
from strake.mcp import run_python

script = """
# 1. Query 10M rows instantly via DataFusion
df = strake.sql("SELECT * FROM user_events").to_pandas()

# 2. Aggregate in Python to prevent context bloat
summary = df.groupby('feature_flag')['latency'].median()

# 3. Print exactly what the LLM needs
print(summary.to_json())
"""

# Runs isolated with OS Sandboxing or Firecracker VMs
result = await run_python(script)
print(result)

Quick Start (5-Minute Setup)

If you're building agents that need to query Postgres, S3, and a REST API in a single operation — without context overflow and without leaking credentials — Strake is the runtime you're missing.

1. Installation

Quick Install (Linux/macOS)

curl -sSfL https://strakedata.com/install.sh | sh

Install via Cargo (Rust)

cargo install --path crates/cli
cargo install --path crates/server

Python Client

pip install strake

2. Configuration (GitOps)

Initialize a new config and validate your sources:

# Initialize a new config
strake-cli init

# Validate configuration
strake-cli validate sources.yaml

# Sync configuration with live database schemas
strake-cli sync

3. Query with Python

First, define your data sources in a sources.yaml file:

sources:
  - name: local_files
    type: csv
    path: "data/*.csv"
    has_header: true
    tables:
      - name: measurements

Then, query using the Strake Python client:

import strake
import polars as pl

# Connect using your source configuration
conn = strake.connect(sources_config="sources.yaml")

# Query across sources using standard SQL
query = "SELECT * FROM measurements LIMIT 5"
data = conn.sql(query)

# Zero-copy integration with Polars/Pandas
df = pl.from_arrow(data)
print(df)

Project Structure

Component Description
strake-runtime Orchestration layer (Federation Engine, Sidecar).
strake-connectors Data source implementations (Postgres, S3, REST, etc).
strake-sql SQL Dialects, Query Optimization, and Substrait generation.
strake-common Shared types, configuration, and telemetry.
strake-error Unified error handling, error codes, and exception types.
strake-server Arrow Flight SQL server implementation.
strake-cli GitOps CLI for managing data mesh configurations.
strake-python Python bindings for high-performance data access.

Contributing

We welcome contributions! Please see our Contributing Guidelines for details on how to get started.

License

Strake is licensed under the Apache 2.0 license.

Metadata

Release files for strake 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distributions (wheels)

Table of built distributions (wheels) for strake 0.3.0
File Interpreter ABI Platform
strake-0.3.0-cp310-abi3-win_amd64.whl CPython 3.10 abi3 Windows x86-64 Details
strake-0.3.0-cp310-abi3-manylinux_2_28_x86_64.whl CPython 3.10 abi3 Linux glibc 2.28+ x86-64 Details
strake-0.3.0-cp310-abi3-macosx_11_0_arm64.whl CPython 3.10 abi3 macOS 11.0+ ARM64 Details

Total release size: 198.7 MB

Release files / strake-0.3.0-cp310-abi3-win_amd64.whl

Download URL strake-0.3.0-cp310-abi3-win_amd64.whl
Size 63.9 MB
Tags CPython 3.10 Windows x86-64 abi3
SHA-256 checksum
How to use checksums
319979fa6112d954b971d711e66f228e52b793ab9f70c5d288760ecf469fa1ac
BLAKE2b-256 checksum
How to use checksums
4eec133c1430336141c9dc446291c2b8d83ad246498fa95a59e6d183c78120d3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 22, 2026.

Transparency log

Release files / strake-0.3.0-cp310-abi3-manylinux_2_28_x86_64.whl

Download URL strake-0.3.0-cp310-abi3-manylinux_2_28_x86_64.whl
Size 72.2 MB
Tags CPython 3.10 Linux glibc 2.28+ x86-64 abi3
SHA-256 checksum
How to use checksums
8d876adf38a006fbe3fc91be7ec446af62e87b364654604cd0cf08ba5c5af627
BLAKE2b-256 checksum
How to use checksums
35a02f51c52c86c0bf17f9ca0ac406577d338fdcb5205b89de81ae2bd8556b7a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 22, 2026.

Transparency log

Release files / strake-0.3.0-cp310-abi3-macosx_11_0_arm64.whl

Download URL strake-0.3.0-cp310-abi3-macosx_11_0_arm64.whl
Size 62.6 MB
Tags CPython 3.10 abi3 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
539d8c2e5368d68fffbcd30ae6d5ebec19f35bdf430e61425ba2834cddf0a746
BLAKE2b-256 checksum
How to use checksums
17b3741cec76dccb0904748c0aaf521dac745d3c6032ba6f5c55c97f04e3ee4d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 22, 2026.

Transparency log
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page