Skip to main content

Charter your data — contract-governed local data exploration, powered by DuckDB

Project description

DataCharter

Query all your data locally — then hand your AI agents exactly the data you choose, and not one column more.

PyPI Python License: Apache-2.0

The big-words version: a local, federated data explorer with governed, regulated agentic data access, powered by DuckDB. Here's what that actually means 👇

🔍 Query all your data, locally — no pipelines, no warehouse, no waiting

  • Local CSV, Parquet, and JSON files
  • Postgres, MySQL, SQL Server, Snowflake, BigQuery — and more
  • JOIN a local CSV → a Snowflake table → a Parquet file in S3, in one SQL statement, all on your laptop
  • Yes, it's as unreasonable as it sounds. You kind of have to try it to believe it.

🤖 Connect an agent — and decide exactly what it's allowed to see

  • Claude Code — runs on your existing subscription, no API key
  • A model running fully local with Ollama
  • Any OpenAI-compatible agent
  • Grant or deny access in the UI or right in your data contracts, at every level: whole sources → individual tables → individual columns
  • PII is auto-detected and defaulted to no agent access — override per field if you really mean to
  • Don't take our word for it: flip on Agent view and see, column by column, exactly what your agent gets back when it runs a query. (Spoiler: the PII comes back •••.)

Wait, there's more!

Beyond local federation and governed agent access, you also get:

  • See answers as you type. Live results preview while you write SQL, one-click auto-charts, and a profiling panel (missing values, distributions, outliers) — no separate BI tool.
  • Safe by design. The engine is read-only by construction — no query can write, delete, or touch the filesystem — so pointing an AI (or a teammate) at your real databases can't do damage.
  • Point other AI tools at your data, too. A governed MCP server exposes the same read-only, PII-masked query tools to Cursor, Cline, or your own agent.
  • Trust every answer. Each result shows exactly which source columns it read — so you always know where a number came from.
  • Save, reuse, export. Snapshot a result as a reusable local table; export to CSV, Parquet, JSON, or XLSX.
  • Governance you can automate. Catch schema/PII drift in CI, diff data across sources, and define certified metrics — all from the command line.

DataCharter — live SQL preview, auto-charts, per-query provenance, and PII masking

Status: pre-release. V1 in development.

Quick start

# Try it instantly on generated demo data — no install, no config:
uvx datacharter serve          # needs `uv` → https://astral.sh/uv
# → serves at http://127.0.0.1:8321 (open it in your browser)

# Or install it:
pip install datacharter        # Python 3.11+

# Start your own workspace:
datacharter init               # scaffolds charter.yaml, queries/, .env.example
# → add a source: edit charter.yaml, or use the "Sources" panel in the UI
datacharter serve              # → http://127.0.0.1:8321

Then, once it's running, drag a CSV, Parquet, or JSON file onto the window to query it instantly — no config needed.

Optional natural-language agent — point it at any OpenAI-compatible endpoint:

export OPENAI_BASE_URL=...     # any OpenAI-compatible API
export OPENAI_API_KEY=...
datacharter serve

…or run fully local — no API key, no data leaves your machine (requires Ollama):

ollama pull qwen3:8b           # once
datacharter serve --local      # qwen3:8b by default (--model to change)

Why DataCharter

  • Your contracts are the catalog. charter.yaml describes sources, tables, and PII fields — the same contract spec your data team already writes, so there's no separate metadata store to maintain.
  • Real federation, not just a shared connection. Filters and projections are pushed down to each source — even across a cross-source join, every leg is filtered where its data lives. (Snowflake runs via connector extract, datacharter[snowflake], with the same pushdown into the extract.)
  • Local-first. One process, your machine, no cloud dependency. The optional --local agent runs a small open model via Ollama — no API key, no data leaves your machine.
  • The workspace is a directory. charter.yaml + queries/*.sql + .env.example — commit it, clone it, datacharter serve. Your team's whole exploration environment travels as a repo; secrets and local state never do.

DataCharter governs and audits your data, not just displays it. The full command set (drift, scan, diff, metric, mcp, and more) is in the CLI reference; the security model is in security.

Built on

DataCharter stands on excellent open-source foundations:

Testing uses VidaiMock, an Apache-2.0 mock LLM server, as the offline agent endpoint in CI.

DuckDB is a trademark of the DuckDB Foundation. DataCharter is an independent project and is not affiliated with or endorsed by the DuckDB Foundation.

License

Apache-2.0

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

datacharter-0.10.0.tar.gz (7.3 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

datacharter-0.10.0-py3-none-any.whl (4.0 MB view details)

Uploaded Python 3

File details

Details for the file datacharter-0.10.0.tar.gz.

File metadata

  • Download URL: datacharter-0.10.0.tar.gz
  • Upload date:
  • Size: 7.3 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for datacharter-0.10.0.tar.gz
Algorithm Hash digest
SHA256 f0f66094af18a40c927495a0083e6abc5dce9916b11fcb2015a92339d8436086
MD5 a75c4a75ff8ee2ac9da1cce4c49358d0
BLAKE2b-256 0ff74442cde7a7a241fb0551ca31ab7a27f003e7f5b0eeaf9b2ac7e744cc4b8a

See more details on using hashes here.

Provenance

The following attestation bundles were made for datacharter-0.10.0.tar.gz:

Publisher: release.yml on datacharter/datacharter

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file datacharter-0.10.0-py3-none-any.whl.

File metadata

  • Download URL: datacharter-0.10.0-py3-none-any.whl
  • Upload date:
  • Size: 4.0 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for datacharter-0.10.0-py3-none-any.whl
Algorithm Hash digest
SHA256 584cfecd31464ae4ccb9fe217a4780d976d7057b46adb8656a6d7eb9e73c66d0
MD5 937b09cc458f34eee9542eb3e6cb4dc9
BLAKE2b-256 fbee3cc44abc7ddf6a008ae8351a1bf246db578b714f927bee24f9be34bf6d50

See more details on using hashes here.

Provenance

The following attestation bundles were made for datacharter-0.10.0-py3-none-any.whl:

Publisher: release.yml on datacharter/datacharter

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page