Skip to main content

Charter your data — contract-governed local data exploration, powered by DuckDB

Project description

DataCharter

Explore all your data locally, in one place — contract-governed data exploration, powered by DuckDB

PyPI Python License: Apache-2.0

Contract-governed local data exploration, powered by DuckDB. Define your sources as data contracts (ODCS-compatible YAML), query them through DuckDB's SQL federation engine with real source pushdowns, and explore in a local web UI — SQL editor with live preview, auto-charts, profiling. Every answer shows which source columns it read; ask questions in plain language or serve the whole thing to an AI agent over MCP, with PII masked from the model — all on your machine.

What you can do

  • Drop a file, query it instantly. Drag a CSV, Parquet, or JSON onto the window and run SQL on it right away — no import, no schema setup.
  • Join across sources — no pipelines. Query and JOIN a Postgres table, a Parquet file, and a Snowflake table in a single SQL statement. No ETL, no copying everything into a warehouse first.
  • Connect all your data. Postgres, MySQL, SQLite, SQL Server, BigQuery, Snowflake, files on S3/GCS/Azure, and Iceberg/Delta tables — all through one engine.
  • See answers as you type. Live results preview while you write SQL, one-click auto-charts, and a profiling panel (missing values, distributions, outliers) — no separate BI tool.
  • Ask in plain English (optional). Turn a question into SQL and an answer with the built-in agent — bring your own model, or run one fully local with no API key.
  • Keep sensitive data away from the AI. Mark PII columns once; the agent and any connected AI see masked values (•••) while you still see the real data locally. Flip Agent view to see exactly what the model sees.
  • Safe by design. The engine is read-only by construction — no query can write, delete, or touch the filesystem — so pointing an AI (or a teammate) at your real databases can't do damage.
  • Point AI tools at your data, safely. A governed MCP server exposes read-only, PII-masked query tools to Cursor, Cline, or your own agent.
  • Trust every answer. Each result shows exactly which source columns it read — so you always know where a number came from.
  • Save, reuse, export. Snapshot a result as a reusable local table; export to CSV, Parquet, JSON, or XLSX.
  • Governance you can automate. Catch schema/PII drift in CI, auto-detect PII columns, diff data across sources, and define certified metrics — from the command line.

DataCharter — live SQL preview, auto-charts, per-query provenance, and PII masking

Status: pre-release. V1 in development.

Quick start

# Try it instantly on generated demo data — no install, no config:
uvx datacharter serve          # needs `uv` → https://astral.sh/uv
# → serves at http://127.0.0.1:8321 (open it in your browser)

# Or install it:
pip install datacharter        # Python 3.11+

# Start your own workspace:
datacharter init               # scaffolds charter.yaml, queries/, .env.example
# → add a source: edit charter.yaml, or use the "Sources" panel in the UI
datacharter serve              # → http://127.0.0.1:8321

Then, once it's running, drag a CSV, Parquet, or JSON file onto the window to query it instantly — no config needed.

Optional natural-language agent — point it at any OpenAI-compatible endpoint:

export OPENAI_BASE_URL=...     # any OpenAI-compatible API
export OPENAI_API_KEY=...
datacharter serve

…or run fully local — no API key, no data leaves your machine (requires Ollama):

ollama pull qwen3:8b           # once
datacharter serve --local      # qwen3:8b by default (--model to change)

Why DataCharter

  • Your contracts are the catalog. charter.yaml describes sources, tables, and PII fields — the same contract spec your data team already writes, so there's no separate metadata store to maintain.
  • Real federation, not just a shared connection. Filters and projections are pushed down to each source — even across a cross-source join, every leg is filtered where its data lives. (Snowflake runs via connector extract, datacharter[snowflake], with the same pushdown into the extract.)
  • Local-first. One process, your machine, no cloud dependency. The optional --local agent runs a small open model via Ollama — no API key, no data leaves your machine.
  • The workspace is a directory. charter.yaml + queries/*.sql + .env.example — commit it, clone it, datacharter serve. Your team's whole exploration environment travels as a repo; secrets and local state never do.

DataCharter governs and audits your data, not just displays it. The full command set (drift, scan, diff, metric, mcp, and more) is in the CLI reference; the security model is in security.

Built on

DataCharter stands on excellent open-source foundations:

Testing uses VidaiMock, an Apache-2.0 mock LLM server, as the offline agent endpoint in CI.

DuckDB is a trademark of the DuckDB Foundation. DataCharter is an independent project and is not affiliated with or endorsed by the DuckDB Foundation.

License

Apache-2.0

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

datacharter-0.3.1.tar.gz (7.2 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

datacharter-0.3.1-py3-none-any.whl (4.0 MB view details)

Uploaded Python 3

File details

Details for the file datacharter-0.3.1.tar.gz.

File metadata

  • Download URL: datacharter-0.3.1.tar.gz
  • Upload date:
  • Size: 7.2 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for datacharter-0.3.1.tar.gz
Algorithm Hash digest
SHA256 f5afc8834c0c9f09e9ede4e1d98e9040345299f06db817307d5de743b0e90e8c
MD5 eea8b196c21176602287f5b2ee697ce9
BLAKE2b-256 d2f5ea93b8903c460e8b62168609a2a2e535a552eb8153c1e4d25f90bf7ef6f2

See more details on using hashes here.

Provenance

The following attestation bundles were made for datacharter-0.3.1.tar.gz:

Publisher: release.yml on datacharter/datacharter

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file datacharter-0.3.1-py3-none-any.whl.

File metadata

  • Download URL: datacharter-0.3.1-py3-none-any.whl
  • Upload date:
  • Size: 4.0 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for datacharter-0.3.1-py3-none-any.whl
Algorithm Hash digest
SHA256 1622611832b107ac5e28b89d49cca44ac2e3065a2b1b4bacaf6794438e19ec4e
MD5 aa65052e211d521f01e2b487b1ae4f67
BLAKE2b-256 0a9492ae6222834b9ac552662a7189bbad72d8919338ae47c70879b52633508f

See more details on using hashes here.

Provenance

The following attestation bundles were made for datacharter-0.3.1-py3-none-any.whl:

Publisher: release.yml on datacharter/datacharter

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page