Skip to main content

Tabalyst

Tools for unfamiliar data.

Tabalyst is an open-source, local-first toolkit for understanding and working with structured data.

Its first tool is Tabalyst Report. The current CSV implementation, Tabalyst CSV Report, analyzes a CSV file and produces both a structured JSON profile and a self-contained interactive HTML report.

Tabalyst is in active alpha development. Its interfaces may still change while the shared toolkit architecture is being established.

Install

Tabalyst supports Python 3.11, 3.12, 3.13, and 3.14.

pip install --upgrade tabalyst

The same command installs Tabalyst or updates it. Tabalyst is in alpha and changes often: update it before each new test.

Tabalyst Report

Use case Command Destination
One file, automatic name tabalyst report data.csv data.html beside the source
One file, custom name tabalyst report data.csv -o report.html The file report.html
Several files, automatic names tabalyst report *.csv Beside each source
Several files, one directory tabalyst report *.csv -d reports/ The directory reports/

The simplest command keeps the source filename:

tabalyst report customers.csv

It creates these files beside the source:

customers.csv
customers.html
customers.json
executions.json

Use -o to choose a different HTML filename for one source:

tabalyst report customers.csv -o customer-analysis.html

Multiple files

Report several CSV files at once:

tabalyst report *.csv

Each report is created beside its source and keeps the source stem:

customers.csv → customers.html
orders.csv    → orders.html
products.csv  → products.html

Use -d to place all reports in one directory:

tabalyst report *.csv -d reports/

This produces:

reports/
├── customers.html
├── customers.json
├── orders.html
├── orders.json
├── products.html
├── products.json
└── executions.json

Both -d reports/ and -d reports are accepted. Quotes are only needed when a path contains spaces.

-o always names one output file and therefore accepts only one input. -d always names an output directory and accepts one or many inputs.

This is invalid because several inputs cannot share one output file:

tabalyst report *.csv -o report.html

Tabalyst rejects the command before processing any file. Use -d reports/ instead.

Safe batch behavior

Before processing begins, Tabalyst resolves every input and planned output. It stops the entire batch if output names collide or if an output already exists. Use --force only when replacing all matching report artifacts is intentional:

tabalyst report *.csv -d reports/ --force

If one CSV is malformed during analysis, Tabalyst reports that error, continues with the remaining files, and returns a non-zero exit code at the end.

Interactive terminals show accurate file and phase progress:

[2/8] orders.csv - Analyzing

Progress and diagnostics use standard error. Progress is disabled automatically outside a terminal and can be disabled explicitly with --no-progress or --quiet.

Useful options

tabalyst report data.csv --delimiter ";"
tabalyst report data.csv --encoding cp1252
tabalyst report data.csv --config tabalyst.json
tabalyst report data.csv --verbose
tabalyst report --help
tabalyst --version

python -m tabalyst accepts the same commands.

Tabalyst Sample

Create a smaller CSV without modifying the source:

tabalyst sample customers.csv --sample-method random --rows 1000 --seed 42

The default output is customers.sample.csv beside the source. Use -o to name the output for one input, or -d to sample several files into one directory:

tabalyst sample customers.csv --sample-method first --rows 100 -o test.csv
tabalyst sample *.csv --sample-method random --percent 5 -d samples/

Available methods are first, last, random and stratified. Stratified sampling approximately preserves the distribution of a selected field:

tabalyst sample customers.csv --sample-method stratified --field province --rows 1000 --seed 42

Sampling reads CSV records as a stream. Random and stratified sampling keep only the requested sample, plus stratum counts, in memory. Existing outputs require --force, and an input file is never overwritten.

What the report analyzes

  • Dataset dimensions, missing cells, duplicates, and quality observations.
  • Physical and semantic types with confidence and error rates.
  • Numeric, date, string-length, normalization, and value distributions.
  • Distinct values, representative examples, enum candidates, and date formats.
  • CSV record widths, quoting, encoding, and delimiter configuration.
  • A bounded raw-data preview while every record is analyzed.

The HTML report is self-contained and works without a CDN or network connection. The JSON profile contains the same canonical analysis result for scripts and future Tabalyst tools.

Python API

The same operations are available without the CLI:

import tabalyst

result = tabalyst.analyze(
    "customers.csv",
    "customers.html",
    separator=";",
)

batch = tabalyst.generate_reports(
    ["*.csv"],
    output_dir="reports",
)

sample = tabalyst.sample_csv(
    "customers.csv",
    method="random",
    rows=1000,
    seed=42,
)

analyze() returns the JSON-serializable profile for one report. generate_reports() returns the complete batch plan, successes, and failures. Expected failures derive from tabalyst.TabalystError.

Existing artifacts are never replaced silently. Pass force=True when replacement is intentional.

Configuration

Configuration files are strict JSON. A minimal file is:

{
  "csv": {
    "delimiter": ";",
    "encoding": "cp1252"
  }
}

Explicit CLI or Python arguments override the configuration file, which overrides Tabalyst defaults. See the configuration reference for all analysis settings.

Local-first behavior and current limits

Tabalyst performs analysis locally and adds no telemetry or remote processing. Generated JSON and HTML may contain source values and should be shared accordingly.

Report analysis currently loads the complete CSV into memory. Sampling uses a streaming reader with bounded row storage. File and phase progress is available for reports; row percentages, throughput estimates, recursive directory input, and parallel batch execution will require later ingestion work.

Examples and development

The repository includes small and synthetic public examples under examples/. See the examples README for regeneration commands.

python -m pytest
python -m ruff check .
python -m build

Additional documentation:

Please report defects and feature requests through the GitHub issue tracker.

License

Tabalyst is released under the MIT License.

Created by Gregory Borelli — Catalyseur Numérique.

Metadata

Release files for tabalyst 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tabalyst 0.3.0
File Size Uploaded
tabalyst-0.3.0.tar.gz 421.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tabalyst 0.3.0
File Interpreter ABI Platform
tabalyst-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 612.6 kB

Release files / tabalyst-0.3.0.tar.gz

Download URL tabalyst-0.3.0.tar.gz
Size 421.4 kB
Tags Source
SHA-256 checksum
How to use checksums
a69ee772b0fbf966f9f7c2f174c623468ba876aa59ef1948c06e54c914ca81af
BLAKE2b-256 checksum
How to use checksums
aa64fe5405dd9681e63568cb44a222ec8e6afc008ad823f33f38b96136369600
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 26, 2026.

Transparency log

Release files / tabalyst-0.3.0-py3-none-any.whl

Download URL tabalyst-0.3.0-py3-none-any.whl
Size 191.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
d4a89af33a8b97dc90ffb49b9d03f5ffb4d4e95de4815e32c16e541dec3fd285
BLAKE2b-256 checksum
How to use checksums
7e4a30ace222cde7fde70d8c547c5c58e154e58f2e6d36f0a8e3854d5a4eddb1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 26, 2026.

Transparency log

Release history Release notifications | RSS feed

0.5.1

2 release files

0.5.0

2 release files

0.4.4

2 release files

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

This release

0.3.0 This release

2 release files

0.2.1

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page