Tabalyst
Tools for unfamiliar data.
Tabalyst is an open-source, local-first toolkit for understanding and working with structured data.
Its first tool is Tabalyst Report. The current CSV implementation, Tabalyst CSV Report, analyzes a CSV file and produces both a structured JSON profile and a self-contained interactive HTML report.
Tabalyst is in active alpha development. Its interfaces may still change while the shared toolkit architecture is being established.
Install
Tabalyst supports Python 3.11, 3.12, 3.13, and 3.14.
pip install --upgrade tabalyst
The same command installs Tabalyst or updates it. Tabalyst is in alpha and changes often: update it before each new test.
Tabalyst Report
| Use case | Command | Destination |
|---|---|---|
| One file, automatic name | tabalyst report data.csv |
data.html beside the source |
| One file, custom name | tabalyst report data.csv -o report.html |
The file report.html |
| Several files, automatic names | tabalyst report *.csv |
Beside each source |
| Several files, one directory | tabalyst report *.csv -d reports/ |
The directory reports/ |
| One JSON file | tabalyst report data.json |
data.report.html beside the source |
| From a scan document | tabalyst report --scan data.scan.json |
data.html beside the scan |
The simplest command keeps the source filename:
tabalyst report customers.csv
It creates these files beside the source:
customers.csv
customers.html
customers.json
executions.json
Use -o to choose a different HTML filename for one source:
tabalyst report customers.csv -o customer-analysis.html
JSON files
A JSON file gets a report too, named <stem>.report.html so its profile
<stem>.report.json never replaces the source. Each collection of records,
such as the customers array of {"customers": [...]}, is a dataset of the
report, and nested fields are columns named by their path, such as
address.city:
tabalyst report orders.json
Multiple files
Report several CSV files at once:
tabalyst report *.csv
Each report is created beside its source and keeps the source stem:
customers.csv → customers.html
orders.csv → orders.html
products.csv → products.html
Use -d to place all reports in one directory:
tabalyst report *.csv -d reports/
This produces:
reports/
├── customers.html
├── customers.json
├── orders.html
├── orders.json
├── products.html
├── products.json
└── executions.json
Both -d reports/ and -d reports are accepted. Quotes are only needed when a
path contains spaces.
-o always names one output file and therefore accepts only one input. -d
always names an output directory and accepts one or many inputs.
This is invalid because several inputs cannot share one output file:
tabalyst report *.csv -o report.html
Tabalyst rejects the command before processing any file. Use -d reports/
instead.
Safe batch behavior
Before processing begins, Tabalyst resolves every input and planned output. It
stops the entire batch if output names collide or if an output already exists.
Use --force only when replacing all matching report artifacts is intentional:
tabalyst report *.csv -d reports/ --force
If one CSV is malformed during analysis, Tabalyst reports that error, continues with the remaining files, and returns a non-zero exit code at the end.
Interactive terminals show accurate file and phase progress:
[2/8] orders.csv - Analyzing
Progress and diagnostics use standard error. Progress is disabled automatically
outside a terminal and can be disabled explicitly with --no-progress or
--quiet.
Useful options
tabalyst report data.csv --delimiter ";"
tabalyst report data.csv --encoding cp1252
tabalyst report data.csv --config tabalyst.json
tabalyst report data.csv --verbose
tabalyst report data.csv --workers 4
tabalyst report --help
tabalyst --version
python -m tabalyst accepts the same commands.
Tabalyst Sample
Create a smaller CSV without modifying the source:
tabalyst sample customers.csv --sample-method random --rows 1000 --seed 42
The default output is customers.sample.csv beside the source. Use -o to
name the output for one input, or -d to sample several files into one
directory:
tabalyst sample customers.csv --sample-method first --rows 100 -o test.csv
tabalyst sample *.csv --sample-method random --percent 5 -d samples/
Available methods are first, last, random and stratified. Stratified
sampling approximately preserves the distribution of a selected field:
tabalyst sample customers.csv --sample-method stratified --field province --rows 1000 --seed 42
Sampling reads CSV records as a stream. Random and stratified sampling keep
only the requested sample, plus stratum counts, in memory. Existing outputs
require --force, and an input file is never overwritten.
Tabalyst Scan
Describe every field of a CSV or JSON file in one JSON document:
tabalyst scan customers.csv
tabalyst scan orders.json --collection "$.customers[]"
The default output is customers.scan.json beside the source; -o, -d and
--force work as for the other commands. Tabalyst Scan reads the file once as
a stream, with memory bounded by configurable limits, and records for each
field its presence, native types, missing values, frequencies, exact
statistics, normalization variants, technical type and the result of every
detector: numbers with decimal commas, dates, booleans, enumerations, email
addresses, URLs, phone numbers, postal codes, currency amounts, percentages,
quantities, UUIDs and IP addresses. Values of sensitive fields, such as email
addresses, are masked by default. Results are written atomically: an
interrupted scan never leaves a partial file. Files of 16 MiB or more are
analyzed by several worker processes, with the same result; --workers
chooses their number, --workers 1 keeps one process.
Build the report from a scan document without reading the source again:
tabalyst report --scan customers.scan.json
Tabalyst refuses a scan whose source changed since it was written, or whose
settings differ from the scan settings of --config.
See Scan CSV and JSON files and the scan format.
What the report analyzes
- Dataset dimensions, missing cells, duplicates, and quality observations.
- Physical and semantic types with confidence and error rates.
- Numeric, date, string-length, normalization, and value distributions.
- Distinct values, representative examples, date formats, and semantic types such as enumerations, email addresses, phone numbers, and postal codes.
- CSV record widths, quoting, encoding, and delimiter configuration.
- A bounded raw-data preview while every record is analyzed.
The report is built on Tabalyst Scan: it reads the CSV once as a stream and masks values of sensitive columns, such as email addresses, by default. Ambiguous dates stay ambiguous; the report shows the evidence of the column without applying it. The HTML report is self-contained and works without a CDN or network connection. The JSON profile contains the same canonical analysis result for scripts and future Tabalyst tools.
Python API
The same operations are available without the CLI:
import tabalyst
result = tabalyst.analyze(
"customers.csv",
"customers.html",
separator=";",
)
batch = tabalyst.generate_reports(
["*.csv"],
output_dir="reports",
)
sample = tabalyst.sample_csv(
"customers.csv",
method="random",
rows=1000,
seed=42,
)
scan = tabalyst.scan("orders.json")
scans = tabalyst.generate_scans(["data/*.json"], output_dir="scans")
reports = tabalyst.generate_reports(["scans/*.scan.json"], from_scan=True)
analyze() returns the JSON-serializable profile for one report.
generate_reports() returns the complete batch plan, successes, and failures.
scan() returns a scan result without writing anything; generate_scans()
writes one .scan.json document per source; from_scan=True builds reports
from such documents.
Expected failures derive from tabalyst.TabalystError.
Existing artifacts are never replaced silently. Pass force=True when
replacement is intentional.
Configuration
Configuration files are strict JSON. A minimal file is:
{
"scan": {
"csv": {
"delimiter": ";",
"encoding": "cp1252"
},
"values": {"null_markers": ["N/A"]}
}
}
Analysis settings, CSV reading included, go in the scan object, shared by
tabalyst report and tabalyst scan. Top-level settings shape the report
presentation, and csv configures tabalyst sample. Explicit CLI or Python
arguments override the configuration file, which overrides Tabalyst defaults.
No configuration file is loaded unless it is passed with --config. See the
configuration reference
for all analysis settings.
Local-first behavior and current limits
Tabalyst performs analysis locally and adds no telemetry or remote processing. Generated JSON and HTML may contain source values and should be shared accordingly.
Reports, samples, and scans read files as a stream, with memory bounded by
configurable limits, including duplicate-row detection
(scan.limits.max_tracked_records). Reports and scans show the share of the file read;
throughput estimates, recursive directory input, and parallel batch execution
will require later work.
Examples and development
The repository includes small and synthetic public examples under examples/.
See the examples README
for regeneration commands.
python -m pytest
python -m ruff check .
python -m build
Additional documentation:
Please report defects and feature requests through the GitHub issue tracker.
License
Tabalyst is released under the MIT License.
Created by Gregory Borelli — Catalyseur Numérique.
Metadata
Release files for tabalyst 0.4.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| tabalyst-0.4.2.tar.gz | 543.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| tabalyst-0.4.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 862.2 kB
Release files / tabalyst-0.4.2.tar.gz
| Download URL | tabalyst-0.4.2.tar.gz |
|---|---|
| Size | 543.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7311c7bac484bf799309dee18f296af1a39fff6b6767c04ef273cd4d2b8f31e6
|
|
BLAKE2b-256 checksum How to use checksums |
b7cb422cc6d38d0758177373b64200e64ae3d6b051500926c7856fa46a9b2848
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency logRelease files / tabalyst-0.4.2-py3-none-any.whl
| Download URL | tabalyst-0.4.2-py3-none-any.whl |
|---|---|
| Size | 318.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7897cb36228fe81d5c24aa78feebc39f30773887fbf693d12836cf5cabed8510
|
|
BLAKE2b-256 checksum How to use checksums |
bd4d877bf1966b2491047cb6ce6f4d3bf1aaee14bea0ca4e0e5bfaffa445d2e5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency log