Datacrafter - NoSQL ETL Tool
Datacrafter is an open-source NoSQL ETL (Extract, Transform, Load) tool designed for data extraction, transformation, and loading with a focus on NoSQL data formats. It provides a command-line interface for building data pipelines that extract data from various sources, process it, and load it into different destinations.
Note: This project is in alpha stage. Code migration from a closed repository is in progress, and documentation is being continuously improved.
Features
- NoSQL-first: JSON Lines and BSON are the native intermediate formats
- Compressed inputs:
.gz/.bz2/.xz/.zstfiles are read transparently by stream-capable sources - CLI-first YAML projects: declare extract → process → load in
datacrafter.yml - File and URL extraction: CSV, JSON, JSONL, XML, XLS/XLSX, ZIP+XML, patterned HTML indexes, RSS/Atom, DCAT catalogs, APIBackuper, trusted Python
collect()scripts - Record transforms:
keymap,typemap, custom Pythonprocess(record), plus optionalautotype(sample-based type inference) andautoid(stable_id) - Inspect:
datacrafter schemaanddatacrafter metricsread JSONL inoutput/(orcurrent/) - Dry-run:
datacrafter run --dry-runvalidates config and prints a plan without downloading or writing - Destinations: JSONL, BSON, CSV, Parquet (optional
pyarrow), MongoDB, ArangoDB, CouchDB, Meilisearch - Open-data packaging:
datapackage.jsonbeside file output;${MONGO_URI}/${VAR:-default}in YAML - Multiple extractors:
extractors:list sharing one processor and destination
Reserved in CLI but not implemented yet: builds / push / ui, automatic docs generation.
Installation
The distribution is published on PyPI as datacrafter-etl (the datacrafter
name there belongs to an unrelated project). The Python module and the CLI
command stay datacrafter.
Using pip (recommended)
pip install datacrafter-etl
From a GitHub release
pip install \
https://github.com/apicrafter/datacrafter/releases/download/v2.0.1/datacrafter_etl-2.0.1-py3-none-any.whl
From git
pip install git+https://github.com/apicrafter/datacrafter.git
From source
git clone https://github.com/apicrafter/datacrafter.git
cd datacrafter
pip install -e .
Requirements
- Python 3.10 or higher
- See requirements.txt for full dependency list
Quick Start
1. Initialize a Project
datacrafter init my-project
cd my-project
This creates a new project directory with a datacrafter.yml configuration file.
2. Configure Your Pipeline
Edit datacrafter.yml to define your data pipeline:
version: "1"
project-name: "my-project"
project-id: "unique-id"
extractor:
mode: "singlefile"
type: "file-csv"
method: "url"
config:
url: "https://example.com/data.csv"
processor:
config:
autotype: true
autoid: true
error_strategy: "skip"
keymap:
type: "names"
fields:
old_column: "new_column"
destination:
type: "file-jsonl"
fileprefix: "output"
3. Run Your Pipeline
datacrafter run
datacrafter run --dry-run
datacrafter schema
datacrafter metrics
4. Check Status
datacrafter status
Command Reference
Main Commands
datacrafter init [DIRECTORY] [--path PATH] [--name NAME]- Initialize a new projectdatacrafter run [--path PATH] [--verbose] [--quiet] [--dry-run]- Execute the data pipeline (or print a plan)datacrafter status [--path PATH]- Show status of latest pipeline executiondatacrafter check [--path PATH]- Validate configuration and environmentdatacrafter clean [--path PATH] [--storage]- Remove temporary filesdatacrafter log [--path PATH] [--lines N]- Show log of latest operationsdatacrafter schema [--path PATH]- Infer field types from output JSONLdatacrafter metrics [--path PATH]- Record counts and field histogramsdatacrafter version- Show version information
Configuration Commands
datacrafter config validate [--path PATH]- Validate project configurationdatacrafter config schema- Show expected configuration file schema
Planned Commands
datacrafter builds- Manage builds (create, remove, list)datacrafter push- Push data to remote storagedatacrafter ui- Launch web user interface
Core Concepts
Extractors
Extractors pull data from various sources:
- Local or remote files: CSV, JSON, XML, XLS/XLSX, BSON, JSONL, ZIP
- APIs and catalogs:
- APIBackuper ✅
- RSS/Atom feeds ✅
- DCAT catalogs ✅
- REST API (generic HTTP beyond URL/file download: Work in progress)
- CMS (Planned): WordPress, Microsoft SharePoint
- Common APIs (Planned): Email, FTP, SFTP
- Online services (Planned): Yandex Metrika, Yandex.Webmaster
Sources
Sources are files or databases created by extractors:
File Sources:
- JSON Lines ✅
- CSV ✅
- BSON ✅
- XLS/XLSX ✅
- XML ✅
- JSON ✅
- YAML (Work in progress)
- SQLite (Work in progress)
Database Sources (Planned):
- SQL databases via SQLAlchemy
- PostgreSQL, ClickHouse
- MongoDB, ArangoDB, ElasticSearch/OpenSearch
Processors
Processors transform data during the pipeline:
- Mappers: Map data fields from one schema to another
keymap: Replace key/column names ✅typemap: Convert data types ✅
- Custom code: Python scripts under the project directory ✅
- Custom tools: Command-line tools for data manipulation (Work in progress)
- Enrichers: Data and metadata enrichment (Planned)
Destinations
Destinations store the processed data:
File Destinations:
- BSON ✅
- JSON Lines ✅
- CSV ✅
- Parquet ✅
- Frictionless Data Package (
datapackage.jsonbeside file output) ✅ - JSON (Work in progress)
- YAML (Planned)
Database Destinations:
- MongoDB ✅
- ArangoDB ✅
- CouchDB ✅
- Meilisearch ✅
- ClickHouse (Planned)
- Any SQL via SQLAlchemy (Planned)
Storage Options (Planned):
- Local filesystem ✅
- S3, FTP, SFTP
- WebDAV, Google Drive, Dropbox, Yandex.Disk
Buzzers
Alerting mechanisms (Planned):
- Email alerts
- Other notification methods
Configuration
Project Structure
A datacrafter project typically has this structure:
my-project/
├── datacrafter.yml # Project configuration
├── current/ # Extracted data
├── output/ # Processed output
├── state.json # Execution state
└── datacrafter.log # Execution logs
Configuration Schema
See the full configuration schema:
datacrafter config schema
Or check the example below:
version: "1"
project-name: "my-project"
project-id: "unique-id"
extractor:
mode: "singlefile" # singlefile, api, code
type: "file-csv" # file-csv, file-json, file-xml, etc.
method: "url" # url, urlbypattern, apibackuper
force: true # Force re-download
config:
url: "https://example.com/data.csv"
processor:
config:
error_strategy: "skip" # skip, fail, retry
max_retries: 3
keymap: # Optional field mapping
type: "names"
fields:
old_name: "new_name"
typemap: # Optional type conversion
field_name: "int" # int, float, date, datetime, bool
custom: # Optional custom code
type: "script"
code: "path/to/script.py"
destination:
type: "file-jsonl" # file-jsonl, file-csv, file-bson, file-parquet, mongodb, arangodb, couchdb, meilisearch
fileprefix: "output"
compress: "gz" # Optional: gz, bz2, xz, zip, zst
Examples
Starter datacrafter.yml recipes live in examples/ (CSV URL, Excel, ZIP+XML, APIBackuper, RSS, DCAT). More recipes: https://github.com/apicrafter/datacrafter-examples
Example: Extract CSV and Convert to JSONL
version: "1"
project-name: "csv-to-jsonl"
project-id: "example-1"
extractor:
mode: "singlefile"
type: "file-csv"
method: "url"
config:
url: "https://example.com/data.csv"
processor:
config:
error_strategy: "skip"
destination:
type: "file-jsonl"
fileprefix: "output"
Example: Extract from API and Store in MongoDB
version: "1"
project-name: "api-to-mongo"
project-id: "example-2"
extractor:
mode: "api"
type: "api"
method: "apibackuper"
config:
endpoint: "https://api.example.com/data"
processor:
keymap:
type: "names"
fields:
api_id: "_id"
api_name: "name"
destination:
type: "mongodb"
connstr: "mongodb://localhost:27017"
dbname: "mydb"
tablename: "mydata"
Development
Running Tests
# Install dev dependencies (includes pytest, coverage, mypy, pip-audit)
pip install -r requirements-dev.txt
pip install -e .
# Fast local test run (no coverage artifacts)
pytest
# With coverage — CI enforces the 80% floor from .coveragerc
pytest --cov=datacrafter --cov-report=term-missing
# Audit dependencies for known vulnerabilities
pip-audit -r requirements.txt
Code Quality
# Linting with pylint
pylint datacrafter/
# Linting with ruff (also run via pre-commit)
ruff check datacrafter tests
# Type checking (whole package, also gated in CI)
mypy datacrafter
Building & Packaging
The project uses modern PEP 621 packaging (pyproject.toml); setup.py is kept
only as a compatibility shim. Runtime dependencies are sourced from
requirements.txt (single source of truth).
python -m build # produces wheel + sdist in dist/
CI runs the test matrix (Python 3.10–3.13), pip-audit, and publishes to PyPI on tag via Trusted Publishing. See CONTRIBUTING.md for details.
Contributing
Contributions are welcome! Please see CONTRIBUTING.md for setup, testing, linting, branching conventions, and the pull-request process.
- Fork the repository
- Create a feature branch (
feat/...) or fix branch (fix/...) - Make your changes
- Add tests if applicable
- Submit a pull request
Documentation
The documentation website lives in docs/ (Docusaurus). Preview locally
with cd docs && npm install && npm start. After GitHub Pages is enabled it
deploys to https://apicrafter.github.io/datacrafter/.
- Getting started - First pipeline
- CONTRIBUTING.md - How to set up and contribute
- Dependencies - Dependency management guide
- CHANGELOG.md - Version history
Security
Trust model. Datacrafter runs configuration files (datacrafter.yml) and
code-type extractor / custom processor scripts as trusted input — these can
execute arbitrary Python (via runpy) and should only come from a source you
control. Scripts MUST resolve inside the project directory. URLs, filenames,
and downloaded data are treated as untrusted and are never passed to a shell.
Secrets belong in the environment (connstr: ${MONGO_URI}), not in committed YAML.
- TLS certificate verification is enabled by default for all HTTPS downloads. Disable it only for trusted endpoints with a known self-signed cert (a warning is logged).
- To report a security vulnerability, please open a private advisory via GitHub Security Advisories rather than a public issue.
License
Licensed under the Apache License 2.0. See LICENSE for details.
Support
- Issues: https://github.com/apicrafter/datacrafter/issues
- Examples: https://github.com/apicrafter/datacrafter-examples
Author
Ivan Begtin
Status: Alpha - Active development in progress
Metadata
Release files for datacrafter-etl 2.0.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| datacrafter_etl-2.0.1.tar.gz | 104.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| datacrafter_etl-2.0.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 182.4 kB
Release files / datacrafter_etl-2.0.1.tar.gz
| Download URL | datacrafter_etl-2.0.1.tar.gz |
|---|---|
| Size | 104.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
bd59fb46d33127968db02b19cc662f282b8a4bc4489d3c0a01bbda5eb530ff56
|
|
BLAKE2b-256 checksum How to use checksums |
d57c0fa25a60fb94a95c1d245de94f861365c5a794ef45fa50021d9a9f974ad7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.7
|
Release files / datacrafter_etl-2.0.1-py3-none-any.whl
| Download URL | datacrafter_etl-2.0.1-py3-none-any.whl |
|---|---|
| Size | 78.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7fb5f6471360e7caa0b242d0aa7662c97c4016cae278019a7a06d5232759345c
|
|
BLAKE2b-256 checksum How to use checksums |
f0965b754ef601d87e1dc14cd3326b4ab49f0f1508b29713cddb09a814e0fc8c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.7
|