Skip to main content

Professional data format converter with powerful query language and library API for JSON, YAML, TOML, XML, and CSV

Reason this release was yanked:

bugged xml

Project description

dataconv

Professional data format converter with powerful query language and library API

Version Python Tests License

dataconv is a versatile tool for converting data between multiple formats (JSON, YAML, TOML, XML, CSV) with advanced filtering, JSONPath extraction, and boolean query capabilities. Use it as a library in your Python projects or as an interactive CLI tool.


Key Features

  • Multi-Format Support - JSON, YAML, TOML, XML, CSV with auto-detection
  • Library API - Clean Pythonic interface for programmatic use
  • Interactive CLI - MySQL-style REPL with live command execution
  • JSONPath Queries - Extract nested data with $.users[*].name syntax
  • Boolean Filtering - Complex WHERE clauses with AND, OR, NOT, XOR operators
  • Type-Safe - Full mypy compliance with comprehensive type hints
  • Format Validation - Built-in validators for each format
  • Runtime Configuration - Customize behavior with options system
  • Battle-Tested - 226 comprehensive tests, 100% passing
  • Performance - Fast JSON with orjson (3x speedup), optional C-YAML for 5x boost

Installation

Choose Your Installation

dataconv offers flexible installation options based on your needs:

Library Only

For embedding in your Python projects without the interactive CLI:

pip install dataconv

Includes:

  • Core data conversion engine
  • Query parser (Lark)
  • JSONPath support (jsonpath-ng)
  • All format support (JSON, YAML, TOML, XML, CSV)
  • Fast JSON processing (orjson - 3x faster)
  • Library API (load, save, convert, query, filter, extract_path)

Excludes: Rich (CLI terminal output), interactive REPL
Size: ~7 MB

CLI Application (Recommended)

Includes beautiful terminal output for the interactive REPL:

pip install dataconv[cli]

Includes: Everything from library + Rich terminal output
Use for: Interactive data conversion, command-line workflows
Size: ~10 MB

Full Installation

Complete installation with all extras and development tools:

pip install dataconv[full]

Includes: CLI + dev tools (pytest, mypy, black, ruff)
Use for: Development, contributing to the project
Size: ~30 MB

From Source

git clone https://github.com/thaisya/dataconv
cd dataconv

# Library only
pip install -e .

# With CLI
pip install -e ".[cli]"

# Full development setup
pip install -e ".[full]"

Requirements: Python 3.10+


Quick Start

As a Library

from dataconv import load, save, convert, query, filter, extract_path

# Load any format (auto-detected)
data = load("config.json")
users = load("data.yaml")

# Convert between formats
convert("input.json", "output.yaml")

# JSONPath extraction
users = load("data.json[$.users[*]]")
names = extract_path(data, "$.users[*].name")

# Filter with conditions
active_users = filter(users, "age > 18 and status == \"active\"")

# Complex queries
result = query('from data.json[$.users[*]] where age > 25 and premium == true')

As a CLI

# Start interactive REPL
dataconv

# Or run directly
python -m dataconv

Interactive Session

DataConv> from data.json to output.yaml
[+] Successfully converted data.json → output.yaml

DataConv> from users.json[$.users[*]] where age > 25 to adults.yaml
[+] Filtered 15 records → adults.yaml

DataConv> load employees.csv
[+] Loaded employees.csv (247 records)

DataConv> show
{
  "employees": [...]
}

Library API Reference

load()

Load data from any supported format with auto-detection.

def load(path: str | Path, **options: Any) -> dict | list

Features:

  • Auto-detects format from extension
  • Supports JSONPath extraction in path
  • Handles literal brackets in filenames
  • Configurable encoding, separators

Examples:

# Basic loading
data = load("config.json")
data = load("data.yaml", encoding="utf-16")

# JSONPath extraction
users = load("data.json[$.users[*]]")
names = load("data.json[$.users[*].name]")

# Literal brackets in filename (file exists)
archive = load("backup[2024].json")

save()

Save data to any format with atomic writes.

def save(data: dict | list, path: str | Path, **options: Any) -> None

Features:

  • Atomic file writes (rename, not overwrite)
  • Auto-detects format from extension
  • Pretty-printing with configurable indentation
  • Custom encoding support

Examples:

# Basic saving
save(data, "output.json")
save(data, "config.yaml", indent=4)

# Custom options
save(data, "data.json", sort_keys=True, ensure_ascii=False)

convert()

One-step format conversion with optional JSONPath extraction.

def convert(source: str | Path, dest: str | Path, **options: Any) -> None

Features:

  • Auto-detects source and destination formats
  • Supports JSONPath in source path
  • Preserves data structure
  • Configurable conversion options

Examples:

# Simple conversion
convert("data.json", "data.yaml")
convert("config.toml", "config.json")

# Convert with extraction
convert("data.json[$.users[*]]", "users.csv")
convert("nested.yaml[$.items[*]]", "items.toml")

extract_path()

Apply JSONPath expression to data.

def extract_path(data: dict | list, path: str) -> Any

Examples:

data = {"users": [{"name": "Alice", "age": 30}, {"name": "Bob", "age": 25}]}

# Extract all users
users = extract_path(data, "$.users[*]")
# [{"name": "Alice", "age": 30}, {"name": "Bob", "age": 25}]

# Extract specific fields
names = extract_path(data, "$.users[*].name")
# ["Alice", "Bob"]

# Complex paths
first_user = extract_path(data, "$.users[0]")

filter()

Filter data using WHERE clause conditions.

def filter(data: dict | list, conditions: str) -> dict | list

Features:

  • Comparison operators: ==, !=, <, >, <=, >=
  • Boolean operators: and, or, not, xor
  • Standalone field checks: where active, where !deleted
  • Nested field access: user.profile.age > 18

Examples:

users = [
    {"name": "Alice", "age": 30, "active": True},
    {"name": "Bob", "age": 25, "active": False},
    {"name": "Charlie", "age": 35, "active": True}
]

# Simple conditions
adults = filter(users, 'age >= 30')
active = filter(users, 'active == true')

# Complex boolean logic
result = filter(users, 'age > 25 and active == true')
result = filter(users, '(age < 30 or age > 40) and active')

# Standalone field checks
active_users = filter(users, 'active')  # Truthy check
inactive = filter(users, '!active')     # Falsy check

query()

Execute full query language (source, extraction, filtering).

def query(query_str: str, **options: Any) -> dict | list

Syntax:

from <source>[optional_jsonpath] [where conditions]

Examples:

# Load and filter
result = query('from data.json where age > 25')

# Extract and filter
result = query('from data.json[$.users[*]] where active == true')

# Complex queries
result = query('''
    from employees.csv[$.data[*]]
    where (department == "Engineering" and salary > 100000)
       or (department == "Sales" and sales > 50000)
''')

CLI Reference

Available Commands

Query Execution

Syntax:

from <source> [to <dest>] [where conditions]

Examples:

from data.json to output.yaml
from users.json[$.users[*]] where age > 25 to adults.yaml
from config.toml to config.json where env == "production"

Helper Commands

  • load <file> - Load and display file contents
  • save <file> - Save current data to file
  • show - Display currently loaded data
  • validate <file> - Check file format validity
  • options - View current configuration
  • set <option> <value> - Change runtime configuration
  • help - Show command help
  • clear - Clear screen
  • exit / quit - Exit REPL

Query Language

JSONPath Syntax

Extract nested data using JSONPath expressions:

$.root                    # Root level
$.users[*]                # All users
$.users[0]                # First user
$.users[*].name           # All names
$.items[?(@.price > 10)]  # Filter in JSONPath

WHERE Clause

Filter data with boolean expressions:

Comparison Operators:

  • == - Equal
  • != - Not equal
  • < - Less than
  • > - Greater than
  • <= - Less than or equal
  • >= - Greater than or equal

Boolean Operators:

  • and - Logical AND
  • or - Logical OR
  • not - Logical NOT
  • xor - Exclusive OR
  • () - Grouping/precedence

Standalone Field Checks:

  • where active - Truthy check
  • where !deleted - Falsy check

Examples:

where age > 18
where status == "active" and premium == true
where (age < 25 or age > 65) and not deleted
where department == "Sales" xor region == "West"
where active and not archived

Configuration Options

Configure behavior at runtime or via API:

Option Type Default Description
atomic bool true Use atomic file writes
indent int 2 JSON/YAML indentation
sort_keys bool false Sort dictionary keys
encoding str "utf-8" File encoding
ensure_ascii bool false Escape non-ASCII in JSON
allow_unicode bool true Allow Unicode in YAML
xml_pretty bool true Pretty-print XML
array_strategy str "horizontal" CSV array handling

Usage in Library:

# Pass as keyword arguments
data = load("file.json", encoding="utf-16", indent=4)
save(data, "out.json", sort_keys=True, ensure_ascii=False)

Usage in CLI:

DataConv> set indent 4
[+] Set indent = 4

DataConv> set sort_keys true
[+] Set sort_keys = true

DataConv> options
Current Configuration:
  atomic: true
  indent: 4
  sort_keys: true
  ...

Supported Formats

Format Read Write Notes
JSON Fast with orjson (3x)
YAML Optional C extensions (5x)
TOML stdlib tomllib on Python 3.11+
XML Via xmltodict
CSV Nested structure support

Testing

Run the test suite:

# Run all tests
python test_runner.py

# With pytest
pytest tests/

# With coverage
pytest --cov=src --cov-report=term-missing

Test Coverage:

  • 226 tests passing
  • 100% API coverage
  • Cross-platform compatibility (Windows, macOS, Linux)

Project Structure

DataConverter/
├── dataconv/
│   ├── api.py           # Public API functions
│   ├── cli.py           # Interactive REPL
│   ├── grammar.py       # Query language grammar
│   ├── parser.py        # Query parser
│   ├── processor.py     # Data processing engine
│   ├── io.py            # File I/O operations
│   ├── validation.py    # Format validators
│   └── options.py       # Configuration management
├── tests/
│   ├── api_test.py
│   ├── cli_test.py
│   ├── grammar_test.py
│   ├── parser_test.py
│   ├── processor_test.py
│   ├── io_test.py
│   └── validation_test.py
├── main.py              # CLI entry point
├── test_runner.py       # Test suite runner
├── pyproject.toml       # Project configuration
├── README.md
├── CHANGELOG.md
└── LICENSE

Development

Setup

# Clone repository
git clone https://github.com/thaisya/dataconv.git
cd dataconv

# Install with dev dependencies
pip install -e ".[full]"

# Run tests
python test_runner.py

# Format code
black src/ tests/

# Lint
ruff check src/ tests/

# Type check
mypy src/

Contributing

Contributions are welcome! Please:

  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/amazing-feature)
  3. Make your changes
  4. Run tests (python test_runner.py)
  5. Commit your changes (git commit -m 'Add amazing feature')
  6. Push to the branch (git push origin feature/amazing-feature)
  7. Open a Pull Request

Changelog

See CHANGELOG.md for version history and migration guides.


License

This project is licensed under the MIT License - see the LICENSE file for details.


Acknowledgments


Made with ❤️ by thaisya

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

dataconv-1.0.2.tar.gz (35.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

dataconv-1.0.2-py3-none-any.whl (33.5 kB view details)

Uploaded Python 3

File details

Details for the file dataconv-1.0.2.tar.gz.

File metadata

  • Download URL: dataconv-1.0.2.tar.gz
  • Upload date:
  • Size: 35.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.2

File hashes

Hashes for dataconv-1.0.2.tar.gz
Algorithm Hash digest
SHA256 4434d174d4a635e4892845ecd673dc96a5dc17db94b423a22ebeab281b1ac452
MD5 5de4a58eff1a01770cd1b1127b5da5a4
BLAKE2b-256 3278c8764c2de95c085add2e910c3dc0b994de6e56e9b58e1e6b45ce2b8ed9df

See more details on using hashes here.

File details

Details for the file dataconv-1.0.2-py3-none-any.whl.

File metadata

  • Download URL: dataconv-1.0.2-py3-none-any.whl
  • Upload date:
  • Size: 33.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.2

File hashes

Hashes for dataconv-1.0.2-py3-none-any.whl
Algorithm Hash digest
SHA256 9ff04d75c943eb9fbec52238a54f035adcbe6e30d32f3de1369e854a4dddf359
MD5 008c62dcba8d5c89fe43777934fe6e56
BLAKE2b-256 9fb4a39368bb999987dd38beb268a3b6747aab24b2bd4f4dfd7969f8d5339079

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page