Skip to main content

A utility library for working with DSV (Delimited String Values) files

Project description

splurge-dsv

PyPI version Python versions License: MIT

CI Coverage Ruff mypy

A robust Python library for parsing and processing delimited-separated value (DSV) files with advanced features for data validation, streaming, and error handling.

Features

  • Multi-format DSV Support: Parse CSV, TSV, pipe-delimited, and custom delimiter separated value files/objects
  • Configurable Parsing: Flexible options for delimiters, quote characters, escape characters, header/footer row(s) handling
  • Memory-Efficient Streaming: Process large files without loading entire content into memory
  • Security & Validation: Comprehensive path validation and file permission checks
  • Unicode Support: Full Unicode character and encoding support
  • Type Safety: Full type annotations with mypy validation
  • Deterministic Newline Handling: Consistent handling of CRLF, CR, and LF newlines across platforms
  • CLI Tool: Command-line interface for quick parsing and inspection of DSV files
  • Robust Error Handling: Clear and specific exceptions for various error scenarios
  • Modern API: Object-oriented API with Dsv and DsvConfig classes for easy configuration and reuse
  • Event Publishing: Integrated event publishing using splurge-pub-sub for monitoring parsing progress and events
  • Comprehensive Documentation: In-depth API reference and usage examples
  • Exhaustive Testing: 297 tests with 87% code coverage including property-based testing, edge case testing, and cross-platform compatibility validation

Installation

pip install splurge-dsv

Quick Start

CLI Usage

# Parse a CSV file
python -m splurge_dsv data.csv --delimiter ,

# Stream a large file
python -m splurge_dsv large_file.csv --delimiter , --stream --chunk-size 1000

YAML configuration file

You can place CLI-equivalent options in a YAML file and pass it to the CLI using --config (or -c). CLI arguments override values found in the YAML file. Example config.yaml:

delimiter: ","
strip: true
bookend: '"'
encoding: utf-8
skip_header_rows: 1
skip_footer_rows: 0
skip_empty_lines: false
detect_columns: true
chunk_size: 500
max_detect_chunks: 5
raise_on_missing_columns: false
raise_on_extra_columns: false

Usage with CLI:

python -m splurge_dsv data.csv --config config.yaml --delimiter "|"
# The CLI delimiter '|' overrides the YAML delimiter

Example using the shipped example config in the repository:

# Use the example file provided at examples/config.yaml
python -m splurge_dsv data.csv --config examples/config.yaml

API Usage

from splurge_dsv import DsvHelper

# Parse a CSV string
data = DsvHelper.parse("a,b,c", delimiter=",")
print(data)  # ['a', 'b', 'c']

# Parse a CSV file
rows = DsvHelper.parse_file("data.csv", delimiter=",")

Modern API

from splurge_dsv import Dsv, DsvConfig

# Create configuration and parser
config = DsvConfig.csv(skip_header=1)
dsv = Dsv(config)

# Parse files
rows = dsv.parse_file("data.csv")

Documentation

License

This project is licensed under the MIT License - see the LICENSE file for details.

This library enforces deterministic newline handling for text files. The reader normalizes CRLF (\r\n), CR (\r) and LF (\n) to LF internally and returns logical lines. The writer utilities normalize any input newlines to LF before writing. This avoids platform-dependent differences when reading files produced by diverse sources.

Recommended usage:

  • When creating files inside the project, prefer the open_text_writer context manager or SafeTextFileWriter which will normalize to LF.
  • When reading unknown files, the open_text / SafeTextFileReader will provide deterministic normalization regardless of the source.

Development

Testing Suite

splurge-dsv features a comprehensive testing suite designed for robustness and reliability:

Test Categories

  • Unit Tests: Core functionality testing (300+ tests)
  • Integration Tests: End-to-end workflow validation (50+ tests)
  • Property-Based Tests: Hypothesis-driven testing for edge cases (50+ tests)
  • Edge Case Tests: Malformed input, encoding issues, filesystem anomalies
  • Cross-Platform Tests: Path handling, line endings, encoding consistency

Running Tests

# Run all tests
pytest tests/ -v

# Run with coverage report
pytest tests/ --cov=splurge_dsv --cov-report=html

# Run specific test categories
pytest tests/unit/ -v                    # Unit tests only
pytest tests/integration/ -v            # Integration tests only
pytest tests/property/ -v               # Property-based tests only
pytest tests/platform/ -v               # Cross-platform tests only

# Run with parallel execution
pytest tests/ -n 4 --cov=splurge_dsv

# Run performance benchmarks
pytest tests/ --durations=10

Test Quality Standards

  • 87%+ Code Coverage: All public APIs and critical paths covered
  • Property-Based Testing: Hypothesis framework validates complex scenarios
  • Cross-Platform Compatibility: Tests run on Windows, Linux, and macOS
  • Performance Regression Detection: Automated benchmarks prevent slowdowns
  • Zero False Positives: All property tests pass without spurious failures

Testing Best Practices

  • Tests use pytest-mock for modern mocking patterns
  • Property tests use Hypothesis strategies for comprehensive input generation
  • Edge case tests validate error handling and boundary conditions
  • Cross-platform tests ensure consistent behavior across operating systems

Code Quality

The project follows strict coding standards:

  • PEP 8 compliance
  • Type annotations for all functions
  • Google-style docstrings
  • 85%+ coverage gate enforced via CI
  • Comprehensive error handling

Changelog

See the CHANGELOG for full release notes.

License

This project is licensed under the MIT License - see the LICENSE file for details.

More Documentation

Contributing

Contributions are welcome! Please see our Contributing Guide for detailed information on:

  • Development setup and workflow
  • Coding standards and best practices
  • Testing requirements and guidelines
  • Pull request process and review criteria

For major changes, please open an issue first to discuss what you would like to change.

Support

For support, please open an issue on the GitHub repository or contact the maintainers.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

splurge_dsv-2025.6.0.tar.gz (81.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

splurge_dsv-2025.6.0-py3-none-any.whl (101.7 kB view details)

Uploaded Python 3

File details

Details for the file splurge_dsv-2025.6.0.tar.gz.

File metadata

  • Download URL: splurge_dsv-2025.6.0.tar.gz
  • Upload date:
  • Size: 81.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.9

File hashes

Hashes for splurge_dsv-2025.6.0.tar.gz
Algorithm Hash digest
SHA256 cba6e1a8a4e1b802f9695a43d5670146527abfffd58f94d0875c27908e83edbf
MD5 ad5514bca344796520940826f40aa45c
BLAKE2b-256 c399d9e0f8cfcfb9337966c2369284854b59700e955da6a19100ac1780658f47

See more details on using hashes here.

File details

Details for the file splurge_dsv-2025.6.0-py3-none-any.whl.

File metadata

  • Download URL: splurge_dsv-2025.6.0-py3-none-any.whl
  • Upload date:
  • Size: 101.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.9

File hashes

Hashes for splurge_dsv-2025.6.0-py3-none-any.whl
Algorithm Hash digest
SHA256 9c9ebc80b89ae23fa8f0afc2ee463b1922931306916e42f82a85868edee420a6
MD5 b912b923864e5e662dd20ca280c7f7f4
BLAKE2b-256 ba6f4f31dce36b8e7201c754f52f8f5b7009bbdf39cb8430d17d513298da43bf

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page