A utility library for working with DSV (Delimited String Values) files
Project description
splurge-dsv
A robust Python library for parsing and processing delimited-separated value (DSV) files with advanced features for data validation, streaming, and error handling.
Features
- Multi-format DSV Support: Parse CSV, TSV, pipe-delimited, and custom delimiter separated value files/objects
- Configurable Parsing: Flexible options for delimiters, quote characters, escape characters, header/footer row(s) handling
- Memory-Efficient Streaming: Process large files without loading entire content into memory
- Security & Validation: Comprehensive path validation and file permission checks
- Unicode Support: Full Unicode character and encoding support
- Type Safety: Full type annotations with mypy validation
- Deterministic Newline Handling: Consistent handling of CRLF, CR, and LF newlines across platforms
- CLI Tool: Command-line interface for quick parsing and inspection of DSV files
- Robust Error Handling: Clear and specific exceptions for various error scenarios
- Modern API: Object-oriented API with
DsvandDsvConfigclasses for easy configuration and reuse - Event Publishing: Integrated event publishing using
splurge-pub-subfor monitoring parsing progress and events - Comprehensive Documentation: In-depth API reference and usage examples
- Exhaustive Testing: 297 tests with 87% code coverage including property-based testing, edge case testing, and cross-platform compatibility validation
Installation
pip install splurge-dsv
Quick Start
CLI Usage
# Parse a CSV file
python -m splurge_dsv data.csv --delimiter ,
# Stream a large file
python -m splurge_dsv large_file.csv --delimiter , --stream --chunk-size 1000
YAML configuration file
You can place CLI-equivalent options in a YAML file and pass it to the CLI
using --config (or -c). CLI arguments override values found in the
YAML file. Example config.yaml:
delimiter: ","
strip: true
bookend: '"'
encoding: utf-8
skip_header_rows: 1
skip_footer_rows: 0
skip_empty_lines: false
detect_columns: true
chunk_size: 500
max_detect_chunks: 5
raise_on_missing_columns: false
raise_on_extra_columns: false
Usage with CLI:
python -m splurge_dsv data.csv --config config.yaml --delimiter "|"
# The CLI delimiter '|' overrides the YAML delimiter
Example using the shipped example config in the repository:
# Use the example file provided at examples/config.yaml
python -m splurge_dsv data.csv --config examples/config.yaml
API Usage
from splurge_dsv import DsvHelper
# Parse a CSV string
data = DsvHelper.parse("a,b,c", delimiter=",")
print(data) # ['a', 'b', 'c']
# Parse a CSV file
rows = DsvHelper.parse_file("data.csv", delimiter=",")
Modern API
from splurge_dsv import Dsv, DsvConfig
# Create configuration and parser
config = DsvConfig.csv(skip_header=1)
dsv = Dsv(config)
# Parse files
rows = dsv.parse_file("data.csv")
Documentation
- Detailed Documentation: Complete API reference, CLI options, and examples
- API Reference: In-depth documentation of classes and methods
- Testing Best Practices: Comprehensive testing guidelines and patterns
- Hypothesis Usage Patterns: Property-based testing guide
- Changelog: Release notes and migration guides
License
This project is licensed under the MIT License - see the LICENSE file for details.
This library enforces deterministic newline handling for text files. The reader
normalizes CRLF (\r\n), CR (\r) and LF (\n) to LF internally and
returns logical lines. The writer utilities normalize any input newlines to LF
before writing. This avoids platform-dependent differences when reading files
produced by diverse sources.
Recommended usage:
- When creating files inside the project, prefer the
open_text_writercontext manager orSafeTextFileWriterwhich will normalize to LF. - When reading unknown files, the
open_text/SafeTextFileReaderwill provide deterministic normalization regardless of the source.
Development
Testing Suite
splurge-dsv features a comprehensive testing suite designed for robustness and reliability:
Test Categories
- Unit Tests: Core functionality testing (300+ tests)
- Integration Tests: End-to-end workflow validation (50+ tests)
- Property-Based Tests: Hypothesis-driven testing for edge cases (50+ tests)
- Edge Case Tests: Malformed input, encoding issues, filesystem anomalies
- Cross-Platform Tests: Path handling, line endings, encoding consistency
Running Tests
# Run all tests
pytest tests/ -v
# Run with coverage report
pytest tests/ --cov=splurge_dsv --cov-report=html
# Run specific test categories
pytest tests/unit/ -v # Unit tests only
pytest tests/integration/ -v # Integration tests only
pytest tests/property/ -v # Property-based tests only
pytest tests/platform/ -v # Cross-platform tests only
# Run with parallel execution
pytest tests/ -n 4 --cov=splurge_dsv
# Run performance benchmarks
pytest tests/ --durations=10
Test Quality Standards
- 87%+ Code Coverage: All public APIs and critical paths covered
- Property-Based Testing: Hypothesis framework validates complex scenarios
- Cross-Platform Compatibility: Tests run on Windows, Linux, and macOS
- Performance Regression Detection: Automated benchmarks prevent slowdowns
- Zero False Positives: All property tests pass without spurious failures
Testing Best Practices
- Tests use
pytest-mockfor modern mocking patterns - Property tests use Hypothesis strategies for comprehensive input generation
- Edge case tests validate error handling and boundary conditions
- Cross-platform tests ensure consistent behavior across operating systems
Code Quality
The project follows strict coding standards:
- PEP 8 compliance
- Type annotations for all functions
- Google-style docstrings
- 85%+ coverage gate enforced via CI
- Comprehensive error handling
Changelog
See the CHANGELOG for full release notes.
License
This project is licensed under the MIT License - see the LICENSE file for details.
More Documentation
- Detailed docs: docs/README-DETAILS.md
Contributing
Contributions are welcome! Please see our Contributing Guide for detailed information on:
- Development setup and workflow
- Coding standards and best practices
- Testing requirements and guidelines
- Pull request process and review criteria
For major changes, please open an issue first to discuss what you would like to change.
Support
For support, please open an issue on the GitHub repository or contact the maintainers.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file splurge_dsv-2025.6.0.tar.gz.
File metadata
- Download URL: splurge_dsv-2025.6.0.tar.gz
- Upload date:
- Size: 81.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.9
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
cba6e1a8a4e1b802f9695a43d5670146527abfffd58f94d0875c27908e83edbf
|
|
| MD5 |
ad5514bca344796520940826f40aa45c
|
|
| BLAKE2b-256 |
c399d9e0f8cfcfb9337966c2369284854b59700e955da6a19100ac1780658f47
|
File details
Details for the file splurge_dsv-2025.6.0-py3-none-any.whl.
File metadata
- Download URL: splurge_dsv-2025.6.0-py3-none-any.whl
- Upload date:
- Size: 101.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.9
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
9c9ebc80b89ae23fa8f0afc2ee463b1922931306916e42f82a85868edee420a6
|
|
| MD5 |
b912b923864e5e662dd20ca280c7f7f4
|
|
| BLAKE2b-256 |
ba6f4f31dce36b8e7201c754f52f8f5b7009bbdf39cb8430d17d513298da43bf
|