Skip to main content

fileslicer

PyPI - Version PyPI - Python Version pre-commit.ci status


fileslicer is a lightweight Python library for efficiently reading and splitting large files using memory mapping. It allows you to iterate over lines within a file slice and split files into chunks without loading the entire file into memory, making it ideal for processing very large files.


Features

  • Memory-efficient line iteration using mmap.
  • Split large files into chunks while respecting newline boundaries.
  • Simple and Pythonic API.
  • Works with files of arbitrary size.

Installation

Install via pip:

pip install fileslicer

Usage

Basic Example: Iterate over a file

from fileslicer import FileSlice

# Create a FileSlice for an entire file
slice = FileSlice.from_file("large_file.txt")

# Iterate over lines in the slice
for line in slice.iter_lines():
    print(line.decode().strip())

Split a File into Chunks

from fileslicer import FileSlice

# Split a file into 4 chunks
chunks = FileSlice.split_file("large_file.txt", splits=4)

for chunk in chunks:
    print(f"Processing bytes {chunk.start_offset}-{chunk.end_offset}")
    for line in chunk.iter_lines():
        print(line.decode().strip())

Create a Custom File Slice

from fileslicer import FileSlice

# Only read bytes 1000 to 5000
slice = FileSlice("large_file.txt", 1000, 5000)

for line in slice.iter_lines():
    print(line.decode().strip())

API

FileSlice

  • FileSlice(file_path: str, start_offset: int, end_offset: int): Represents a slice of a file.

  • iter_lines() -> Generator[bytes]: Iterate over lines in the file slice as bytes.

  • @staticmethod from_file(file_path: str) -> FileSlice: Create a FileSlice covering the entire file.

  • @staticmethod split_file(file_path: str, splits: int) -> list[FileSlice]: Split a file into multiple slices, aligned to newline boundaries.


Why Use fileslicer?

Processing extremely large files with standard file reading can be slow and memory-intensive. fileslicer uses memory mapping to efficiently slice and iterate over file data without reading everything into memory. Inspired by the "1 Billion Row Challenge" in Python, it is perfect for data processing pipelines, log analysis, and ETL tasks.


License

fileslicer is distributed under the terms of the MIT license.

Metadata

Release files for fileslicer 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for fileslicer 0.1.0
File Size Uploaded
fileslicer-0.1.0.tar.gz 10.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for fileslicer 0.1.0
File Interpreter ABI Platform
fileslicer-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 16.5 kB

Release files / fileslicer-0.1.0.tar.gz

Download URL fileslicer-0.1.0.tar.gz
Size 10.3 kB
Tags Source
SHA-256 checksum
How to use checksums
6198f736662b49f8273b163f7140052f142924a7f273d4bb29c80ae81910c1a8
BLAKE2b-256 checksum
How to use checksums
f68b6a658ce67f71058141be2cff7e46f52fe26dfd160adc3b9e880b4c1c33f2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 21, 2025.

Transparency log

Release files / fileslicer-0.1.0-py3-none-any.whl

Download URL fileslicer-0.1.0-py3-none-any.whl
Size 6.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
338c5448387c879b56862f458426b5754f98341d6583348ca6e86b7a0b9288db
BLAKE2b-256 checksum
How to use checksums
510a8fa4cd80a22333e579ff56ae66b9e948efeaee21a53fcd86e5e505c72906
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 21, 2025.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page