Skip to main content

MX8 File system

This library provides environment agnostic file system access across local and AWS, including:

  • File / IO
  • List / Glob
  • Locking
  • Caching
  • Comparing Dictionaries

API Reference

Below are the functions and classes exported by the library (mx8fs.__all__). Paths can be local (e.g. /tmp/file.txt) or S3 (s3://bucket/key). HTTPS reads are supported in specific helpers.

File I/O and Paths

  • read_file(path: str) -> str: Read text from S3, HTTPS, or local using UTF‑8.

    • S3: GetObject and decode body.
    • HTTPS: simple GET; raises FileNotFoundError for network or 404.
    • Local: standard open/read.

    Example:

    from mx8fs import read_file
    
    text = read_file('/tmp/example.txt')
    s3_text = read_file('s3://my-bucket/path/data.txt')
    https_text = read_file('https://example.com/info.txt')
    
  • read_file_with_version(path: str) -> tuple[str, str]: Read text and a version identifier.

    • S3: returns (content, etag) (quotes stripped).
    • Local: returns (content, mtime-as-string).

    Example:

    from mx8fs import read_file_with_version
    
    content, version = read_file_with_version('s3://my-bucket/app/config.json')
    
  • write_file(path: str, data: str) -> None: Write UTF‑8 text to S3 or local. Creates parent directories for local writes.

    Example:

    from mx8fs import write_file
    
    write_file('/tmp/output.txt', 'hello world')
    write_file('s3://my-bucket/logs/run.txt', 'done')
    
  • update_file_if_version_matches(path: str, data: str, version: str) -> None: Conditional write.

    • S3: uses IfMatch=<etag>. Raises VersionMismatchError on mismatch, FileNotFoundError if key missing.
    • Local: acquires a FileLock, compares mtime string; raises VersionMismatchError or FileNotFoundError similarly.

    Example (optimistic concurrency):

    from mx8fs import read_file_with_version, update_file_if_version_matches, VersionMismatchError
    
    path = 's3://my-bucket/app/config.json'
    data, ver = read_file_with_version(path)
    new_data = data + "\n" + "# updated"
    try:
        update_file_if_version_matches(path, new_data, ver)
    except VersionMismatchError:
        # somebody else updated it — reload & retry as needed
        pass
    
  • file_exists(path: str) -> bool: True if file exists (HEAD on S3; os.path.exists locally).

    Example:

    from mx8fs import file_exists
    if file_exists('s3://my-bucket/data.csv'):
        ...
    
  • delete_file(path: str) -> None: Delete from S3 or local. Local deletion ignores missing files.

    Example:

    from mx8fs import delete_file
    delete_file('/tmp/temp.txt')
    delete_file('s3://my-bucket/tmp/stale.json')
    
  • delete_files(paths: list[str], max_workers: int = 500) -> None: Batch delete mixed S3/local paths.

    • S3: groups by bucket; uses DeleteObjects in chunks of 1000.
    • Local: deletes concurrently with a thread pool.

    Example:

    from mx8fs import delete_files
    delete_files([
        '/tmp/a.txt',
        's3://my-bucket/path/old.log',
    ])
    
  • copy_file(src: str, dst: str, chunk_size: int = 131072) -> None: Copy between S3, local storage, and HTTPS sources with bounded memory. S3→S3 uses CopyObject; other S3 destinations use sequential multipart uploads, and local destinations are replaced atomically after a successful transfer.

    Example:

    from mx8fs import copy_file
    copy_file('/tmp/input.bin', '/tmp/output.bin')
    copy_file('s3://my-bucket/a.txt', 's3://my-bucket/archive/a.txt')
    
  • move_file(src: str, dst: str) -> None: Copy then delete source.

    Example:

    from mx8fs import move_file
    move_file('s3://my-bucket/new.txt', 's3://my-bucket/archive/new.txt')
    
  • get_files(root_path: str, prefix: str = "", cutoff_date: datetime | None = None, cutoff_earlier: bool = True) -> list[str]: List files under a root.

    • S3: returns keys relative to the root prefix; optional prefix filter and cutoff_date (defaults to older than) on LastModified.
    • Local: returns relative file paths under the directory; supports the same filters (mtime in UTC).

    Example:

    from datetime import datetime, timezone
    from mx8fs import get_files
    
    files = get_files('/data/reports/', prefix='2025-')
    old = get_files('s3://my-bucket/reports/', cutoff_date=datetime.now(timezone.utc))
    
  • list_files(root_path: str, file_type: str, prefix: str = "") -> list[str]: List files with a given extension (without dot).

    • Returns basenames without extension, optionally filtered by prefix.

    Example:

    from mx8fs import list_files
    names = list_files('/data', 'json', prefix='user_')  # ['user_1', 'user_2']
    
  • get_folders(root_path: str, prefix: str = "") -> list[str]: List immediate subfolders (non‑recursive).

    • S3: uses Delimiter='/' and returns top‑level folder names under the prefix.
    • Local: returns directory names directly under root_path.

    Example:

    from mx8fs import get_folders
    top = get_folders('s3://my-bucket/data/')  # e.g. ['a', 'b']
    
  • most_recent_timestamp(root_path: str, file_type: str) -> float: Latest modification time for files of a type.

    • S3: scans with a paginator and returns the newest LastModified as epoch seconds.
    • Local: max of os.path.getmtime for *.ext, or 0 when none.

    Example:

    from mx8fs import most_recent_timestamp
    ts = most_recent_timestamp('/data', 'csv')
    
  • get_public_url(path: str, expires_in: int = 3600, method: str = "get_object") -> str: For S3, returns a presigned URL for GET/PUT; for local, returns the input path.

    Example:

    from mx8fs import get_public_url
    get_url = get_public_url('s3://my-bucket/file.txt')
    put_url = get_public_url('s3://my-bucket/file.txt', method='put_object')
    
  • purge_folder(root_path: str, dry_run: bool = True, max_workers: int = 500, cutoff_date: datetime | None = None) -> list[str]: List and optionally delete all files under a folder/prefix.

    • Returns the full paths that were (or would be) deleted; respects cutoff_date when provided.

    Example:

    from mx8fs import purge_folder
    paths = purge_folder('s3://my-bucket/tmp/', dry_run=True)
    deleted = purge_folder('/tmp/workdir', dry_run=False)
    

Stream Handlers

  • class BinaryFileHandler(path: str, mode: str = "rb", content_type: str | None = None): Context manager for binary reads/writes.

    • S3: rb downloads to an in‑memory buffer; wb uploads on exit (uses content_type if provided).
    • HTTPS: supports rb read‑only; writing or text modes raise NotImplementedError.
    • Local: proxies to built‑in file object; ensures directories exist for writes.
    • Usage:
    from mx8fs import BinaryFileHandler
    with BinaryFileHandler('s3://bucket/key.bin', 'wb', content_type='application/octet-stream') as f:
        f.write(b'data')
    
  • GzipFileHandler(path: str, mode: str = "rb", encoding: str | None = None): Context manager for gzip‑compressed files.

    • Modes: rb, wb, rt, wt (text modes require encoding). Wraps BinaryFileHandler under the hood.

    Example:

    from mx8fs import GzipFileHandler
    # write text
    with GzipFileHandler('/tmp/data.gz', 'wt', encoding='utf-8') as f:
        f.write('hello')
    # read binary
    with GzipFileHandler('/tmp/data.gz', 'rb') as f:
        blob = f.read()
    

Locking

  • class Waiter(wait_period: float, time_out_seconds: float): Helper for periodic waits with a timeout.

    • Methods: start_timeout(), check_timeout(), wait(), timed_out(); usable as a context manager.

    Example:

    from mx8fs import Waiter
    with Waiter(0.1, 5) as w:
        while not some_condition():
            w.check_timeout()
    
  • class FileLock(file: str, wait_period: float = 0.1, time_out_seconds: int = 840, maximum_age: int = 900): Cross‑process/client lock via lock files.

    • Creates lock files named {file}.{timestamp}.{random}.lock and coordinates access; cleans up on exit.
    • Use a higher wait_period for eventually‑consistent backends.
    • Usage:
    from mx8fs import FileLock, write_file
    with FileLock('s3://bucket/data.txt'):
        write_file('s3://bucket/data.txt', 'content')
    

Caching

  • get_cache_filename(path: str, name: str, extension: str, expiration_seconds: int = 0, **kwargs) -> str: Build a deterministic, optionally time‑scoped cache filename based on args/kwargs.

    Example:

    from mx8fs import get_cache_filename
    fname = get_cache_filename('/tmp/cache', 'fetch_users', 'txt', expiration_seconds=3600, region='us', extra_args=(1,2))
    
  • cache_to_disk(path: str, expiration_seconds: int = 0, log_group: str = "", ignore_kwargs: list[str] | None = None): Decorator for caching text results.

    • Stores .txt payloads using read_file/write_file; logs cache hits when log_group is set.

    Example:

    from mx8fs import cache_to_disk
    
    @cache_to_disk('/tmp/mycache', expiration_seconds=600)
    def compute(user_id: str) -> str:
        return f"hello {user_id}"
    
    print(compute('42'))
    
  • cache_to_disk_binary(path: str, expiration_seconds: int = 0, log_group: str = "", ignore_kwargs: list[str] | None = None): Decorator for caching arbitrary Python objects via pickle.

    • Stores .pickle payloads using BinaryFileHandler.

    Example:

    from mx8fs import cache_to_disk_binary
    
    @cache_to_disk_binary('/tmp/bin_cache')
    def heavy() -> dict:
        return {'a': 1}
    data = heavy()
    

JSON Storage

Install JSON storage dependencies with pip install "mx8fs[json-storage]".

  • class JsonFileStorage(base_path: str, randomizer: Callable | None = None): Base class for simple JSON model storage.

    • Methods: list(), read(key), read_many(keys, max_workers=None), write(model), write_dict(dict, key=None), update(model), delete(key), get_lock(key, wait_period=0.1, time_out_seconds=840, maximum_age=900).
    • Implements unique key generation and defers serialization to subclass hooks.

    Example (using factory below for a Pydantic model): see next section.

  • json_file_storage_factory(extension: str, model: Any, key_field: str = "key") -> type[JsonFileStorage]: Generates a concrete JsonFileStorage for a Pydantic model.

    • The resulting class knows how to serialize/deserialize the given model and stores files as <key>.<extension>.

    Example:

    from pydantic import BaseModel
    from mx8fs import json_file_storage_factory
    
    class User(BaseModel):
        key: str | None = None
        name: str
    
    UserStorage = json_file_storage_factory('json', User, key_field='key')
    store = UserStorage('/tmp/users')
    
    u = store.write(User(name='Ada'))     # auto-keyed
    with store.get_lock(u.key):
        u = store.update(User(key=u.key, name='Ada Lovelace'))
    got = store.read(u.key)
    all_keys = store.list()
    store.delete(u.key)
    

Indexed JSON Storage

json_index_factory adds a SQL secondary index while keeping JSON files canonical. Local storage uses a hidden SQLite database at <base_path>/.mx8fs-index.sqlite3; indexed S3 storage uses Aurora DSQL.

Use pip install "mx8fs[indexed-json-storage]" for indexed storage. It uses SQLite locally and Aurora DSQL for S3 storage.

from datetime import datetime
from pydantic import BaseModel
from mx8fs import json_file_storage_factory, json_index_factory

class Job(BaseModel):
    key: str | None = None
    status: str
    created_at: datetime
    not_before: datetime | None = None

JobIndex = json_index_factory(
    Job,
    fields=["status", "created_at", "not_before"],
    table_name="jobs",  # optional; otherwise derived from the model
)

JobStorage = json_file_storage_factory("json", Job, index=JobIndex)
jobs = JobStorage("/tmp/jobs")

query = (
    jobs.query()
    .where(jobs.index.status == "pending")
    .order_by(jobs.index.created_at.desc())
    .page(1, 50)
)

rows = query.all()       # SQLAlchemy RowMapping objects: key + indexed fields
keys = query.keys()      # ordered keys only
total = query.count()    # count before limit/offset
models = query.models()  # concurrently load and validate the JSON models

Every index definition is fingerprinted from the selected Pydantic fields. Missing or incompatible index tables are logged, recreated under a shared SQL lease, and rebuilt automatically from the canonical JSON files. Different storage paths using the same definition share a physical table but remain isolated by an internal namespace. Older fingerprinted table versions are retained for compatibility with older deployments.

Indexed S3 storage requires MX8FS_DSQL_ENDPOINT. MX8FS_DSQL_USER defaults to admin, and MX8FS_DSQL_DATABASE defaults to postgres. AWS credentials and Region resolution use the normal AWS chain.

rebuild_index() can be called to force a full reconciliation. Indexed mutations already acquire re-entrant per-key file locks, so an explicit read-modify-write lock remains safe:

with jobs.get_lock("job-key"):
    job = jobs.read("job-key")
    jobs.update(job.model_copy(update={"status": "complete"}))

Comparison Utilities

  • class ResultsComparer(ignore_keys: list[str] | None, create_test_data: bool = False, obfuscate_regex: str | None = None): Compare text/JSON with optional obfuscation and test‑data generation.

    • Methods: compare_dicts, get_text_differences, get_dict_differences, get_api_response_differences.
    • Obfuscates sensitive content by default (passwords, tokens, keys, etc.).

    Example:

    from mx8fs import ResultsComparer
    
    rc = ResultsComparer(ignore_keys=['timestamp'])
    diffs = rc.compare_dicts({'a': 1, 'timestamp': 1}, {'a': 2, 'timestamp': 2})
    if diffs:
        print(diffs)
    

Exceptions

  • class VersionMismatchError(FileNotFoundError): Raised by update_file_if_version_matches on conditional write mismatches.

    Example (catching):

    from mx8fs import VersionMismatchError
    try:
        raise VersionMismatchError('example')
    except VersionMismatchError:
        pass
    

Notes

  • S3 settings: the client uses environment variables BOTO_MAX_CONNECTIONS, BOTO_CONNECT_TIMEOUT, BOTO_READ_TIMEOUT, BOTO_MAX_RETRIES, and BOTO_RETRY_MODE to tune behavior.
  • Time handling: cutoff_date is normalized to timezone‑aware UTC for consistent comparisons.
  • HTTPS reads: read_file and BinaryFileHandler (rb) support https:// sources; writing to HTTPS is not supported.

Pre-commit hooks

We use precommit to run formatting checks, so whenever you clone a project run:

pre-commit install

Before you do anything else.

You can run this at any time using:

pre-commit run --all-files

Setting up the development environment

You can install the full dev requirements by running setup.sh to

  1. Install the current repo and the python lib
  2. Run the pre-commit hooks on all files

The project should open reasonable well in vs.code and includes three launch configurations for running unit tests and the debug server.

Code conventions and structure

We use python type hinting with pylance and flake8 for linting. Unit tests are created to 100% branch coverage.

The code is structured as follows:

  • The mx8fs folder contains the full library.
  • Tests are stored in the tests folder and run using pytest.

License

Copyright © 2025 MX8 Labs

Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the “Software”), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED “AS IS”, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.

Development setup

Run ./setup.sh to install dependencies and configure the environment.

Tooling

Use black, ruff, and mypy for formatting/linting/type checking. Run pytest for tests.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mx8fs-1.1.1.35.tar.gz (37.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

mx8fs-1.1.1.35-py3-none-any.whl (37.9 kB view details)

Uploaded Python 3

File details

Details for the file mx8fs-1.1.1.35.tar.gz.

File metadata

  • Download URL: mx8fs-1.1.1.35.tar.gz
  • Upload date:
  • Size: 37.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for mx8fs-1.1.1.35.tar.gz
Algorithm Hash digest
SHA256 6f07be2e285bc86637a35f3e818914c3b7e12eebc2f2557b71a5b85419500760
MD5 252409883432c4116b935187d61f7bf5
BLAKE2b-256 c667d633e5c14ff0c2340324824a9a1eef069eeb2dccfc6426caa9ce2e1aefc0

See more details on using hashes here.

Provenance

The following attestation bundles were made for mx8fs-1.1.1.35.tar.gz:

Publisher: deploy_pypi.yml on mx8-io/mx8fs

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mx8fs-1.1.1.35-py3-none-any.whl.

File metadata

  • Download URL: mx8fs-1.1.1.35-py3-none-any.whl
  • Upload date:
  • Size: 37.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for mx8fs-1.1.1.35-py3-none-any.whl
Algorithm Hash digest
SHA256 f3ca7390bf67bcad7b10bb331ca4f84820167ec931194fd479b7fa4fb2c31dbb
MD5 ae1852b665ce06312f135d58b5f48c5d
BLAKE2b-256 dc269415781ba1432db9ccff2dc6a63438ddd2e132b564dad6bf2e83f6ed9669

See more details on using hashes here.

Provenance

The following attestation bundles were made for mx8fs-1.1.1.35-py3-none-any.whl:

Publisher: deploy_pypi.yml on mx8-io/mx8fs

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

1.1.1.42

2 files

1.1.1.41

2 files

1.1.1.40

2 files

1.1.1.39

2 files

1.1.1.38

2 files

1.1.1.37

2 files

1.1.1.36

2 files

This release

1.1.1.35 This release

2 files

1.1.1.34

2 files

1.1.0.33

2 files

1.1.0.32

2 files

1.1.0.31

2 files

1.1.0.30

2 files

1.1.0.29

2 files

1.1.0.28

2 files

1.1.0.27

2 files

1.1.0.26

2 files

1.1.0.25

2 files

1.1.0.24

2 files

1.1.0.23

2 files

1.1.0.22

2 files

1.1.0.21

2 files

1.1.0.20

2 files

1.1.0.19

2 files

1.1.0.18

2 files

1.1.0.17

2 files

1.1.0.16

2 files

1.1.0.15

2 files

1.1.0.14

2 files

1.1.0.13

2 files

1.1.0.12

2 files

1.1.0.11

2 files

1.1.0.10

2 files

1.1.0.9

2 files

1.1.0.8

2 files

1.1.0.7

2 files

1.1.0.6

2 files

1.1.0.5

2 files

1.1.0.4

2 files

1.0.0.12

2 files

1.0.0.8

2 files

1.0.0.7

2 files

1.0.0.6

2 files

1.0.0.2

2 files

1.0.0.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page