Skip to main content

watfile

Classify files with a decision-making AI model and sort them into category folders.

watfile sends each document's text (title/abstract-grade extract) to a TypeSafe AI Jev (System One) Choice question, gets back a typed answer with a selected category, per-category probabilities and confidence, then moves the file into the matching folder. A Classifier abstraction keeps the backend pluggable — local MLX (laya) and other backends slot in later.

Install

Requires Python 3.12+ and uv.

From PyPI

# one-off run, no install
uvx watfile --help

# persistent CLI on your PATH
uv tool install watfile
watfile --help

Updates:

uv tool upgrade watfile      # or: watfile --self-update
uvx watfile@latest ...       # one-off runs always fetch the newest version

Configuration

watfile resolves its TypeSafe API key (create one at https://console.typesafe.ai/) with this precedence — first match wins:

  1. TYPESAFE_API_KEY environment variable
  2. .env file in the current directory (gitignored; TYPESAFE_API_KEY=...)
  3. ~/.config/watfile/config.toml (api_key = "...", also base_url, model; $WATFILE_CONFIG or $XDG_CONFIG_HOME can relocate it)
export TYPESAFE_API_KEY=...        # option 1
echo 'TYPESAFE_API_KEY=...' > .env # option 2
cat > ~/.config/watfile/config.toml <<'EOF'   # option 3
api_key = "..."
EOF

Usage

Point watfile at files or folders, and either give a comma-separated category list (-c) or a target folder whose subfolders are the categories (-d):

# explicit categories, files moved into ./sorted/<category>/
uv run watfile ~/Downloads/invoice.pdf -c invoice,donation,apartment

# folder input, recursive; categories = existing subfolders of -d
mkdir -p ~/docs/{invoice,donation,apartment}
uv run watfile ~/Downloads -r -d ~/docs

# preview without touching anything
uv run watfile ~/Downloads -r -d ~/docs -n

# actually move the files (default is symlinking into the category folders)
uv run watfile ~/Downloads -r -d ~/docs -m

# copy instead
uv run watfile ~/Downloads -r -d ~/docs --copy

# custom output root with -c
uv run watfile *.pdf -c computerscience,biology -o ~/sorted

Output per file:

bill.pdf: invoice (conf 0.94) -> symlink to ~/docs/invoice/bill.pdf

Files that can't be classified (unsupported extension, no extractable text) are skipped with a warning; name collisions get a _1, _2… suffix.

Supported inputs

  • Text formats (read directly): .txt .md .markdown .rst .log .csv .json
  • PDF (via liteparse): only the first 2 pages are parsed, OCR disabled — enough for classification, ~1000x faster than a full parse. Scanned/image-only PDFs are skipped.

Backends

  • jev (default)TypeSafe AI Jev, cloud API. Needs TYPESAFE_API_KEY. Highest accuracy (4/4 on the arXiv fixtures). Batches automatically: documents are packed into one system_one call (~256 tokens each, up to ~100 files per call in the 30k-token window), so classifying a folder costs one API call, not one per file.

  • laya — local typed-decision model from the Laya family; runs on any platform. The runtime is auto-detected from what's installed (LAYA_RUNTIME overrides):

    extra runtime where speed
    pip install watfile[laya] MLX (GPU) Apple Silicon ~13ms/decision
    pip install watfile[laya-coreml] Core ML (ANE) Apple Silicon ~5ms, 2.8× lower energy
    pip install watfile[laya-torch] PyTorch (CPU/GPU) any OS ~45–450ms (CPU)

    No API key needed; checkpoints download once and then run offline. Default checkpoints: multilingual where available (torch/coreml → handles non-English documents out of the box). Because laya's context window is small (512–1024 tokens), the backend defaults to adaptive multi-chunk classification: the extract is split into ~200-token chunks; chunk 1 decides if its probability is decisive (≥0.5), otherwise further chunks are classified and probabilities aggregated until the decision is decisive (max 10). On the arXiv fixtures: 3/4 (Jev 4/4).

watfile ~/Downloads -r -d ~/docs --backend laya
# checkpoint via config or env:
#   ~/.config/watfile/config.toml -> laya_model = "..."
#   or LAYA_MODEL=... / LAYA_RUNTIME=torch|mlx|coreml

Options (main)

watfile --help shows only these:

usage: watfile [-h] [-r] (-c CATEGORIES | -d DIRECTORY) [-o OUTPUT]
               [--backend {jev,laya}] [-n] [-m | --copy | --symlink]
               [--help-all]
               inputs [inputs ...]

main options:
  -h, --help            show this help message and exit
  --help-all            show advanced options too
  -r, --recursive       recurse into folder inputs
  -v, --version         print version and exit
  -c CATEGORIES         comma-separated categories
  -d DIRECTORY          target folder whose existing subfolders are the categories
  -o OUTPUT             output root for sorted files (default: same as -d, or ./sorted with -c)
  --backend {jev,laya}  classifier backend (default: jev)
  -n, --dry-run         print decisions without placing files
  -m, --move            move files into the category folder (default: symlink)
  --copy                copy files instead of symlinking
  --symlink             create symlinks in category folders (default)
  --self-update         update watfile in place (uv tool / pipx aware)

Advanced options

Shown by watfile --help-all:

  --batch N             cap files per API call (default: automatic — jev packs
                        everything that fits the 30k-token window, ~100 docs;
                        laya doesn't batch)
  --no-batch            disable batching, one API call per file (debugging)
  --chunk-tokens N      per-document token budget (default: backend-specific)
  --chunks N            split each document into N chunks, aggregate
                        probabilities; 0 = adaptive. Default: 0 for laya,
                        1 for jev

Batching example (jev batches by default; the flag just caps batch size):

# 100 files: ~2 API calls instead of 100 (256 tokens/doc, 30k window)
uv run watfile ~/Downloads -r -d ~/docs

# cap batch size, e.g. to keep batches small for debugging
uv run watfile ~/Downloads -r -d ~/docs --batch 25

Development

From source

git clone <repo> && cd watfile
uv sync            # create venv + install deps (typesafe-sdk, liteparse, laya-mlx)
uv run watfile --help

# or install the local checkout as a tool
uv tool install --from . watfile

Tests

uv sync
uv run pytest              # unit tests; live API tests skip without TYPESAFE_API_KEY

tests/fixture/ contains 4 real arXiv PDFs with ground-truth categories (derived from their arXiv subject tags) used by the integration tests.

Roadmap

  • laya local backend (MLX via OpenAI-compatible HTTP) done — native laya-mlx
  • batching: classify 25/50/100 files in a single API call done — jev batches automatically into the 30k-token window (default on, --no-batch to disable)

Release files for watfile 0.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for watfile 0.2.1
File Size Uploaded
watfile-0.2.1.tar.gz 16.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for watfile 0.2.1
File Interpreter ABI Platform
watfile-0.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 37.0 kB

Release files / watfile-0.2.1.tar.gz

Download URL watfile-0.2.1.tar.gz
Size 16.7 kB
Tags Source
SHA-256 checksum
How to use checksums
600df7d718f6f06002261880f7af2c8fc34723623fd2c1a7a2e74c5bdd89db23
BLAKE2b-256 checksum
How to use checksums
c88f2bdeb0d67fd9c4e4f97bb6102981d030b19358c3a8150ea2157430f73a65
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 20, 2026.

Transparency log

Release files / watfile-0.2.1-py3-none-any.whl

Download URL watfile-0.2.1-py3-none-any.whl
Size 20.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
87e68f0e4572d86f8594d869342e8e3da074b144c65324ae6589d3d9b9e98168
BLAKE2b-256 checksum
How to use checksums
a53f99df5646490cc8cfbc410d71d6f823fb2d4b5761a53e6452e8b5b2252bcb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 20, 2026.

Transparency log

Release history Release notifications | RSS feed

0.3.0

2 release files

This release

0.2.1 This release

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page