Skip to main content

SnipeX logo

SnipeX

Single-GPU, oNe-week Idea-to-Prototype EXecution.

PyPI version GitHub stars Python versions MIT license

SnipeX is a collection of small AI algorithm prototypes. Algorithm-specific dependencies are provided through optional package extras. Every algorithm prototype in SnipeX is designed to run on a single consumer GPU and reach a working result within one week.


⚡ Quick start

Installing the base package provides the snipex command, package version, and help for the available algorithm families without installing PyTorch:

pip install snipex
snipex --help

🧠 LLM prototype (V0.1.0)

The first prototype is a NanoChat-inspired decoder-only language model with RoPE, RMSNorm and QK norm, ReLU² MLPs, document-aware ClimbMix packing, AdamW with linear warmup and cosine decay, MPS/CUDA support, and CUDA DDP.

Install the LLM prototype from PyPI:

pip install "snipex[llm]"

From a source checkout, create the development environment with:

uv sync --extra llm

FA3 is optional and CUDA-only. auto falls back to PyTorch SDPA when the community kernel cannot be loaded. It is an implementation optimization, not an algorithm extra, so source checkouts install it through a dependency group:

uv sync --extra llm --group flash

PyPI users who explicitly want FA3 can install snipex[llm] and kernels separately. The normal snipex[llm] installation uses SDPA without kernels.

📚 Data and tokenizer

Keep every split in a separate directory. SnipeX never infers a split from a filename and never reads the test directory during tokenizer or model training:

/path/to/climbmix/
├── train/*.{parquet,jsonl}
├── val/*.{parquet,jsonl}
└── test/*.{parquet,jsonl}

Base-model JSONL uses one document per line:

{"text": "One pretraining document."}

The equivalent Parquet schema has one string column named text. Mixing Parquet and JSONL files inside one split is supported.

Train and inspect the 32K RustBPE tokenizer:

uv run snipex tok train \
  --train-data /path/to/climbmix/train \
  --out-dir artifacts/tokenizer

uv run snipex tok encode --tokenizer artifacts/tokenizer "Hello SnipeX"
uv run snipex tok decode --tokenizer artifacts/tokenizer 123 456
uv run snipex tok eval --tokenizer artifacts/tokenizer "Hello SnipeX"
uv run snipex tok eval \
  --tokenizer artifacts/tokenizer \
  --val-data /path/to/climbmix/val

🚀 Pretraining

Use a small model for an MPS smoke test:

uv run snipex llm pretrain \
  --train-data /path/to/climbmix/train \
  --val-data /path/to/climbmix/val \
  --tokenizer artifacts/tokenizer \
  --output runs/mps-smoke \
  --device mps --compile no --attention sdpa \
  --depth 2 --sequence-length 128 --device-batch-size 2 --steps 10

Use one CUDA GPU by running the command normally:

uv run snipex llm pretrain \
  --train-data /path/to/climbmix/train \
  --val-data /path/to/climbmix/val \
  --tokenizer artifacts/tokenizer \
  --output runs/base-single \
  --device cuda --depth 12 --sequence-length 512 \
  --device-batch-size 8 --total-batch-tokens 131072

Use two GPUs with DDP through torchrun:

uv run torchrun --standalone --nproc-per-node=2 -m snipex.cli llm pretrain \
  --train-data /path/to/climbmix/train \
  --val-data /path/to/climbmix/val \
  --tokenizer artifacts/tokenizer \
  --output runs/base-dual \
  --device cuda --depth 12 --sequence-length 512 \
  --device-batch-size 8 --total-batch-tokens 131072

--device-batch-size is sequences per GPU. --total-batch-tokens is the desired global token count per optimizer step across every GPU and accumulation step; SnipeX reports the derived accumulation count and rejects non-divisible values. Omit both total batch and accumulation options to use one micro-batch per step. Omit --steps to derive the run length from the parameter/data ratio. Each data option names a directory; SnipeX recursively reads every supported file in that directory. During training, train updates the model and val provides periodic and final validation metrics. The test directory is not read.

Training writes run.json, TCurve CSV/plots under metrics/, periodic checkpoints when requested, and an always-present final.pt. Resume with:

uv run snipex llm pretrain <same-data-and-output-options> \
  --resume runs/base-dual/step_001000.pt

Evaluate or generate from a base checkpoint:

uv run snipex llm eval \
  --checkpoint runs/base-dual/final.pt \
  --test-data /path/to/climbmix/test \
  --tokenizer artifacts/tokenizer

uv run snipex llm gen \
  --checkpoint runs/base-dual/final.pt \
  --tokenizer artifacts/tokenizer \
  --prompt "The purpose of a prototype is"

💬 SFT and chat

SFT accepts Parquet or JSONL with the standard messages structure used by MS-SWIFT. JSONL contains one conversation per line:

{"messages": [{"role": "user", "content": "Hello"}, {"role": "assistant", "content": "Hi"}]}

Only assistant content and its end token contribute to the loss. V0.1.0 data-prep does not convert Alpaca, query/response, ShareGPT, tool, or multimodal schemas.

Normalize one directory of Parquet files at a time. data-prep does not infer repository layouts or splits; every source Parquet is converted to an output Parquet with the same filename:

uv run snipex llm data-prep \
  --dataset mmlu \
  --input /path/to/raw-mmlu-train \
  --output sft-data/train \
  --tokenizer artifacts/tokenizer \
  --sequence-length 512

uv run snipex llm data-prep \
  --dataset mmlu \
  --input /path/to/raw-mmlu-val \
  --output sft-data/val \
  --tokenizer artifacts/tokenizer \
  --sequence-length 512

--tokenizer and --sequence-length must be supplied together. With both, rows whose complete rendered conversation exceeds the model window are filtered. With neither, data-prep only normalizes the schema and SFT will reject any over-length row rather than truncate it.

The row converters are public Python functions. A custom converter with the same row -> {"messages": [...]} contract can be passed directly to the loader:

from snipex.llm.sft_data import iter_sft_batches
from snipex.llm.sft_prep import prep_mmlu

batches = iter_sft_batches(..., preprocess=prep_mmlu)
uv run snipex llm sft \
  --train-data smoltalk/train mmlu/train gsm8k/train \
  --val-data smoltalk/val mmlu/val gsm8k/val \
  --tokenizer artifacts/tokenizer \
  --checkpoint runs/base-dual/final.pt \
  --output runs/sft \
  --total-batch-tokens 32768 \
  --max-examples 10000

uv run snipex llm chat \
  --checkpoint runs/sft/final.pt \
  --tokenizer artifacts/tokenizer

Generation is deliberately simple in V0.1.0: temperature/top-k autoregressive sampling without a KV cache.

🛠️ Development

Create the environment and run the local checks:

uv sync --extra llm
uv run ruff check .
uv run pytest -m "not gpu"
uv build

Use algo/<name> for an algorithm branch. Put its package under snipex/<name>/, its tests under tests/<name>/, and declare only its required dependencies in a matching extra in pyproject.toml:

[project.optional-dependencies]
example = ["example-dependency>=1"]

Install or test that algorithm with:

uv sync --extra example

Do not add an extra until a real algorithm needs it.


❤️ Support

If SnipeX helps you turn an idea into a working prototype, consider giving it a star on GitHub.

SnipeX is released under the MIT License.

Metadata

Release files for snipex 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for snipex 0.1.0
File Size Uploaded
snipex-0.1.0.tar.gz 43.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for snipex 0.1.0
File Interpreter ABI Platform
snipex-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 89.2 kB

Release files / snipex-0.1.0.tar.gz

Download URL snipex-0.1.0.tar.gz
Size 43.7 kB
Tags Source
SHA-256 checksum
How to use checksums
7faf1d73eccd3f4aa2fbb6bbf4b603545f90e10ad26721f1e3428b04eaa82318
BLAKE2b-256 checksum
How to use checksums
016c658210afb8b169f403803bb5cbdefe98e5ab51977a915343759482c104dd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.10.12

Release files / snipex-0.1.0-py3-none-any.whl

Download URL snipex-0.1.0-py3-none-any.whl
Size 45.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e400eefad1eebc0bf498220d859b185b38324b2be5a83f7b9705f1a8bdf39c13
BLAKE2b-256 checksum
How to use checksums
30ffd05267b16009f43c070344fa6fac35fc1d977a2863dcd765346e30f40700
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.10.12

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page