SnipeX
Single-GPU, oNe-week Idea-to-Prototype EXecution.
SnipeX is a collection of small AI algorithm prototypes. Algorithm-specific dependencies are provided through optional package extras. Every algorithm prototype in SnipeX is designed to run on a single consumer GPU and reach a working result within one week.
⚡ Quick start
Installing the base package provides the snipex command, package version, and
help for the available algorithm families without installing PyTorch:
pip install snipex
snipex --help
🧠 LLM prototype (V0.1.0)
The first prototype is a NanoChat-inspired decoder-only language model with RoPE, RMSNorm and QK norm, ReLU² MLPs, document-aware ClimbMix packing, AdamW with linear warmup and cosine decay, MPS/CUDA support, and CUDA DDP.
Install the LLM prototype from PyPI:
pip install "snipex[llm]"
From a source checkout, create the development environment with:
uv sync --extra llm
FA3 is optional and CUDA-only. auto falls back to PyTorch SDPA when the
community kernel cannot be loaded. It is an implementation optimization, not
an algorithm extra, so source checkouts install it through a dependency group:
uv sync --extra llm --group flash
PyPI users who explicitly want FA3 can install snipex[llm] and kernels
separately. The normal snipex[llm] installation uses SDPA without kernels.
📚 Data and tokenizer
Keep every split in a separate directory. SnipeX never infers a split from a filename and never reads the test directory during tokenizer or model training:
/path/to/climbmix/
├── train/*.{parquet,jsonl}
├── val/*.{parquet,jsonl}
└── test/*.{parquet,jsonl}
Base-model JSONL uses one document per line:
{"text": "One pretraining document."}
The equivalent Parquet schema has one string column named text. Mixing
Parquet and JSONL files inside one split is supported.
Train and inspect the 32K RustBPE tokenizer:
uv run snipex tok train \
--train-data /path/to/climbmix/train \
--out-dir artifacts/tokenizer
uv run snipex tok encode --tokenizer artifacts/tokenizer "Hello SnipeX"
uv run snipex tok decode --tokenizer artifacts/tokenizer 123 456
uv run snipex tok eval --tokenizer artifacts/tokenizer "Hello SnipeX"
uv run snipex tok eval \
--tokenizer artifacts/tokenizer \
--val-data /path/to/climbmix/val
🚀 Pretraining
Use a small model for an MPS smoke test:
uv run snipex llm pretrain \
--train-data /path/to/climbmix/train \
--val-data /path/to/climbmix/val \
--tokenizer artifacts/tokenizer \
--output runs/mps-smoke \
--device mps --compile no --attention sdpa \
--depth 2 --sequence-length 128 --device-batch-size 2 --steps 10
Use one CUDA GPU by running the command normally:
uv run snipex llm pretrain \
--train-data /path/to/climbmix/train \
--val-data /path/to/climbmix/val \
--tokenizer artifacts/tokenizer \
--output runs/base-single \
--device cuda --depth 12 --sequence-length 512 \
--device-batch-size 8 --total-batch-tokens 131072
Use two GPUs with DDP through torchrun:
uv run torchrun --standalone --nproc-per-node=2 -m snipex.cli llm pretrain \
--train-data /path/to/climbmix/train \
--val-data /path/to/climbmix/val \
--tokenizer artifacts/tokenizer \
--output runs/base-dual \
--device cuda --depth 12 --sequence-length 512 \
--device-batch-size 8 --total-batch-tokens 131072
--device-batch-size is sequences per GPU. --total-batch-tokens is the desired
global token count per optimizer step across every GPU and accumulation step;
SnipeX reports the derived accumulation count and rejects non-divisible values.
Omit both total batch and accumulation options to use one micro-batch per step.
Omit --steps to derive the run length from the parameter/data ratio.
Each data option names a directory; SnipeX recursively reads every supported
file in that directory. During training, train updates the model and val
provides periodic and final validation metrics. The test directory is not read.
Training writes run.json, TCurve CSV/plots under metrics/, periodic
checkpoints when requested, and an always-present final.pt. Resume with:
uv run snipex llm pretrain <same-data-and-output-options> \
--resume runs/base-dual/step_001000.pt
Evaluate or generate from a base checkpoint:
uv run snipex llm eval \
--checkpoint runs/base-dual/final.pt \
--test-data /path/to/climbmix/test \
--tokenizer artifacts/tokenizer
uv run snipex llm gen \
--checkpoint runs/base-dual/final.pt \
--tokenizer artifacts/tokenizer \
--prompt "The purpose of a prototype is"
💬 SFT and chat
SFT accepts Parquet or JSONL with the standard messages structure used by
MS-SWIFT. JSONL contains one conversation per line:
{"messages": [{"role": "user", "content": "Hello"}, {"role": "assistant", "content": "Hi"}]}
Only assistant content and its end token contribute to the loss. V0.1.0
data-prep does not convert Alpaca, query/response, ShareGPT, tool, or
multimodal schemas.
Normalize one directory of Parquet files at a time. data-prep does not infer
repository layouts or splits; every source Parquet is converted to an output
Parquet with the same filename:
uv run snipex llm data-prep \
--dataset mmlu \
--input /path/to/raw-mmlu-train \
--output sft-data/train \
--tokenizer artifacts/tokenizer \
--sequence-length 512
uv run snipex llm data-prep \
--dataset mmlu \
--input /path/to/raw-mmlu-val \
--output sft-data/val \
--tokenizer artifacts/tokenizer \
--sequence-length 512
--tokenizer and --sequence-length must be supplied together. With both,
rows whose complete rendered conversation exceeds the model window are
filtered. With neither, data-prep only normalizes the schema and SFT will
reject any over-length row rather than truncate it.
The row converters are public Python functions. A custom converter with the
same row -> {"messages": [...]} contract can be passed directly to the loader:
from snipex.llm.sft_data import iter_sft_batches
from snipex.llm.sft_prep import prep_mmlu
batches = iter_sft_batches(..., preprocess=prep_mmlu)
uv run snipex llm sft \
--train-data smoltalk/train mmlu/train gsm8k/train \
--val-data smoltalk/val mmlu/val gsm8k/val \
--tokenizer artifacts/tokenizer \
--checkpoint runs/base-dual/final.pt \
--output runs/sft \
--total-batch-tokens 32768 \
--max-examples 10000
uv run snipex llm chat \
--checkpoint runs/sft/final.pt \
--tokenizer artifacts/tokenizer
Generation is deliberately simple in V0.1.0: temperature/top-k autoregressive sampling without a KV cache.
🛠️ Development
Create the environment and run the local checks:
uv sync --extra llm
uv run ruff check .
uv run pytest -m "not gpu"
uv build
Use algo/<name> for an algorithm branch. Put its package under
snipex/<name>/, its tests under tests/<name>/, and declare only its required
dependencies in a matching extra in pyproject.toml:
[project.optional-dependencies]
example = ["example-dependency>=1"]
Install or test that algorithm with:
uv sync --extra example
Do not add an extra until a real algorithm needs it.
❤️ Support
If SnipeX helps you turn an idea into a working prototype, consider giving it a star on GitHub.
SnipeX is released under the MIT License.
Metadata
Release files for snipex 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| snipex-0.1.0.tar.gz | 43.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| snipex-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 89.2 kB
Release files / snipex-0.1.0.tar.gz
| Download URL | snipex-0.1.0.tar.gz |
|---|---|
| Size | 43.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7faf1d73eccd3f4aa2fbb6bbf4b603545f90e10ad26721f1e3428b04eaa82318
|
|
BLAKE2b-256 checksum How to use checksums |
016c658210afb8b169f403803bb5cbdefe98e5ab51977a915343759482c104dd
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.10.12
|
Release files / snipex-0.1.0-py3-none-any.whl
| Download URL | snipex-0.1.0-py3-none-any.whl |
|---|---|
| Size | 45.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e400eefad1eebc0bf498220d859b185b38324b2be5a83f7b9705f1a8bdf39c13
|
|
BLAKE2b-256 checksum How to use checksums |
30ffd05267b16009f43c070344fa6fac35fc1d977a2863dcd765346e30f40700
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.10.12
|