Skip to main content

aphrody-train

Training surface for aphrody (P10, docs/plans/codex-oss.md):

  • triples: Generates dataset triples (query, positive, negative) from distinct source chunks, with exact-hash gold-leak gates, optional MinHash near-duplicate clustering before split assignment and mined hard negatives (see below). A supplied --gold-file must be readable and nonempty.
  • embed: Fine-tunes dense text embeddings with MultipleNegativesRankingLoss / Matryoshka and exports to ONNX.
  • rerank: Fine-tunes cross-encoder rerankers and exports to ONNX.
  • lora: Fine-tunes a PEFT LoRA adapter from guarded dataset rollouts output, then converts it to a verified GGUF adapter with llama.cpp. Explicit NF4 QLoRA is available for qualified CUDA training hosts; DPO training is not implemented.
  • dry_run: Runs actual CPU all-MiniLM training on 50 deterministic probe triples and verifies its ONNX export; it requires the full training dependencies.
  • dataset rollouts|dpo: Exports redacted JSONL from Rust rollout records. Successful turns require a final message and successful tool results. DPO pairs require a failed/denied call followed by a successful same-tool repair in one turn. Both require --gold-file; rows that hash to a gold query, sentence, title or URL are excluded, rows gain text_sha256 of the prompt, and identical content from different sessions is kept once.
  • hashing gold: Writes a copy of a gold JSONL with additive text_sha256 / target_sha256 fields for registry anti-joins.
  • protocol: Emits streaming JSONL events (start, progress, eval, checkpoint, artifact, done, error) to stdout.

Install --extra full before embed or rerank. Both commands require a nonempty triples JSONL file, train a checkpoint, export ONNX and verify its output against the trained model before emitting artifact and done.

Run these commands from the selected Aphrody py/ workspace. Set APHRODY_PY_ROOT when the deployed workspace uses another declared path.

LoRA preserves the full precision base model by default. It requires a local Hugging Face base model, a prepared directory containing rollouts.jsonl from dataset rollouts, the untouched gold JSONL, and an official llama.cpp checkout. Set APHRODY_LLAMA_CPP_CONVERTER to its convert_lora_to_gguf.py path, then run with uv run --package aphrody-train --extra full python -m aphrody_train.lora --rollout-dir <prepared-dir> --gold-file <gold.jsonl> --base-model <local-model-dir> --output-dir <output-dir>. The output adapter.gguf is a LoRA adapter and needs the matching base model in GGUF at inference time; it is not a standalone merged model. The trainer rechecks gold overlap, redacts known secret forms, and verifies nonzero finite adapter weights and GGUF tensor pairs before reporting success. Model quality still requires an independent evaluation before promotion.

The official converter runs with a qualified CPython entry point. Standalone Python uses its own sys.executable; native Bun embedding reports Bun there, so set APHRODY_LORA_CONVERTER_PYTHON (or --converter-python) to the absolute Python executable from the locked training environment. This selection is validated before loading the model. It does not replace the shared interpreter used for training or change another SDK consumer's sys.executable.

For an exclusive GPU training job, --quantization nf4 (or APHRODY_LORA_QUANTIZATION=nf4) loads the same local Hugging Face base through bitsandbytes NF4 with double quantization, prepares it with PEFT's prepare_model_for_kbit_training, and uses a paged 8-bit optimizer. The qualified GPU environment installs the locked full and qlora extras (bitsandbytes==0.50.0); this option fails if CUDA or that dependency is absent. It never silently trains on CPU or downloads another base. The host scheduler must acquire an exclusive GPU lease and suspend its inference allocation first. An 8 GiB GPU is a candidate to qualify, not a memory capacity guarantee.

Completed output directories are refused. Each successful export adds a private training-manifest.json with base config, dataset, untouched gold and adapter SHA-256, quantization mode and step count. promoted remains false; training does not change the active model. Compare the candidate against the current Shenron adapter on untouched gold before promotion.

Implementation follows the Transformers bitsandbytes guide and PEFT quantization guide.

The source-only training container is declared in tools/config/container/gpu-fleet/Dockerfile.train. It requires an immutable, already qualified native Bun/Buv/PyJS/CPython/Aphrody-train runtime image and consumes py/uv.lock with --locked --extra full --extra qlora. Builds belong to the declared VPS factory; datasets, model weights, provider stores and run output are separate protected runtime mounts. A source Dockerfile or a passing metadata check does not qualify an Omar GPU training run.

The same trainer is callable through the native Python SDK as sdk.invoke('aphrody_train.lora', 'run_lora_job', [request]). The JSON request names absolute rollout_dir, gold_file, base_model, converter and output_dir paths, plus job_id, positive integer steps and explicit quantization (none or nf4). The function emits the existing job protocol, returns the verified export manifest only after success, and keeps the same gold, secret, converter and completed-output guards. The host scheduler must hold the exclusive GPU allocation for the entire invocation; runtime/package qualification remains separate from this source adapter. The optional converter_python names that same qualified absolute CPython entry point; otherwise the deployed environment supplies it.

CPU receipt of 2026-10-09: Bun 1.4.3-aphrody.4 (41211b568) and CPython 3.12.15 execute the SDK job in the same PID, reject gold overlap before loading Torch, and select the existing training virtualenv's CPython entry point for conversion. Its separate CPU probe reports that virtualenv prefix and CPython 3.12.15. This evidence covers the native boundary and interpreter selection; it does not qualify GPU training or a GGUF conversion.

Dataset extraction is available as uv run --package aphrody-train python -m aphrody_train.dataset rollouts --rollout-dir <dir> --gold-file <gold.jsonl> --output-dir <dir> (or dpo). Redaction covers credential fields and common token forms; review generated datasets before external publication because arbitrary secrets cannot be recognized by patterns alone.

Triples, leak and dedup gates

python -m aphrody_train.triples runs, in order (D11 of docs/decisions/infra/sql-convergence.md in the Aphrody monorepo; background in docs/research/ai/sqllm/training.md there):

  1. Hashing. aphrody_train.hashing normalises text with NFC, collapses every run of Unicode whitespace to one space and strips it, then takes the SHA-256 hex of the UTF-8 bytes. casefold=True folds case first. Emitted rows carry text_sha256 (no casefold) of their anchor text, the triple query or the rollout prompt; gold rows hashed with hashing gold use the same rule on the gold query, so the registry leak check is JOIN gold USING (text_sha256). Rows also carry text_casefold_sha256.
  2. Exact dedup. Chunks whose casefolded passage hash was already seen are dropped (exact_duplicates).
  3. Gold anti-join (aphrody_train.gates.GoldIndex). In memory, a row leaks when one of its keys is a gold key. Keys are the casefolded hashes of the whole query, passage and title, of each sentence or line of at least 12 characters, and of each URL (fragment and trailing slash removed) plus its path. Gold queries, titles and URLs are keyed the same way. Leaked chunks are not reused as negatives. --legacy-substring-guard adds the former substring heuristic on top; it is off by default.
  4. Near duplicates. --near-dup auto|on|off (default auto: on when datasketch is installed, otherwise reported as unavailable) clusters passages with MinHash LSH: word 5-gram shingles, --minhash-perm 128, --near-dup-threshold 0.8 (Jaccard; at 128 permutations LSH accepts up to 0.98). Clustering runs before split assignment. --val-fraction sends whole clusters, chosen by a seeded hash of the cluster id (--split-seed), to triples.val.jsonl. --near-dup-drop keeps one member per cluster.
  5. Negatives. --negatives auto|mined|positional. Mining calls sentence_transformers.util.mine_hard_negatives per split with range_max=50, max_score=0.8, relative_margin=0.05 and sampling_strategy="top". The miner is --miner-model, else --base-model, else intfloat/multilingual-e5-small; E5 models get query: / passage: prompts. Candidates from the same document (doc_id, document_id, source_id, parent_id, doc, else URL, else path) or the same near-duplicate cluster are removed. A row with no remaining candidate gets the positional negative. positional is the former (idx + n/2) % n choice; it moves to the next position when that chunk shares the document or cluster. auto falls back to positional when sentence-transformers, datasets or the model cannot be loaded, and reports the reason. mined fails instead.
  6. Reranker data. With --reranker-model, every training row is also written to triples.flag.jsonl in the FlagEmbedding layout: {"query", "pos", "neg", "pos_scores", "neg_scores", "text_sha256"}. The scores are cross-encoder scores. Candidates that score at or above the positive are dropped as likely false negatives.

Outputs: triples.jsonl keeps the original query, positive, negative and chunk_id fields and adds text_sha256, text_casefold_sha256, positive_sha256, negative_sha256, cluster_id, split, negative_strategy and, for mined rows, negative_score. It holds the training split only. Artifact roles are triples, triples_val and reranker_flag. The done event carries a stats object with every gate count, the negative strategy and the fallback reason.

Extras: dedup installs datasketch 2.0 (MIT). mine and full use the locked sentence-transformers 6.1 (Apache-2.0), datasets and torch for mining and reranker scoring. full also supplies the declared ONNX exporters and PEFT/TRL training stack. Consume these extras through the existing workspace lock rather than installing an independent dependency set.

uv run --package aphrody-train --extra mine python -m aphrody_train.triples \
  --input-chunks chunks.jsonl --gold-file gold.jsonl --output-dir out \
  --negatives mined --val-fraction 0.1 \
  --reranker-model BAAI/bge-reranker-v2-m3

Metadata

Release files for aphrody-train 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for aphrody-train 0.1.0
File Size Uploaded
aphrody_train-0.1.0.tar.gz 110.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for aphrody-train 0.1.0
File Interpreter ABI Platform
aphrody_train-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 166.2 kB

Release files / aphrody_train-0.1.0.tar.gz

Download URL aphrody_train-0.1.0.tar.gz
Size 110.5 kB
Tags Source
SHA-256 checksum
How to use checksums
a7df8e39911830550b70c2c6b61c3a79f6c64a204eb4fa6e5de32a8b104119ea
BLAKE2b-256 checksum
How to use checksums
b6ad46a473f45ab9d5d4cbbdccb8f7a5aebccfa8b2b12ee701c46cdde36a4b7a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.24 {"installer":{"name":"uv","version":"0.12.24","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / aphrody_train-0.1.0-py3-none-any.whl

Download URL aphrody_train-0.1.0-py3-none-any.whl
Size 55.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3eadf139065bf4c843edabf53535537cbfdffe1d21adc19874a459c16c436ac2
BLAKE2b-256 checksum
How to use checksums
f819297bd3d3c620ea7f192f42809d1b13f59da54a1637d0611d71095c580595
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.24 {"installer":{"name":"uv","version":"0.12.24","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page