Skip to main content

tailcyclenet

Finetune a tracktail point tracker into an animal pose estimator. Three settings, one model: 3D multiview, 3D single-view, 2D single-view.

The pipeline detects animals, crops them, and decodes per-keypoint poses through a single window loop. It reads one annotation format (docs/annotation_format.md) that serves both hand annotation and bulk training across datasets of differing keypoint sets, camera counts, and dimensionality. This README is the committed reference: what this is, how to run it, and the invariants a contributor must not break.


Setup

From a checkout (contributors)

Dependencies are managed with pixi. posetail==0.4.1 is pinned from PyPI.

pixi install
pixi run python -c "import posetail, tailcyclenet"   # sanity check
pixi run test                                        # test suite
pixi run lint                                        # ruff
  • The LD_LIBRARY_PATH prepend in pyproject.toml is load-bearing — the env ships a newer libstdc++ than some hosts, and without it import scipy.optimize dies naming only CXXABI.
  • posetail >= 0.4.1 ships every behaviour this repo once monkeypatched (per-frame camera offsets, crop_box_for_points, scene_features=/input_size= on the tracker forward); there is no patch layer anymore.

From pip (training on your own machine)

tailcyclenet is on PyPI, but two more things are needed before it can train — read both notes, they are not optional extras:

# 1. torch, with the CUDA build matching your host -- the plain PyPI wheel is often CPU-only
#    or the wrong CUDA version. cu128 matches what this repo's own pixi env pins; use whatever
#    matches your driver.
pip install torch --index-url https://download.pytorch.org/whl/cu128

# 2. tailcyclenet itself
pip install tailcyclenet

# 3. aniposelib's PYTORCH branch, not the plain PyPI release posetail's own dependency pulls in.
#    tailcyclenet.format/crop.py/the detector's cross-view association all construct
#    aniposelib.cameras.Camera and need the pytorch branch's nn.Module version (GPU
#    projection/triangulation/Jacobians) for training and inference in EVERY setting -- 2D,
#    3D single-view, 3D multiview -- not just multiview. This is a live dependency on
#    lambdaloop/anipose-lib's `pytorch` branch continuing to exist at this URL, not a pinned
#    release; a known fragility, not an oversight.
pip install "aniposelib @ git+https://github.com/lambdaloop/anipose-lib.git@pytorch"

Then:

tailcyclenet train --data /path/to/your/dataset-root

configs/base.toml (2D and 3D alike) ships inside the package, so --config is optional and defaults to the shipped recipe unmodified; --data must already be in the format docs/annotation_format.md specifies. tailcyclenet train-detector, tailcyclenet infer, and tailcyclenet eval are the pip equivalents of the scripts/*.py invocations documented below — same flags, same defaults, no repo checkout required. scripts/convert_*.py (turning a labelling tool's own export into that format) stay repo-only; there is no pip-installed onboarding path for unconverted data yet.

A COCO-pretrained detector backbone ([model].pretrained = "coco", not the shipped default) is fetched once from a tagged Megvii GitHub release and cached at ~/.cache/tailcyclenet/weights/ (override with --weights-dir or $TAILCYCLENET_CACHE_DIR) — pre-populate that directory yourself on a host with no internet access.

A CPU-only torch install still imports and runs the test suite / tiny smoke checks, just not real training — don't install CUDA torch expecting speed you then don't need for that.


Layout

tailcyclenet/     library: format, dataset, crop rule, model, inference, metrics, detector
                  train.py/train_detector.py/eval.py are the CLI bodies (`tailcyclenet <cmd>`)
scripts/          train.py  train_detector.py  infer.py  eval.py  convert_*.py -- thin
                  dispatchers into tailcyclenet/ for train/train_detector/infer/eval; the
                  convert_*.py/combine_roots.py/render*.py family stays repo-only
configs/          base.toml + detector.toml + detector/<root>.toml (every config layers over its
                  family's base); shipped inside the pip package too (`tailcyclenet.configs`)
configs/datasets/ per-dataset keypoint and skeleton definitions
docs/             annotation_format.md — the data format spec (human-owned)
tests/            invariants (crop rule, converters, geometry)

Training

One estimator trains across every dataset root under [data].path; a keypoint embedding table is what lets roots with different keypoint sets share a model.

# 3D (multiview / single-view) or 2D (single-view); prob_2d_only adds a train-time true-2D
# image-plane path on 3D sessions, so one config serves both
pixi run python scripts/train.py --config configs/base.toml --data <root>

# one node, N gpus: one item per rank, gradients averaged by DDP
pixi run python scripts/train.py --config configs/base.toml --data <root> --devices 4

# pip install (see Setup): --config is optional, defaults to the packaged base.toml
tailcyclenet train --data <root> [--devices 4]

Facts that are easy to get wrong:

  • cams_to_sample and val_cams_to_sample control camera count; prob_2d_only is a train-only true-2D image-plane draw on 3D sessions, with stored 2D labels preferred and 3D projection as fallback. n_keypoints is derived from the data, never configured.
  • The per-rank batch is structurally 1; --devices N is the only batch dimension this repo has. Every iteration count in a config is a total across ranks (60,000 is 60,000 samples on any gpu count) and the learning rate is scaled by sqrt(N), so a multi-gpu run is two levers off a single-gpu one; provenance.toml records which it was.
  • [model].gridresid_offset has no default and must be stated — the two values load the same tensors, so a mismatch produces numbers rather than an exception.
  • A run folder writes keypoint_registry.toml (the derived keypoint axis) and provenance.toml (commit + dirty flag). A config is not a provenance record.
  • The video encoder unfreezes mid-run per [model].video_encoder_requires_grad (a bool, or an int iteration to unfreeze at). A run started before the shipped default (8 blocks at 10,000) is not comparable to one after; false restores the old arm.

Detector

pixi run python scripts/train_detector.py --config configs/detector.toml

# pip install: --config is optional, defaults to the packaged detector.toml
tailcyclenet train-detector --out runs/det-<name>

The recipe lives in the config, not on the CLI — every default is there with its evidence, and an unknown key raises rather than silently training at a default. Only --out, --iters and --device override.

One detector per dataset, and input_wh defaults to an aspect-matched size rather than a square: a square letterbox on a wide frame wastes most of the canvas and can put the animal below the stride the FPN can represent. The regression target is crop.crop_box_for_points — the detector reproduces the crop the pose model was trained on, so [data].boxes must equal the pose run's [data].box_source.


Inference and eval

# one source session (a dataset root works only if it holds a single session in --split)
pixi run python scripts/infer.py --run runs/<name> --data <session-dir> --split test \
    --detector runs/det-<name> --out pred/

# or, straight off raw footage + an anipose calibration
pixi run python scripts/infer.py --run runs/<name> --out pred/ \
    --videos rec/ --calibration anipose/calibration.toml --cam-regex 'cam([0-9]+)_' \
    --detector runs/det-<name> --max-animals 4

pixi run python scripts/eval.py pred/ --data <root> --split test --chunk 500

# pip install: `tailcyclenet infer`/`tailcyclenet eval` take identical flags
tailcyclenet infer --run runs/<name> --data <session-dir> --split test \
    --detector runs/det-<name> --out pred/
tailcyclenet eval pred/ --data <root> --split test --chunk 500
  • --out is a prediction session directory (session.toml, calibration.toml, groups.pq, points3d.pq, keypoints.pq, instances.pq, windows.pq), written a block at a time so nothing is proportional to clip length. eval.py and render.py both read it; render.py finds its own pixels via the session's [provenance].
  • --data and --videos are exactly-one-of, and a run is one source session (which may hold many groups). For --videos, the camera name is the regex capture group and the session is built in memory — nothing is staged.
  • There is one window loop. Box sources: annotations, a detections npz (--boxes), or a per-dataset detector (--detector). Prompt regimes: none (query-free), carry (previous window's own prediction — what deployment does), self (two passes), labels (an oracle, gated off by default).

Defaults are not the recommendation. The good settings are root-conditional, so sweep them per root. Current values: --anchor carry, --overlap 4, --refine derived (on 3D / off 2D), --track on, --box-prompt auto, --prefetch-windows 1 (bit-exact, performance only), --max-ram derived from the host. In particular --anchor is root-conditional in 2D (a carried prior on a crowded root is often the wrong animal's pose) and --overlap's optimum is seam-count against seam-size — sweep per root.

Four rules that are not root-conditional:

  • --vis-thresh has no meaning in 2D at the shipped default — the visibility head is only trained when [training.losses].vis_loss_2d_weight is nonzero (default 0.0).
  • --box-prompt auto needs a detector or boxes file. A box-model run without one refuses rather than silently falling back to the GT oracle. Pass --detector/--boxes, or --box-prompt none to withhold the box.
  • Always run eval.py with --chunk 500 on long clips — the bootstrap resamples groups, so a single long clip returns DEGENERATE.
  • Use --min-match-kpts 0.5 for deltas and 0 for absolutes.

The largest lever is not a flag: pose accuracy on a ground-truth crop is far better, at full coverage, than through the detector, and on a long clip essentially all coverage loss is no box. Fix the crop path before tuning identity flags.


Reference

  • docs/annotation_format.md — the data format spec (human-owned).

Metadata

Release files for tailcyclenet 0.3.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tailcyclenet 0.3.1
File Size Uploaded
tailcyclenet-0.3.1.tar.gz 721.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tailcyclenet 0.3.1
File Interpreter ABI Platform
tailcyclenet-0.3.1-py3-none-any.whl Python 3 none any Details

Total release size: 1.2 MB

Release files / tailcyclenet-0.3.1.tar.gz

Download URL tailcyclenet-0.3.1.tar.gz
Size 721.0 kB
Tags Source
SHA-256 checksum
How to use checksums
094396742776047aa9922f468f274519a406b5a7db0a5da7be200c9cc77557b1
BLAKE2b-256 checksum
How to use checksums
9346ab24e164c75cd1bef5368f2971d3cf85aafa646f4ca403ee0d6104db701c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.0

Release files / tailcyclenet-0.3.1-py3-none-any.whl

Download URL tailcyclenet-0.3.1-py3-none-any.whl
Size 462.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f637601b9daa656cd186b43ab03793124a6966cd0943e6cd11a0cb2f0a3577de
BLAKE2b-256 checksum
How to use checksums
901e5fcbff02887528f9c3d07fbca5ce655e3b865a125cad0bf923856d4ad231
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.0

Release history Release notifications | RSS feed

0.3.2

2 release files

This release

0.3.1 This release

2 release files

0.3.0

2 release files

0.2.7

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.0

2 release files

0.1.10

2 release files

0.1.9

2 release files

0.1.8

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

0.0.8

2 release files

0.0.7

2 release files

0.0.6

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.3

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page