Skip to main content

tailcyclenet

Finetune a posetail point tracker into an animal pose estimator. Three settings, one model: 3D multiview, 3D single-view, 2D single-view.

The pipeline detects animals, crops them, and decodes per-keypoint poses through a single window loop. It reads one annotation format (docs/annotation_format.md) that serves both hand annotation and bulk training across datasets of differing keypoint sets, camera counts, and dimensionality. This README is the committed reference: what this is, how to run it, and the invariants a contributor must not break.


Setup

Dependencies are managed with pixi. posetail==0.3.5 is pinned from PyPI.

pixi install
pixi run python -c "import posetail, tailcyclenet"   # sanity check
pixi run test                                        # test suite
pixi run lint                                        # ruff
  • The LD_LIBRARY_PATH prepend in pyproject.toml is load-bearing — the env ships a newer libstdc++ than some hosts, and without it import scipy.optimize dies naming only CXXABI.
  • posetail >= 0.3.5 ships every behaviour this repo once monkeypatched (per-frame camera offsets, crop_box_for_points, scene_features=/input_size= on the tracker forward); there is no patch layer anymore.

Layout

tailcyclenet/     library: format, dataset, crop rule, model, inference, metrics, detector
scripts/          train.py  train_detector.py  infer.py  eval.py  convert_*.py
configs/          base.toml + detector.toml + detector/<root>.toml (every config layers over its family's base)
configs/datasets/ per-dataset keypoint and skeleton definitions
docs/             annotation_format.md — the data format spec (human-owned)
tests/            invariants (crop rule, converters, geometry)

Training

One estimator trains across every dataset root under [data].path; a keypoint embedding table is what lets roots with different keypoint sets share a model.

# 3D (multiview / single-view) or 2D (single-view) -- the camera-count keys in base.toml are
# harmless on a one-camera root, so one config serves both
pixi run python scripts/train.py --config configs/base.toml --data <root>

# one node, N gpus: one item per rank, gradients averaged by DDP
pixi run python scripts/train.py --config configs/base.toml --data <root> --devices 4

Facts that are easy to get wrong:

  • The two configs differ in exactly three keys — cams_to_sample, val_cams_to_sample, prob_2d_only — all camera-count questions a one-camera root cannot ask. n_keypoints is derived from the data, never configured.
  • The per-rank batch is structurally 1; --devices N is the only batch dimension this repo has. Every iteration count in a config is a total across ranks (60,000 is 60,000 samples on any gpu count) and the learning rate is scaled by sqrt(N), so a multi-gpu run is two levers off a single-gpu one; provenance.toml records which it was.
  • [model].gridresid_offset has no default and must be stated — the two values load the same tensors, so a mismatch produces numbers rather than an exception.
  • A run folder writes keypoint_registry.toml (the derived keypoint axis) and provenance.toml (commit + dirty flag). A config is not a provenance record.
  • The video encoder unfreezes mid-run per [model].video_encoder_requires_grad (a bool, or an int iteration to unfreeze at). A run started before the shipped default (8 blocks at 10,000) is not comparable to one after; false restores the old arm.

Detector

pixi run python scripts/train_detector.py --config configs/detector.toml

The recipe lives in the config, not on the CLI — every default is there with its evidence, and an unknown key raises rather than silently training at a default. Only --out, --iters and --device override.

One detector per dataset, and input_wh defaults to an aspect-matched size rather than a square: a square letterbox on a wide frame wastes most of the canvas and can put the animal below the stride the FPN can represent. The regression target is crop.crop_box_for_points — the detector reproduces the crop the pose model was trained on, so [data].boxes must equal the pose run's [data].box_source.


Inference and eval

# one source session (a dataset root works only if it holds a single session in --split)
pixi run python scripts/infer.py --run runs/<name> --data <session-dir> --split test \
    --detector runs/det-<name> --out pred/

# or, straight off raw footage + an anipose calibration
pixi run python scripts/infer.py --run runs/<name> --out pred/ \
    --videos rec/ --calibration anipose/calibration.toml --cam-regex 'cam([0-9]+)_' \
    --detector runs/det-<name> --max-animals 4

pixi run python scripts/eval.py pred/ --data <root> --split test --chunk 500
  • --out is a prediction session directory (session.toml, calibration.toml, groups.pq, points3d.pq, keypoints.pq, instances.pq, windows.pq), written a block at a time so nothing is proportional to clip length. eval.py and render.py both read it; render.py finds its own pixels via the session's [provenance].
  • --data and --videos are exactly-one-of, and a run is one source session (which may hold many groups). For --videos, the camera name is the regex capture group and the session is built in memory — nothing is staged.
  • There is one window loop. Box sources: annotations, a detections npz (--boxes), or a per-dataset detector (--detector). Prompt regimes: none (query-free), carry (previous window's own prediction — what deployment does), self (two passes), labels (an oracle, gated off by default).

Defaults are not the recommendation. The good settings are root-conditional, so sweep them per root. Current values: --anchor carry, --overlap 4, --refine derived (on 3D / off 2D), --track on, --box-prompt auto, --prefetch-windows 1 (bit-exact, performance only), --max-ram derived from the host. In particular --anchor is root-conditional in 2D (a carried prior on a crowded root is often the wrong animal's pose) and --overlap's optimum is seam-count against seam-size — sweep per root.

Four rules that are not root-conditional:

  • --vis-thresh has no meaning in 2D at the shipped default — the visibility head is only trained when [training.losses].vis_loss_2d_weight is nonzero (default 0.0).
  • --box-prompt auto needs a detector or boxes file. A box-model run without one refuses rather than silently falling back to the GT oracle. Pass --detector/--boxes, or --box-prompt none to withhold the box.
  • Always run eval.py with --chunk 500 on long clips — the bootstrap resamples groups, so a single long clip returns DEGENERATE.
  • Use --min-match-kpts 0.5 for deltas and 0 for absolutes.

The largest lever is not a flag: pose accuracy on a ground-truth crop is far better, at full coverage, than through the detector, and on a long clip essentially all coverage loss is no box. Fix the crop path before tuning identity flags.


Reference

  • docs/annotation_format.md — the data format spec (human-owned).

Metadata

Release files for tailcyclenet 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tailcyclenet 0.1.1
File Size Uploaded
tailcyclenet-0.1.1.tar.gz 391.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tailcyclenet 0.1.1
File Interpreter ABI Platform
tailcyclenet-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 624.3 kB

Release files / tailcyclenet-0.1.1.tar.gz

Download URL tailcyclenet-0.1.1.tar.gz
Size 391.6 kB
Tags Source
SHA-256 checksum
How to use checksums
aabce0b40329fa01eb9c01f17bbf36dc5b87cbd70ac0569aba5bb7c380eb735e
BLAKE2b-256 checksum
How to use checksums
8ff49bfcea0c150027684a6e7c0eae883d96bd9e0a73bf23fb4ef4d07afeb63a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.15

Release files / tailcyclenet-0.1.1-py3-none-any.whl

Download URL tailcyclenet-0.1.1-py3-none-any.whl
Size 232.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
418a93fb80fc50452266f522ed680f3ac65d464a796c5ff0cec62327ae3d3389
BLAKE2b-256 checksum
How to use checksums
ea4a1afbea45f88d14c12e6b8ba9f98fad058f48185c87b699de327208576866
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.15

Release history Release notifications | RSS feed

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.7

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.0

2 release files

0.1.10

2 release files

0.1.9

2 release files

0.1.8

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

This release

0.1.1 This release

2 release files

0.1.0

2 release files

0.0.8

2 release files

0.0.7

2 release files

0.0.6

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.3

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page